東京大学 · 情報科学
Tao Li教授の研究室は、動画認識と自己教師付き学習に焦点を当てた先端的研究を推進しています。特に、時間的・空間的情報を効果的に捉えるためのコントラスト学習フレームワーク(IIC)や、残差フレームを用いた効率的な動き特徴抽出手法の開発が特徴です。また、手の義手に触覚フィードバックを提供するための実用的で安価なデバイスの開発にも取り組んでおり、臨床応用に向けた実装的アプローチを重視しています。
Figures are computed from collected data and may differ slightly.
We propose a self-supervised method to learn feature representations from videos. A standard approach in traditional self-supervised methods uses positive-negative data pairs to train with contrastive learning strategy. In such a case, different modalities of the same video are treated as positives and video clips from a different video are treated as negatives. Because the spatio-temporal information is important for video representation, we extend the negative samples by introducing intra-nega
In this paper, we propose a self-supervised contrastive learning method to learn video feature representations. In traditional self-supervised contrastive learning methods, constraints from anchor, positive, and negative data pairs are used to train the model. In such a case, different samplings of the same video are treated as positives, and video clips from different videos are treated as negatives. Because the spatio-temporal information is important for video representation, we set the tempo
Recently, 3D convolutional networks yield good performance in action recognition. However, an optical flow stream is still needed for motion representation to ensure better performance, whose cost is very high. In this paper, we propose a cheap but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 35.6% and 26.6% points improvements over top-1 accuracy can be achieved on
Recently, 3D convolutional networks yield good performance in action recognition. However, optical flow stream is still needed to ensure better performance, the cost of which is very high. In this paper, we propose a fast but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 20.5% and 12.5% points improvements over top-1 accuracy can be achieved on the UCF101 and HMDB51
Abstract Human rely profoundly on tactile feedback from fingertips to interact with the environment, whereas most hand prostheses used in clinics provide no tactile feedback. In this study we demonstrate the feasibility to use a tactile display glove that can be worn by a unilateral hand amputee on the remaining healthy hand to display tactile feedback from a hand prosthesis. The main benefit is that users could easily distinguish the feedback for each finger, even without training. The claimed
Pretext tasks and contrastive learning have been successful in self-supervised learning for video retrieval and recognition. In this study, we analyze their optimization targets and utilize the hyper-sphere feature space to explore the connections between them, indicating the compatibility and consistency of these two different learning methods. Based on the analysis, we propose a self-supervised training method, referred as Pretext-Contrastive Learning (PCL), to learn video representations. Ext
Recently, pretext-task based methods are proposed one after another in self-supervised video feature learning. Meanwhile, contrastive learning methods also yield good performance. Usually, new methods can beat previous ones as claimed that they could capture "better" temporal information. However, there exist setting differences among them and it is hard to conclude which is better. It would be much more convincing in comparison if these methods have reached as closer to their performance limits
Abstract When the traditional collaborative filtering algori- thm is applied to drug recommendation, the recommendation effect is not good due to the sparsity of data. In view of the above problems, this paper proposes a collaborative filtering recommendation algorithm based on user behavior and drug semantics (UBDS-CF). Firstly, we construct the purchasing behavior matrix of users and drugs, and use the weighted cosine similarity to calculate the basic similarity between drugs; then construct t
Abstract Creating impressive video content such as movies and advertisements is a very important yet challenging task in business that requires both a sense of creativity and a lot of experience. Even professionals cannot necessarily invoke the impressions and emotions that they have aimed at. Many video advertisements are created and then disappear without giving a large impact on viewers. This paper presents a large-scale dataset of television (TV) advertisements that consists of 14,490 videos
Recently, 3D convolutional networks (3D ConvNets) yield good performance in action recognition. However, optical flow stream is still needed to ensure better performance, the cost of which is very high. In this paper, we propose a fast but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 35.6% and 26.6% points improvements over top-l accuracy can be obtained on the UCF1
For the current deep learning methods in human movement recognition, there are problems of high number of parameters and insufficient feature extraction., this paper proposes a residual network incorporating improved attention mechanism for human behavior, which introduces an attention mechanism into the residual network and the multilayer vector machine for the channel attention module was replaced using a one-dimensional convolutional, which reduces the feature loss and the parameter number at
To adapt the need of the battle command on the condition of tall technique,based on the analysis of the efficitive factor about the battle command,construct the system of the evaluate index,combine the fuzzy intergrated method,set up the model and evaluate,offered credibility theory foundation about the battle command efficiency.
In this companion paper, we provide details of the artifacts to support the replication of "Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework", which was presented at MM'20. The Inter-intra Contrastive (IIC) framework aims to extract more discriminative temporal information by extending intra-negative samples in contrastive self-supervised learning. In this paper, we first summarize our contribution. Then we explain the file structure of the source code and de
Multiculture Counselling Theory(MCT)was brought out with the impelling of the movement of multiculture.MCT emphasizes on paying more attention to the differences in various culture and differences in the various cultural characters of individuals.It will be referenced in our country because there are abundant conventional culture,many folks in China,and the trend of international communication and national modernization is being strengthened.The introduction of MCT will bring an improvement in t
Open papers in the app to read, cite, and organize with AI.