The University of Tokyo · 컴퓨터과학
마사히로 스즈끼 교수의 연구실은 다중모态 기반의 딥 generative 모델, 특히 변분 오토인코더를 활용한 양방향 다중모态 생성 기술에 초점을 맞추고 있습니다. 이미지와 텍스트 간의 상호 교환 가능한 고수준 표현 학습 및 공유 표현 추출을 통해 비정형적이고 이질적인 데이터 간의 통합적 분석을 구현합니다. 또한 신경망 기반의 전이 학습 기법을 응용해 다양한 작업 간의 지식 전이 효율을 높이는 연구도 진행 중입니다.
표시된 성과는 수집된 데이터 기준으로 산출되며, 일부 차이가 있을 수 있습니다.
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation th
Multimodal learning is a framework for building models that make predictions based on different types of modalities. Important challenges in multimodal learning are the inference of shared representations from arbitrary modalities and cross-modal generation via these representations; however, achieving this requires taking the heterogeneous nature of multimodal data into account. In recent years, deep generative models, i.e. generative models in which distributions are parameterized by deep neur
The adenovirus vector (AdV) can carry two transgenes in its genome, the therapeutic gene and a reporter gene, for example. The E3 insertion site has often been used for the expression of the second transgene. A transgene can be inserted at six different sites/orientations: E1, E3 and E4 sites, and right and left orientations. However, the best combination of the insertion sites and orientations as for the titers and the expression levels has not sufficiently been studied. We attempted to constru
The ability to fine-tune the movement of swallowing-related organs and change the swallowing pattern to fit the volume of a bolus, texture and the physical properties of the food to be swallowed is referred to as the swallowing reserve. In other words, it is the response capability of food swallowing to avoid choking and aspiration. Herein, we focus on the coordination of the suprahyoid and infrahyoid muscles activities, which are closely related to swallowing movement, as a first step to develo
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation th
Machine learning is the basis of important advances in artificial intelligence. Unlike the general methods of machine learning, which use the same tasks for training and testing, the method of transfer learning uses different tasks to learn a new task. Among the various transfer learning algorithms in the literature, we focus on the attribute-based transfer learning. This algorithm realizes transfer learning by introducing attributes and transferring the results of training to another task with
This paper proposes and analyzes a methodology of forecasting movements of the analysts’ net income estimates and those of stock prices. We achieve this by applying natural language processing and neural networks in the context of analyst reports. In the pre-experiment, we applied our method to extract opinion sentences from the analyst report while classifying the remaining parts as non-opinion sentences. Then, we performed two additional experiments. First, we employed our proposed method for