The University of Tokyo · Computer Science
Professor Masahiro Suzuki's research lab specializes in multimodal machine learning and deep generative modeling, with a focus on cross-modal generation and joint representation learning across diverse data types such as text, images, and biological signals. The lab develops advanced variational autoencoder frameworks—like the Joint Multimodal Variational Autoencoder (JMVAE)—to enable bidirectional generation and robust inference in heterogeneous data environments. Additionally, the lab explores applications in biomedical engineering, including viral vector design for gene therapy and the evaluation of physiological functions such as swallowing reserve through surface electromyography. These interdisciplinary efforts bridge artificial intelligence with healthcare and life sciences.
Figures are computed from collected data and may differ slightly.
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation th
Multimodal learning is a framework for building models that make predictions based on different types of modalities. Important challenges in multimodal learning are the inference of shared representations from arbitrary modalities and cross-modal generation via these representations; however, achieving this requires taking the heterogeneous nature of multimodal data into account. In recent years, deep generative models, i.e. generative models in which distributions are parameterized by deep neur
The adenovirus vector (AdV) can carry two transgenes in its genome, the therapeutic gene and a reporter gene, for example. The E3 insertion site has often been used for the expression of the second transgene. A transgene can be inserted at six different sites/orientations: E1, E3 and E4 sites, and right and left orientations. However, the best combination of the insertion sites and orientations as for the titers and the expression levels has not sufficiently been studied. We attempted to constru
The ability to fine-tune the movement of swallowing-related organs and change the swallowing pattern to fit the volume of a bolus, texture and the physical properties of the food to be swallowed is referred to as the swallowing reserve. In other words, it is the response capability of food swallowing to avoid choking and aspiration. Herein, we focus on the coordination of the suprahyoid and infrahyoid muscles activities, which are closely related to swallowing movement, as a first step to develo
We investigate deep generative models that can exchange multiple modalities bi-directionally, e.g., generating images from corresponding texts and vice versa. Recently, some studies handle multiple modalities on deep generative models, such as variational autoencoders (VAEs). However, these models typically assume that modalities are forced to have a conditioned relation, i.e., we can only generate modalities in one direction. To achieve our objective, we should extract a joint representation th
Machine learning is the basis of important advances in artificial intelligence. Unlike the general methods of machine learning, which use the same tasks for training and testing, the method of transfer learning uses different tasks to learn a new task. Among the various transfer learning algorithms in the literature, we focus on the attribute-based transfer learning. This algorithm realizes transfer learning by introducing attributes and transferring the results of training to another task with
This paper proposes and analyzes a methodology of forecasting movements of the analysts’ net income estimates and those of stock prices. We achieve this by applying natural language processing and neural networks in the context of analyst reports. In the pre-experiment, we applied our method to extract opinion sentences from the analyst report while classifying the remaining parts as non-opinion sentences. Then, we performed two additional experiments. First, we employed our proposed method for
Open papers in the app to read, cite, and organize with AI.