The University of Tokyo · Computer Science
Professor Naoki Kimura's research lab specializes in human-computer interaction, focusing on silent and wearable speech interfaces, tactile sensing for mobile text entry, and AI-driven sensory augmentation. The lab develops innovative systems like SilentSpeller and TieLent that enable voice-free communication using physiological signals such as tongue movement and subtle facial gestures, emphasizing usability in dynamic, real-world environments. Another key direction involves using deep learning to enhance sensory experiences—such as visual and auditory perception—through context-aware image generation and timbre modeling. The lab also explores low-cost, teacher-free solutions for musical instrument learning using unsupervised representation learning.
Figures are computed from collected data and may differ slightly.
The availability of digital devices operated by voice is expanding rapidly. However, the applications of voice interfaces are still restricted. For example, speaking in public places becomes an annoyance to the surrounding people, and secret information should not be uttered. Environmental noise may reduce the accuracy of speech recognition. To address these limitations, a system to detect a user's unvoiced utterance is proposed. From internal information observed by an ultrasonic imaging sensor
Speech is inappropriate in many situations, limiting when voice control can be used. Most unvoiced speech text entry systems can not be used while on-the-go due to movement artifacts. Using a dental retainer with capacitive touch sensors, SilentSpeller tracks tongue movement, enabling users to type by spelling words without voicing. SilentSpeller achieves an average 97% character accuracy in offline isolated word testing on a 1164-word dictionary. Walking has little effect on accuracy; average o
We propose a system, called ExtVision, to augment visual experiences by generating and projecting context-images onto the periphery of the television or computer screen. A peripheral projection of the context-image is one of the most effective techniques to enhance visual experiences. However, the projection is not commonly used at present, because of the difficulty in preparing the context-image. In this paper, we propose a deep neural network-based method to generate context-images for periphe
Voice control provides hands-free access to computing, but there are many situations where audible speech is not appropriate. Most unvoiced speech text entry systems can not be used while on-the-go due to movement artifacts. SilentSpeller enables mobile silent texting using a dental retainer with capacitive touch sensors to track tongue movement. Users type by spelling words without voicing. In offline isolated word testing on a 1164-word dictionary, SilentSpeller achieves an average 97% charact
With the increased use of smart speakers, silent speech interaction (SSI) is attracting attention. Unfortunately, traditional silent speech interaction methods require the addition of obtrusive sensors and devices around the user's face, making wearability and portability a challenge. Considering that most uses for smart speakers do not require many words, we suggest a more casual approach, TieLent, which can easily be worn between the neck and the chest. TieLent's RGB camera is set away from th
One of the most difficult things in practicing musical instruments is improving timbre. Unlike pitch and rhythm, timbre is a high-dimensional and sensuous concept, and learners cannot evaluate their timbre by themselves. To efficiently improve their timbre control, learners generally need a teacher to provide feedback about timbre. However, hiring teachers is often expensive and sometimes difficult. Our goal is to develop a low-cost learning system that substitutes the teacher. We found that a v
Immersion is an important factor in video experiences. Therefore, various methods and video viewing systems have been proposed. Head-mounted displays (HMDs) are home-friendly pervasive devices, which can provide an immersive video experience owing to their wide field-of-view (FoV) and separation of users from the outside environment. They are often used for viewing panoramic and stereoscopic recorded videos or virtually generated environments, but the demand for viewing standard plane videos wit
Data-driven machine learning approaches have become increasingly used in human-computer interaction (HCI) tasks. However, compared with traditional machine learning tasks, for which large datasets are available and maintained, each HCI project needs to collect new datasets because HCI systems usually propose new sensing or use cases. Such datasets tend to be lacking in amount and lead to low performance or place a burden on participants in user studies. In this paper, taking hand gesture recogni
reactivation is required in patients with past HBV infection receiving abatacept.
LT to the primary lesion in metastatic hormone-sensitive prostate cancer may provide prognostic benefits and especially in patients with low tumor burden.
Immersion is an important factor in video experiences. Therefore, various methods and video viewing systems have been proposed so far. Although head-mounted displays (HMDs) are home-friendly and more available among these devices, they can provide an immersive video experience owing to their wide field-of-view (FoV) and separation of users from the outside environment. They are often used for panoramic and stereoscopic VR videos, but the demand for viewing standard plane videos has increased in
Open papers in the app to read, cite, and organize with AI.