[Paper Review] Towards Ultrasound Tongue Image prediction from EEG during speech production
This study proposes a deep learning approach to predict ultrasound tongue images from non-invasive EEG signals during speech production, using a fully connected deep neural network. Results show a weak but noticeable EEG-UTI relationship, with the model distinguishing between articulated speech and neutral tongue positions, demonstrating feasibility for future speech-based brain-computer interfaces using articulatory feedback.
Previous initial research has already been carried out to propose speech-based BCI using brain signals (e.g. non-invasive EEG and invasive sEEG / ECoG), but there is a lack of combined methods that investigate non-invasive brain, articulation, and speech signals together and analyze the cognitive processes in the brain, the kinematics of the articulatory movement and the resulting speech signal. In this paper, we describe our multimodal (electroencephalography, ultrasound tongue imaging, and speech) analysis and synthesis experiments, as a feasibility study. We extend the analysis of brain signals recorded during speech production with ultrasound-based articulation data. From the brain signal measured with EEG, we predict ultrasound images of the tongue with a fully connected deep neural network. The results show that there is a weak but noticeable relationship between EEG and ultrasound tongue images, i.e. the network can differentiate articulated speech and neutral tongue position.
Motivation & Objective
- To investigate the feasibility of predicting ultrasound tongue images (UTI) from non-invasive EEG signals during speech production.
- To integrate multimodal data—EEG, speech, and real-time ultrasound tongue imaging—into a unified analysis framework.
- To explore whether articulatory kinematics, measured directly via ultrasound, can improve EEG-based speech decoding compared to indirect estimation methods.
- To establish a foundation for future speech neuroprostheses using non-invasive EEG and real-time articulatory feedback.
Proposed method
- Synchronized recordings of EEG, speech, and ultrasound tongue imaging were collected using hardware synchronization to ensure temporal alignment.
- A fully connected deep neural network (FC-DNN) was trained to map 80-dimensional EEG features to 2D ultrasound tongue images.
- EEG data were preprocessed and reduced to 80-dimensional features for input into the FC-DNN.
- The network was trained end-to-end using mean squared error (MSE) loss on synchronized EEG-UTI pairs.
- Model performance was evaluated using MSE on training, validation, and test sets, with qualitative visual inspection of predicted UTI sequences.
- A public Keras implementation of the DNN model is provided for reproducibility and further development.

Experimental results
Research questions
- RQ1Is there a detectable relationship between EEG signals and ultrasound tongue images during speech production?
- RQ2Can a deep neural network learn to predict UTI sequences from non-invasive EEG inputs?
- RQ3How well can the model differentiate between neutral tongue positions and active articulation during speech?
- RQ4Does the inclusion of real-time ultrasound data improve the interpretability or performance of EEG-based speech decoding compared to indirect articulatory estimation?
- RQ5What is the potential of using direct articulatory feedback in future non-invasive speech BCIs?
Key findings
- The FC-DNN achieved a test set mean squared error (MSE) of 0.0055, indicating measurable but limited prediction accuracy.
- Visual inspection revealed that the model could distinguish between articulated speech and neutral tongue positions, similar to voice activity detection.
- Despite low MSE values, the predicted UTI sequences showed limited visual similarity to the original images, suggesting weak but non-random EEG-UTI correspondence.
- The results demonstrate a weak but noticeable relationship between EEG and ultrasound tongue images, supporting the feasibility of EEG-to-UTI prediction.
- The study confirms that direct ultrasound-based articulatory data can be used to enhance EEG-based speech decoding, even with non-invasive recordings.
- The authors released a public Keras implementation of the model to support future research in EEG-to-UTI synthesis.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.