Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell Computing and Information Science
  3. Computing and Information Science
  4. Computing and Information Science Research
  5. Data from SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

Data from SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

File(s)
SoniSpeech_01_of_09.zip (7.97 GB)
SoniSpeech_02_of_09.zip (7.96 GB)
SoniSpeech_03_of_09.zip (7.96 GB)
SoniSpeech_04_of_09.zip (7.96 GB)
SoniSpeech_05_of_09.zip (7.96 GB)
  View More
Permanent Link(s)
https://doi.org/10.7298/xjjr-9m85
https://hdl.handle.net/1813/124578
Collections
Computing and Information Science Research
Author
Zhang, Ruidong
Liu, Jiacheng
Guimbretière, François
Zhang, Cheng
Abstract

These files contain data supporting all results reported in SoniSpeech. Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies. Approaches achieving larger vocabularies require obtrusive hardware such as facial electrodes. We present SoniSpeech, the first large-scale, open-vocabulary, trimodal dataset for wearable SSI using acoustic-sensing eyewear. It contains 34 hours across 18,000 utterances with three synchronized modalities: ultrasound echo profiles, voiced audio, and frontal video, in both voiced and silent modes. The corpus draws from the SODA dialogue dataset, providing contemporary conversational English with 5,356 unique words and full phoneme coverage. A CTC-based ResNet-34 baseline achieves 26.3% word error rate (WER) on open-vocabulary silent speech recognition, the first benchmark for this task.

Description
Please cite as: Ruidong Zhang, Jiacheng Liu, François Guimbretière, Cheng Zhang. (2026) Data from SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces. [Data set] Cornell University Library eCommons Repository. https://doi.org/10.7298/xjjr-9m85
Date Issued
2026
Keywords
silent speech interface
•
acoustic sensing
•
dataset
•
wearable computing
•
smart glasses
Rights
Attribution-NonCommercial 4.0 International
Rights URI
https://creativecommons.org/licenses/by-nc/4.0/
Type
dataset

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance