Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. Cornell University Graduate School
  3. Cornell Theses and Dissertations
  4. From Silence to Voice: Sensing, Recognizing and Restoring Silent Speech

From Silence to Voice: Sensing, Recognizing and Restoring Silent Speech

Access Restricted

Access to this document is restricted. Some items have been embargoed at the request of the author, but will be made publicly available after the "No Access Until" date.

During the embargo period, you may request access to the item by clicking the link to the restricted file(s) and completing the request form. If we have contact information for a Cornell author, we will contact the author and request permission to provide access. If we do not have contact information for a Cornell author, or the author denies or does not respond to our inquiry, we will not be able to provide access. For more information, review our policies for restricted content.

File(s)
Zhang_cornellgrad_0058F_15582.pdf (17.82 MB)
No Access Until
2027-06-22
Permanent Link(s)
https://doi.org/10.7298/yzwk-9588
https://hdl.handle.net/1813/126629
Collections
Cornell Theses and Dissertations
Author
Zhang, Ruidong
Abstract

Speech is arguably the most natural form of human communication, yet millions of people cannot speak aloud, whether by choice in privacy-sensitive situations or by necessity after losing their voice to disease. This dissertation develops wearable systems that capture a person's silent articulatory movements and convert them into practical communication: text for those who choose silence, and restored voice for those who have lost theirs. Six systems form a progressive arc across three dimensions: sensing, recognition, and restoration. The first dimension is sensing. Miniature speakers on a wearable device emit inaudible sound waves toward the user's face. When the user silently mouths words, the movements of their lips, jaw, and tongue reshape the returning echoes into distinctive patterns that encode the articulatory state. SpeeChin, a necklace-mounted infrared camera, first proved that wearable silent speech recognition is feasible with over 50 commands but exposed the structural limits of optical sensing: high power consumption, privacy risk, and sensitivity to lighting. EchoSpeech replaced the camera with near-ultrasonic acoustic sensing on smart glasses, achieving 4.5% word error rate on 31 commands at 73.3 mW, a 74x power reduction. HPSpeech generalized acoustic sensing to commodity headphones, demonstrating that the approach works on devices millions of people already own. The second dimension is recognition. SoniSpeech constructed the first large-scale, open-vocabulary dataset for minimally-obtrusive wearable silent speech: 34.1 hours, 18,000 utterances, 5,356 unique words. A baseline system achieved 26.3% word error rate on silent speech with no language model, on a scaling curve that shows no saturation. WhisperWeave extends acoustic-sensing recognition to laryngectomees by fusing articulatory echo profiles with pseudo-whisper audio on a single pair of eyeglasses, achieving 13.8% word error rate through bimodal complementarity. The third dimension is restoration. ReVoice shifted from text output to voice synthesis, introducing bone conduction probing of the vocal tract and an end-to-end generative adversarial network that synthesizes natural speech directly from silent articulation with 102.7 ms latency. Evaluated with 26 participants including 11 people who had lost their vocal cords to laryngeal cancer, ReVoice achieved 18.9% word error rate and 3.82 mean opinion score for naturalness, with no significant difference between clinical and healthy participants. These results substantially exceed the speech quality of the electrolarynx while providing a hands-free, socially acceptable form factor. Together, these six systems demonstrate that active acoustic sensing on everyday wearables can sense silent articulation with sufficient fidelity. Powered by advanced machine learning, we can recognize silent speech from closed vocabulary to open-vocabulary natural language, generalize to clinical populations through bimodal fusion, and restore expressive voice for people with no vocal cords.

Description
186 pages
Date Issued
2026-05
Committee Chair
Zhang, Cheng
Committee Member
Guimbretiere, Francois
Choudhury, Tanzeem
Degree Discipline
Information Science
Degree Name
Ph. D., Information Science
Degree Level
Doctor of Philosophy
Rights
Attribution-NonCommercial-ShareAlike 4.0 International
Rights URI
https://creativecommons.org/licenses/by-nc-sa/4.0/
Type
dissertation or thesis

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance