EVALUATING AND DESIGNING SPEECH RECOGNITION
Automated Speech Recognition (ASR) systems currently treat transcription as a purely technical problem of objective accuracy, assuming a single “correct” representation of speech. This dissertation argues that this rigid, one-size-fits-all approach marginalizes diverse speech patterns and denies user autonomy. Through a combination of philosophical analysis, empirical evaluation, and human-centered systems design, this work fundamentally reorients the evaluation and design of speech technologies to promote algorithmic equity.First, it reframes ASR bias as a form of epistemic injustice, demonstrating that systemic misrecognition constitutes structural disrespect and imposes unique temporal harms on marginalized communities. Second, it challenges the evaluation paradigm of “reference monism” – the enforcement of a single transcription convention as the ultimate ground truth. It introduces Epistemic Injustice Distance (EID) and the WER-Range metric to evaluate ASR performance equitably across multiple legitimate conventions. Finally, it introduces SpeechSpectrum, a framework that reconceptualizes speech-to-text systems as tools for cross-modal translation rather than mechanical reproduction, granting users explicit control over transcript fidelity to match their contextual needs. Ultimately, this dissertation provides a comprehensive roadmap for building ASR systems that respect linguistic pluralism, distribute representational power equitably, and prioritize user agency.