Cornell University
Library
Cornell UniversityLibrary

eCommons

Help
Log In(current)
  1. Home
  2. College of Arts and Sciences
  3. Linguistics
  4. Linguistics - Monographs, Papers and Research
  5. A web application for filtering and annotating web speech data

A web application for filtering and annotating web speech data

File(s)
Lutz-Cadwallader-Rooth.pdf (325.38 KB)
Article
Permanent Link(s)
https://hdl.handle.net/1813/33464
Collections
Linguistics - Monographs, Papers and Research
Author
Lutz, David
Cadwallader, Parry
Rooth, Mats
Abstract

A vast and growing amount of recorded speech is freely available on the web, including podcasts, radio broadcasts, and posts on media-sharing sites. However, finding specific words or phrases in online speech data remains a challenge for researchers, not least because transcripts of this data are often automatically-generated and imperfect. We have developed a web application, “ezra”, that addresses this challenge by allowing non-expert and potentially remote annotators to filter and annotate speech data collected from the web and produce large, high-quality data sets suitable for speech research. We have used this application to filter and annotate thousands of speech tokens. Ezra is freely available on GitHub1, and development continues.

Sponsorship
NSF 1035151 RAPID: Harvesting
Speech Datasets for Linguistic Research on the
Web (Digging into Data Challenge)
Date Issued
2013-07-22
Publisher
Special Interest Group of the ​Association for Computational Linguistics on Web as Corpus (ACL SIGWAC)
Keywords
corpus
•
speech
•
web interface
•
annotation
•
filtering
•
prosody
Previously Published as
Stefan Evert, Egon Stemle and Paul Rayson (editors). Proceedings of the 8th Web as Corpus Workshop. July, 2013.
Type
article

Site Statistics | Help

About eCommons | Policies | Terms of use | Contact Us

copyright © 2002-2026 Cornell University Library | Privacy | Web Accessibility Assistance