Harvesting speech datasets for linguistic research on the web
Author
Rooth, Mats
Howell, Jonathan
Wagner, Michael
Abstract
This is a white paper for a project that harvested audio and transcribed data from podcasts and news broadcasts on the web. Tools were developed to analyze the different uses of prosody (rhythm, stress and intonation) within spoken communication using phonetic analysis and machine learning.
Sponsorship
NSF 1035151 RAPID: Harvesting Speech Datasets for Linguistic Research on the Web (Digging into Data Challenge) and SSHRC Digging into Data Challenge Grant 869-2009-0004.
Date Issued
2013-10-29
Keywords
Previously Published as
Final project white paper, Digging into Data Challenge
Type
article