ML classification in a benchmark of prosodic minimal pairs
This paper presents the following: (i) A computational methodology for collecting large numbers of utterances of a fixed word string from online sources, using the index of youtube transcriptions at filmot.com. (ii) A prototype benchmark constructed with the methodology consisting of prosodic minimal pairs, which are short word sequences which depending on context and/or lexical identity are pronounced with different prosodies. The benchmark is grouped into pairs, for instance utterances of "much as I did'" with or without focus prosody on the first person subject. (iii) A classification model obtained by tuning wav2vec2-base on the training portion of the benchmark. Presented with an audio that is stipulated to be an utterance of a given word string, the model selects one of two alternative prosodies for the word string. For instance, it can determine whether a given utterance of "much as I did" has focus prosody on the subject, or default prosody. Accuracy of this determination on separate test data is better than 90% for each minimal pair.