Phonetic inventory for an Arabic speech corpus

Halabi, Nawar and Wald, Mike (2016) Phonetic inventory for an Arabic speech corpus. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Slovenia, Slovenia. 23 - 28 May 2016. pp. 734-738 .

Record type: Conference or Workshop Item (Poster)

Abstract

Corpus design for speech synthesis is a well-researched topic in languages such as English compared to Modern Standard Arabic, and there is a tendency to focus on methods to automatically generate the orthographic transcript to be recorded (usually greedy methods). In this work, a study of Modern Standard Arabic (MSA) phonetics and phonology is conducted in order to create criteria for a greedy meth-od to create a speech corpus transcript for recording. The size of the dataset is reduced a number of times using these optimisation methods with different parameters to yield a much smaller dataset with identical phonetic coverage than before the reduction, and this output transcript is chosen for recording. This is part of a larger work to create a completely annotated and segmented speech corpus for MSA.

Text

Arabic Phonetic Vocab 2016.pdf - Accepted Manuscript

Available under License University of Southampton Accepted Manuscript Licence.

Download (732kB)

More information

Published date: 25 May 2016

Venue - Dates: Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016), Slovenia, Slovenia, 2016-05-23 - 2016-05-28

Keywords: phonology, corpus design, corpus evaluation

Organisations: Web & Internet Science