Automatic Speech-to-Text Transcription in Arabic
Abstract
The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Being a highly inflected language, the Arabic language has a very large lexical variety and typically with several possible (generally semantically linked) vowelizations for each written form. This article summarizes research carried out over the last few years on speech-to-text transcription of broadcast data in Arabic. The initial research was oriented toward processing of broadcast news data in Modern Standard Arabic, and has since been extended to address a larger variety of broadcast data, which as a consequence results in the need to also be able to handle dialectal speech. While standard techniques in speech recognition have been shown to apply well to the Arabic language, taking into account language specificities help to significantly improve system performance.
Document Details
- Document Type
- Pub Defense Publication
- Publication Date
- Dec 01, 2009
- Source ID
- 10.1145/1644879.1644885
Entities
People
- Abdelkhalek Messaoudi
- Jean-luc Gauvain
- Lori Lamel
Organizations
- Defense Advanced Research Projects Agency