Automatic Speech-to-Text Transcription in Arabic

Abstract

The Arabic language presents a number of challenges for speech recognition, arising in part from the significant differences in the spoken and written forms, in particular the conventional form of texts being non-vowelized. Being a highly inflected language, the Arabic language has a very large lexical variety and typically with several possible (generally semantically linked) vowelizations for each written form. This article summarizes research carried out over the last few years on speech-to-text transcription of broadcast data in Arabic. The initial research was oriented toward processing of broadcast news data in Modern Standard Arabic, and has since been extended to address a larger variety of broadcast data, which as a consequence results in the need to also be able to handle dialectal speech. While standard techniques in speech recognition have been shown to apply well to the Arabic language, taking into account language specificities help to significantly improve system performance.

Document Details

Document Type: Pub Defense Publication
Publication Date: Dec 01, 2009
Source ID: 10.1145/1644879.1644885

Entities

People

Abdelkhalek Messaoudi
Jean-luc Gauvain
Lori Lamel

Organizations

Defense Advanced Research Projects Agency

Automatic Speech-to-Text Transcription in Arabic

Abstract

Document Details

Entities

People

Organizations

Tags

Fields of Study

Readers

Technology Areas