MediaParl: Bilingual mixed language accented speech database
- Resource Type
- Conference
- Authors
- Imseng, David; Bourlard, Herve; Caesar, Holger; Garner, Philip N.; Lecorve, Gwenole; Nanchen, Alexandre
- Source
- 2012 IEEE Spoken Language Technology Workshop (SLT) Spoken Language Technology Workshop (SLT), 2012 IEEE. :263-268 Dec, 2012
- Subject
- Computing and Processing
General Topics for Engineers
Speech
Databases
Dictionaries
Switches
Training
Standards
Speech recognition
Multilingual corpora
Non-native speech
Mixed language speech recognition
Language identification
- Language
MediaParl is a Swiss accented bilingual database containing recordings in both French and German as they are spoken in Switzerland. The data were recorded at the Valais Parliament. Valais is a bilingual Swiss canton with many local accents and dialects. Therefore, the database contains data with high variability and is suitable to study multilingual, accented and non-native speech recognition as well as language identification and language switch detection. We also define monolingual and mixed language automatic speech recognition and language identifictaion tasks and evaluate baseline systems. The database is publicly available for download.