Falcon-ASR Beats Arabic Speech Benchmark, Targets Emirati Dialect
TII's 1.6B-parameter Falcon-ASR posts 20.92% Arabic WER and 22.73% Emirati WER, while supporting English, French, Spanish and Portuguese with unified weights.
Summary
On October 7, 2026, Abu Dhabi's Technology Innovation Institute introduced Falcon-ASR, a 1.6 billion-parameter speech recognition model prioritizing Arabic and the Emirati dialect. Following the ELM Research Center's Open Universal Arabic ASR Leaderboard protocol, it averaged 20.92% word error rate across six equally weighted test sets, 2.25 percentage points better than the 23.17% leading published result in the September 30 snapshot. On TII's held-out, human-validated Emirati and Gulf recordings, Falcon-ASR achieved 22.73% WER and 10.19% character error rate, the lowest among compared systems, with WER 4.07 points below Qwen3-Omni.
Falcon-ASR was trained on Emirati, Modern Standard Arabic, other Gulf and Arabic dialects, and English, with noise, overlapping speech, music, reverberation, telephony, speed and pitch variations reflecting calls, meetings and everyday recordings. Unified weights transcribe Arabic, English, French, Spanish and Portuguese without a language flag; English averaged 5.74% WER across seven Hugging Face Open ASR Leaderboard tests. The model adds word-level timestamps and builds on Falcon3-Audio. A Hugging Face demo is available, while API access and native applications are planned. TII credits the Falcon-Emirati team and Mikhail Lubinets for foundation-model and computing support.
Positives
- 20.92% average Arabic WER beat the leading published result by 2.25 percentage points across six equally weighted benchmarks.
- 22.73% WER and 10.19% CER were the lowest results among systems in TII's internal Emirati evaluation.
- 5.74% mean English WER was achieved across seven public Hugging Face benchmark sets.
- Five languages share one set of model weights and require no language flag.
- Word-level timestamps connect every transcribed word with its position in the recording.
Risks & concerns
- 20.92% Arabic WER shows substantial transcription errors remain despite the benchmark lead.
- 22.73% Emirati WER was measured internally on held-out recordings rather than through a fully public leaderboard evaluation.
- Dialectal Arabic has fewer transcribed resources than Modern Standard Arabic, complicating training and evaluation.
- Regional variation, conversational speech and telephone audio can still challenge systems that perform well on formal broadcasts.
- API access and native applications remain planned, leaving the Hugging Face demo as the current access route.