TII Unveils Falcon-ASR: State-of-the-Art Speech Recognition for Emirati Arabic
Falcon-ASR is a newly introduced, advanced speech recognition model developed by the Technology Innovation Institute (TII) in Abu Dhabi. This model is specifically engineered for the Arabic language, with a pronounced and critical focus on accurately transcribing the Emirati dialect. Beyond this primary focus, Falcon-ASR demonstrates robust multilingual capabilities, supporting English, French, Spanish, and Portuguese. The model is built upon the architecture and training methodologies established by the Falcon3-Audio work, indicating a sophisticated foundation in audio-language modeling.
One of the most significant aspects of this announcement is the model's performance metrics. In its evaluation across six standardized Arabic test sets, Falcon-ASR achieved an average Word Error Rate (WER) of 20.92%. This figure represents a notable improvement, as it significantly outperforms the best published result recorded in the leaderboard snapshot used for comparison, which stood at 23.17%. Furthermore, the model’s performance on the Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, was evaluated against the standard protocol, confirming that its average WER is 2.25 percentage points better than the best published result in that specific snapshot. The WER and Character Error Rate (CER) are key metrics in speech recognition, where lower values signify superior accuracy, indicating fewer mistakes in transcribing spoken words or characters.
However, the true depth of Falcon-ASR's capability is revealed in its internal Emirati evaluation. Recognizing that Arabic speech is highly variable—differing greatly by region, speaker, and recording environment—TII conducted a rigorous internal assessment. In this specialized Emirati evaluation, Falcon-ASR recorded a WER of 22.73% and a CER of 10.19%. Crucially, the model achieved the lowest WER and CER among all systems compared in this internal benchmark. This performance is particularly impressive because the evaluation conditions were intentionally challenging, including the simulation of background noise, overlapping speech, music, room reverberation, and common telephony effects. This comprehensive testing ensures that the model is reliable for real-world, everyday recordings, such as those encountered in meetings or phone calls, rather than just clean, formal broadcasts.
The development team addressed the inherent difficulties in Arabic speech processing. They highlighted that dialectal Arabic possesses fewer transcribed resources compared to Modern Standard Arabic (MSA), making both training and objective evaluation substantially more complex. To overcome this, Falcon-ASR was trained on a diverse mix of data, encompassing Emirati, MSA, various other Gulf and Arabic dialects, alongside English. The core objective of this extensive training regimen is to accurately transcribe the natural, everyday speech patterns used by people, including the complex linguistic phenomena of dialectal forms and code-switching (the transition between languages during speech).
Beyond the primary Arabic focus, the model’s versatility is a major selling point. It supports English, French, Spanish, and Portuguese, utilizing the same set of weights across all five languages, eliminating the need for a language-specific flag. For English, the model achieved a mean WER of 5.74% across seven public test sets used by the Hugging Face Open ASR Leaderboard. Moreover, the system provides advanced functionality, including word-level timestamps, which links every transcribed word back to its precise position within the original audio file, a feature invaluable for detailed transcription and analysis.
The ability to perform highly accurate transcription across multiple, distinct languages and challenging, real-world acoustic conditions marks a significant leap in the field of AI-powered speech processing. The internal evaluation further demonstrated the model's superior performance, showing that its WER was 4.07 percentage points lower than the next best competitor, Qwen3-Omni. This level of detail in testing—including the use of held-out recordings and human-validated transcripts beyond the public UAE subset—validates the model's robustness and high level of transcription accuracy at both the word and character level. The model's development and deployment through a Hugging Face Demo Space, with planned API access and native applications, signals its immediate availability for developers and enterprises looking to integrate state-of-the-art speech recognition into their products. The overall effort represents a major contribution to making advanced, dialect-aware speech technology accessible to a wider range of users and applications.
Why it matters
- —It sets a new benchmark for ASR accuracy, especially for complex, under-resourced dialects like Emirati Arabic.
- —The rigorous internal testing methodology (noise, overlap, reverberation) proves real-world applicability far beyond standard benchmarks.
- —Its multilingual capability (Arabic, English, French, Spanish, Portuguese) using unified weights makes it highly versatile for global enterprise use.
Key facts
- Achieved 20.92% average WER on six Arabic test sets, beating the published best result by 2.25 percentage points.
- In internal Emirati evaluation, it recorded the lowest WER (22.73%) and CER (10.19%) among compared systems.
- Trained on a diverse mix of dialects (Emirati, MSA, Gulf, etc.) to handle everyday speech and code-switching.
- Supports five languages (Arabic, English, French, Spanish, Portuguese) using the same model weights.
- Provides word-level timestamps, linking transcribed words to their exact position in the audio.
The full text is in the original source. Here we provide a brief summary and key facts.