FFASR Leaderboard Launched — a Benchmark for Evaluating ASR in Real Acoustic Conditions
Treble Technologies
Hugging Face
OpenAI
IBM
Cohere
Meta
SpeechBrain
Treble Technologies and Hugging Face have launched the FFASR Leaderboard, the first open community benchmark for evaluating automatic speech recognition (ASR) systems in far-field conditions. The benchmark shows that at low signal-to-noise ratio (SNR), far-field recognition error rates are several times higher than in near-field. The platform enables evaluation of both accuracy (word error rate, WER) and speed (real-time factor, RTFx) of models.
Treble Technologies and Hugging Face have introduced the FFASR Leaderboard, the first open, community-supported benchmark for evaluating ASR models in realistic far-field conditions with acoustically complex environments: reverberation, noise, and distant microphones. The benchmark uses Treble's hybrid simulator, combining a wave solver for low and mid frequencies with geometric acoustics for high frequencies, enabling accurate modeling of physical phenomena in 14 different rooms ranging from 20 to 470 m³. Evaluation covers nine conditions, four of which form the main ranking, and includes tracks for simulation validation (Lab Measured and Lab Simulated) as well as a beta version with a moving sound source. For each submission, Word Error Rate (WER) for nine conditions and RTFx (audio seconds per inference second) on an NVIDIA L4 GPU are recorded. Results from all submitted models show a consistent gap: WER in far-field with low SNR is several times higher than in near-field, and the Pareto front visualization clearly demonstrates the trade-off between accuracy and speed. The benchmark supports architectures such as Whisper, IBM Granite Speech, Cohere Transcribe, Wav2Vec2, HuBERT, SpeechBrain, and others via Hub, and custom evaluators can be created for complex stacks. The developers plan to add multi-speaker scenarios, microphone array support, and echo cancellation.
Source: Hugging Face blog —
original
