Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Related stories
Introducing the Open FinLLM Leaderboard
Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases
Adding Benchmaxxer Repellant to the Open ASR Leaderboard
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
The Open ASR Leaderboard Adds Its First Global South Language
The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard
Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models
The paper introduces a lightweight ASR head that can be added to full‑duplex speech‑to‑speech models, enabling real‑time user transcription without major architectural changes. The method adds only a few parameters and preserves full‑duplex conversational features such as turn‑taking and barge‑in. Experiments show a streaming WER of 10.21% within the duplex framework and 7.73% when trained as a standalone ASR model, matching state‑of‑the‑art performance.
Open LLM Leaderboard: DROP deep dive
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
arXiv:2606. 19704v1 Announce Type: new Abstract: Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes.
GENEB: Why Genomic Models Are Hard to Compare
arXiv:2606. 04525v1 Announce Type: cross Abstract: Progress in genomic foundation models is difficult to assess due to fragmented benchmarks, incompatible evaluation protocols, and task-specific reporting.