Community Evals: Because we're done trusting black-box leaderboards over the community
Related stories
Introducing the Open FinLLM Leaderboard
Introducing the Red-Teaming Resistance Leaderboard
Introducing the Enterprise Scenarios Leaderboard: a Leaderboard for Real World Use Cases
Open LLM Leaderboard: DROP deep dive
Introducing the Open Ko-LLM Leaderboard: Leading the Korean LLM Evaluation Ecosystem
Deliberative Curation: A Protocol for Multi-Agent Knowledge Bases
arXiv:2606. 00007v1 Announce Type: new Abstract: As AI agents transition from isolated tools to collaborative participants in shared knowledge ecosystems, governing collective knowledge curation becomes a critical challenge.
AI Evaluation Should Work With Humans
arXiv:2608. 13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.
What's going on with the Open LLM Leaderboard?
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv:2608. 07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable.
A shared playbook for trustworthy third party evaluations
OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
arXiv:2607. 24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective.