Hugging Face Trending Papers

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

Read the original on Hugging Face Trending Papers →

On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is toward how much compute that factuality costs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
2d ago

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent evaluations increasingly benchmark LLMs, but rankings can be swayed by evaluation conditions such as scaffolds or tasks, making reliability claim‑dependent. A Bayesian variance‑decomposition framework applied to 22 benchmarks shows that reliability varies with the measurement goal: fixed model‑scaffold systems rank reliably, while underlying‑model rankings are less stable. Scaffold choice can alter conclusions, and adding more tasks only modestly improves reliability when scaffold coverage is limited; however, pooling diverse benchmarks can substantially raise cross‑task ranking reliability and reduce cost.

By Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
arXiv AI
Aug 28

AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling

AgentJudgeBench is a new benchmark that evaluates the reliability of large language model (LLM) judges on agentic tool‑calling tasks involving workflow directed acyclic graphs (DAGs). It contains 3,808 instances across six DAG topologies and three difficulty tiers, tested with five generators (3B–70B open‑weight models and GPT‑5.4) and six judges (20B to frontier scale) under both paired‑with‑and‑without‑ground‑truth conditions. The study finds that judge alignment degrades with task difficulty, ground‑truth exposure can sometimes hurt alignment, and structured evaluation rubrics provide modest improvements, revealing a structural ceiling that model capacity alone cannot surpass.

By Abhigya Verma, Amit Kumar Saha, Seganrasan Subramanian, Sai Harshitha Aluru