Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,030 stories · RSS feed

arXiv AI
Aug 10

Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

arXiv:2511. 07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure.

By Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher R\'e
arXiv AI
Aug 10

A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy

arXiv:2608. 07427v1 Announce Type: new Abstract: LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (NTSDA), where raw multivariate KPI windows from 4G/5G cell sites expand into thousands of floating-point tokens.

By Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury
arXiv AI
Aug 10

Social World Models

arXiv:2509. 00559v3 Announce Type: replace Abstract: Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited information.

By Xuhui Zhou, Jiarui Liu, Akhila Yerukola, Hyunwoo Kim, Maarten Sap
arXiv Machine Learning
Aug 10

Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions

arXiv:2409. 18804v3 Announce Type: replace-cross Abstract: Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond.

By Iskander Azangulov, George Deligiannidis, Judith Rousseau
arXiv Machine Learning
Aug 10

Fast and Accurate: An Adaptive VLA Inference Framework through Environment-aware Model Selection

arXiv:2608. 06434v1 Announce Type: cross Abstract: Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness.

By Yuewei Sun, Lang Qin, Zechuan Tian, Jingwen Li, Guiqin Wang, Shengzeng Huo, Wenxin Ren, Tao Fang, Xiaochen Zhang, Guanqing Deng, Xiang Wang, Xiaowen Dong, Qinghai Guo, Yuxin Ma