Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

16,906 stories · RSS feed

arXiv AI
Jul 15

TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale

arXiv:2607. 13028v1 Announce Type: cross Abstract: Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse enough to cover the safety-critical long tail that logged data rarely contains.

By Zhouchonghao Wu, Akshay Rangesh, Weixin Li, Wei-Jer Chang, Zachary Lee, Tim Wang, Wei Zhan
arXiv Machine Learning
Jul 15

Exploring Zero-Shot Foundation Models for Multivariate Time Series Anomaly Detection

arXiv:2607. 12454v1 Announce Type: new Abstract: Multivariate Time Series Anomaly Detection (MTSAD) is essential for reliability and safety in domains such as industrial process monitoring and financial risk management, yet conventional approaches rely on application-specific models that are costly to train and hard to scale.

By Martin Uray, Saverio Messineo, Roland Kwitt, Stefan Huber
arXiv AI
Jul 15

An Empirical Study for Android-to-OpenHarmony GUI Test Migration

arXiv:2607. 11245v2 Announce Type: replace-cross Abstract: To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existing GUI test cases has become a critical problem.

By Yakun Zhang, Xinjia Chen, Yiyun Chen, Yuxia Zhang, Mingyi Zhou, Xiang Gao, Shaokun Zhang, Li Li, Yunming Ye
arXiv AI
Jul 15

Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions

arXiv:2607. 12406v1 Announce Type: new Abstract: The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model.

By Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Sirui Zhang, Hanyu Yang, Changxuan Fan, Zhongwei Xie, Hongyu Luo, Wun Yu Chan, Wei Fan, Haoran Li, Yangqiu Song
arXiv Machine Learning
Jul 15

MUSA-PINN: Multi-scale Weak-form Physics-Informed Neural Networks for Fluid Flow in Complex Geometries

arXiv:2603. 08465v3 Announce Type: replace Abstract: While Physics-Informed Neural Networks (PINNs) offer a mesh-free approach to solving fluid-flow PDEs, standard point-wise residual minimization suffers from convergence pathologies in topologically complex domains like Triply Periodic Minimal Surfaces (TPMS).

By Weizheng Zhang, Xunjie Xie, Hao Pan, Xiaowei Duan, Bingteng Sun, Qiang Du, Lin Lu
arXiv AI
Jul 15

A Neurosymbolic Approach to Natural Language Formalization and Verification

arXiv:2511. 09008v2 Announce Type: replace-cross Abstract: Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits their adoption in regulated industries like finance and health-care that operate under strict policies.

By Chenyang An, Sam Bayless, Stefano Buliani, Darion Cassel, Byron Cook, Duncan Clough, R\'emi Delmas, Nafi Diallo, Ferhat Erata, Nick Feng, Dimitra Giannakopoulou, Aman Goel, Aditya Gokhale, Joe Hendrix, Victor Heorhiadi, Marc Hudak, Dejan Jovanovi\'c, Andrew M. Kent, Benjamin Kiesl-Reiter, Jeffrey J. Kuna, Nadia Labai, Joseph Lilien, Divya Raghunathan, Zvonimir Rakamari\'c, Niloofar Razavi, Michael Tautschnig, Ali Torkamani, Nathaniel Weir, Michael W. Whalen, Jianan Yao
arXiv AI
Jul 15

Evidence-Grounded AI for Musculoskeletal Care

arXiv:2607. 12527v1 Announce Type: new Abstract: Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation.

By Wenjie Li, Yujie Zhang, Fanrui Zhang, Haoran Sun, Renhao Yang, Junjun He, Weiran Huang, Yuanfeng Ji, Chenrun Wang, Kailing Wang, Hongcheng Gao, Kaipeng Zhang, Hanyu Wang, Angela Lin Wang, Xingqi He, Yilin Huang, Shiyi Yao, Lilong Wang, Yankai Jiang, Yirong Chen, Chenglong Ma, Jiyao Liu, Ming Hu, Gen Li, Yidong Xu, Chengyu Zhuang, Jiawei Liu, Yin Zhang, Lequan Yu, Lu Chen, Yinpeng Dong, Lei Liu, Carlos Gutierrez Sanroman, Yu Qiao, Weijie Ma, Xiaosong Wang, Lei Wang