Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

12,573 stories · RSS feed

arXiv Machine Learning
5d ago

SPEAR: Structure Property Explainability with Attention Regularization

arXiv:2608. 13826v1 Announce Type: cross Abstract: Machine learning is increasingly used to learn structure property relationships from spectroscopic and diffraction data, yet its adoption in materials discovery is often limited by poor interpretability of model predictions.

By Aditya Raghavan, Utkarsh Pratiush, Dalton A. Pearl, Jade Holliman Jr, Katharine Page, Philip D Rack, Sergei V Kalinin
arXiv Machine Learning
5d ago

CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

arXiv:2608. 14332v1 Announce Type: cross Abstract: Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them.

By Anisa Saleem, Duksu Kim
arXiv AI
5d ago

HAM-RAG: Hierarchy-Aware Multimodal RAG for Structure-Faithful Interleaved Generation

arXiv:2608. 14032v1 Announce Type: cross Abstract: Existing multimodal RAG methods often flatten structured documents into isolated text and image units, weakening the source organization and local text-image logic needed for faithful evidence selection and placement.

By Yin Li, Ziyang Hu, Zhiyu Guo, Xiangyu Liu, Wenbin Li, Boo-Ho Yang, Rav Lawana, Ziyue Li, Wei Zeng, Fugee Tsung
arXiv AI
5d ago

ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction

arXiv:2608. 13622v1 Announce Type: new Abstract: Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting.

By Yongqi Tong, Tan Li Hui Faith, Choy Zhen Wen Marcus, Zhou Jin, Kewei Fu, Jiang-Ming Yang, Jianshe Li, Xin Zhang
arXiv Machine Learning
5d ago

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

arXiv:2608. 13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions.

By Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar, Mounya Elhilali, Zeyu Jin
arXiv AI
5d ago

ArGEnT: Arbitrary Geometry-encoded Transformer for Operator Learning

arXiv:2602. 11626v3 Announce Type: replace-cross Abstract: Learning solution operators on arbitrary geometries remains a central challenge in scientific machine learning, especially for many-query simulation, physics-informed learning, and evolving geometries requiring accurate, geometry-aware predictions at arbitrary spatial locations.

By Wenqian Chen, Zhi-Feng Wei, Yucheng Fu, Michael Penwarden, Pratanu Roy, Panos Stinis
arXiv AI
5d ago

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

arXiv:2608. 14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret.

By Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova, Poorani Gnanasambandan, Dilshad Hussain, Muhammad Uzair Khan, Nkembeng Kevin Nkengfoa, Paul Praveen J., Fabio Priante, Sjoerd Franciscus van der Werf, Thomas Frederik Jan van Roeden