Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

16,065 stories · RSS feed

arXiv Machine Learning
Jul 22

Uncertainty Quantification for AI-Driven Crash Simulation Surrogates: A Comparative Study of Monte Carlo Dropout and Deep Ensemble on Open-Source Bumper Beam Benchmark

arXiv:2607. 18294v1 Announce Type: new Abstract: Machine learning surrogate models are increasingly being explored in engineering product development to augment simulation-driven design, offering near-instantaneous predictions that complement computationally expensive high-fidelity analyses.

By Sudeep Chavare
arXiv Machine Learning
Jul 22

Is EEG-to-Text Feasible in Real-World Scenarios? An In-Depth Analysis Using a Neuropsychology-Inspired Benchmark

arXiv:2607. 18749v1 Announce Type: new Abstract: Translating brain signals into text could restore communication for people with severe paralysis, yet practically usable systems to date rely on invasive electrocorticography (ECoG).

By Zihan Zhang (Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology), Yu Bao (Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology, Shanghai Innovation Institute), Xiao Ding (Research Center for Social Computing and Interactive Robotics, Harbin Institute of Technology), Tianyi Jiang (State Key Laboratory for Novel Software Technology, Nanjing University), Kai Xiong (Zhongguancun Laboratory)
arXiv AI
Jul 22

BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data

arXiv:2607. 18241v1 Announce Type: new Abstract: Large language models (LLMs) excel at analyzing individual documents but break down on exhaustive, cross-entity analytical questions over enterprise-scale datasets due to context overflow, loss of per-entity attribution, and linear latency from sequential tool calls.

By Anupreet Walia
arXiv Machine Learning
Jul 22

Local Label-Informed Feature Transfer for Generating Ground-Truth Medical Images: A Comparison of GAN- and Diffusion-Based Approaches

arXiv:2607. 18882v1 Announce Type: cross Abstract: Validating Explainable Artificial Intelligence (XAI) methods in medical imaging requires ground-truth data with known locations of informative features.

By Rick Wilming, Irem Ozseker, Luca Matteo Cornils, Ahc\`ene Boubekki, Benedict Clark, Danny Panknin, Stefan Haufe
arXiv AI
Jul 22

MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers

arXiv:2607. 18252v1 Announce Type: new Abstract: Machine learning methods have shown that data-driven policies can accelerate mixed-integer linear programming (MILP) solvers, but many such approaches remain difficult to inspect, adapt, and deploy because the learned policy is represented as an external predictor or other opaque model.

By Jinbiao Nie, Kewei Feng, Xiaoyuan Zhang, Shan Yin, Zizhuo Wang, Bin Dong
arXiv AI
Jul 22

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

arXiv:2607. 05462v2 Announce Type: replace-cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse.

By Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman