Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

16,906 stories · RSS feed

arXiv Machine Learning
Jul 14

Inter-Stop Energy Prediction and Causal Driver Quantification for Dual-Source Trolleybuses via a Time-Aware Tabular Deep Learning Architecture

arXiv:2607. 11349v1 Announce Type: cross Abstract: Dual-source trolleybuses alternate between overhead catenary supply and on-board battery operation, creating energy-use patterns driven by route attributes, high-frequency trajectories, and hourly weather.

By Wentao Zeng (School of Management, Foshan University, Foshan, China a School of Management, Foshan University, Foshan, China, School of Mechanical and Electrical Engineering and Automation, Foshan University, Foshan, China), Zijian Huang (School of Artificial Intelligence, South China Normal University, Guangzhou, China), Yiming Bie (School of Transportation, Jilin University, Changchun, China), Jiabin Wu (School of Management, Foshan University, Foshan, China a School of Management, Foshan University, Foshan, China), Jun Gong (Department of Civil Engineering, The University of Hong Kong, Hong Kong, China)
arXiv Machine Learning
Jul 14

Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion

arXiv:2506. 22036v2 Announce Type: replace Abstract: With the increasing multimodal knowledge privatization requirements, multimodal knowledge graphs in different institutes are usually decentralized, lacking of effective collaboration system with both stronger reasoning ability and transmission safety guarantees.

By Ying Zhang, Yu Zhao, Xuhui Sui, Baohang Zhou, Xiangrui Cai, Li Shen, Xiaojie Yuan, Dacheng Tao
arXiv AI
Jul 14

Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation

arXiv:2607. 11412v1 Announce Type: cross Abstract: Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applications in urban planning, forest monitoring, and climate policy, where both accuracy and reliability are critical.

By Ritu Yadav, Andrea Nascetti, Yifang Ban
arXiv AI
Jul 14

Longitudinal Multi-View Breast Cancer Risk Prediction

arXiv:2607. 11343v1 Announce Type: cross Abstract: Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detection.

By Solveig Thrun, Zijun Sun, Suaiba A. Salahuddin, Kristoffer Wickstr{\o}m, Elisabeth Wetzer, Stine Hansen, Robert Jenssen, Michael Kampffmeyer
arXiv AI
Jul 14

Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

arXiv:2603. 22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-object compositional scenarios.

By Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun
arXiv AI
Jul 14

The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions

arXiv:2607. 10911v1 Announce Type: cross Abstract: In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications.

By Filip Klubicka, Vasudevan Nedumpozhimana, Sneha Rautmare, Bora Caglayan, Mingxue Wang, John D. Kelleher
arXiv AI
Jul 14

The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

arXiv:2607. 11175v1 Announce Type: new Abstract: The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping medical agents, moving them from task specific predictors toward autonomous systems that perceive, reason, plan, remember, and act in clinical environments.

By Chunzheng Zhu, Lei Tian, Bohan Tan, Ziqi Zhou, Yuxuan Sun, Yijun Wang, Chengchao Lv, Yilin Wen, Yijun He, Jinghao Lin, Yihang Chen, Cheewei Tan, Qianshan Wei, Lei Zhao, Bin Pu, Kenli Li, Yuan Xue, Jianxin Lin