Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,257 stories · RSS feed

arXiv Machine Learning
Jul 7

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

arXiv:2605. 17758v2 Announce Type: replace Abstract: Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information.

By Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran, Tu Nguyen, Arshia Harish Puthran, Muhjaazee Love, Aadi Sharma, Mahdi Bagheri, Ian Harris, Amir M. Rahmani
arXiv AI
Jul 7

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

arXiv:2607. 03600v1 Announce Type: cross Abstract: Adversarial robustness in Unsupervised Domain Adaptation (UDA) remains a significant challenge due to noisy pseudo labels and inherent distributional shifts between the clean source and adversarially perturbed target domains.

By Sushant Dagaji Desale, Rahul Mishra, Ashutosh Kumar Sinha
arXiv Machine Learning
Jul 7

GeoFlow: Geo-Aware Modeling of Inter-Area Relationships in Origin-Destination Flow Prediction and Generation

arXiv:2607. 05257v1 Announce Type: new Abstract: Origin-destination (OD) flow modeling underpins urban planning and mobility analysis, but prevailing graph-based methods often neglect salient geographic attributes, limiting their ability to model long-range and multi-area dependencies.

By Zherui Huang, Guanjie Zheng, Hao Xue, Linghe Kong
arXiv Machine Learning
Jul 7

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

arXiv:2605. 03677v2 Announce Type: replace Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model.

By Wenjin Hou, Shangpin Peng, Weinong Wang, Zheng Ruan, Yue Zhang, Zhenglin Zhou, Mingqi Gao, Yifei Chen, Kaiqi Wang, Hongming Yang, Chengquan Zhang, Zhuotao Tian, Han Hu, Yi Yang, Fei Wu, Hehe Fan
arXiv AI
Jul 7

Graph Representation Learning of Longitudinal Medical Imaging Trajectories for Treatment Response Prediction

arXiv:2607. 04912v1 Announce Type: cross Abstract: In patients with breast cancer, pathological complete response (pCR) has been established as a clinically meaningful surrogate marker for long-term outcomes.

By Johannes Kiechle, Richard Osuala, Daniel M. Lang, Stefan M. Fischer, Ivana Jan\'i\v{c}kov\'a, Karim Lekadir, Julia A. Schnabel, Jan C. Peeken
arXiv AI
Jul 7

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

arXiv:2605. 29976v2 Announce Type: replace-cross Abstract: We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather forecasting and evaluated up to a 10-day lead time.

By Renu Singh, Robert Brunstein, Antonia Jost, Yana Hasson, Thomas Rackow, Claire Monteleoni, Christian Lessig, Guillaume Couairon