Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

17,521 stories · RSS feed

arXiv AI
Jul 7

Graph Representation Learning of Longitudinal Medical Imaging Trajectories for Treatment Response Prediction

arXiv:2607. 04912v1 Announce Type: cross Abstract: In patients with breast cancer, pathological complete response (pCR) has been established as a clinically meaningful surrogate marker for long-term outcomes.

By Johannes Kiechle, Richard Osuala, Daniel M. Lang, Stefan M. Fischer, Ivana Jan\'i\v{c}kov\'a, Karim Lekadir, Julia A. Schnabel, Jan C. Peeken
arXiv AI
Jul 7

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

arXiv:2607. 04557v1 Announce Type: cross Abstract: Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles.

By Dongmin Bang, Sugyun An, Inyoung Sung, Ilho Yun, Sun Kim, Sangseon Lee
arXiv AI
Jul 7

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.

By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li
arXiv AI
Jul 7

Spectral Rewiring for Exploration, Purification, and Model Merging

arXiv:2607. 03065v1 Announce Type: cross Abstract: Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two deployment-relevant bottlenecks: suppressed reasoning performance, often reflected by premature saturation of test-time scaling, and interference when consolidating multiple capabilities through multi-domain training or model merging.

By Zhilong Zhang, Hongli Yu, Huan-ang Gao, Hanlin Wu, Yuxuan Song, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou
arXiv AI
Jul 7

Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models

arXiv:2607. 04546v1 Announce Type: cross Abstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation.

By Riccardo O. Feingold, Davide Liconti, Chenyu Yang, Robert K. Katzschmann
arXiv AI
Jul 7

Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration

arXiv:2511. 02200v2 Announce Type: replace Abstract: The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling diverse agents to integrate unique expertise, collaborate flexibly, and address challenges unattainable for individual models.

By Jingbo Wang, Sendong Zhao, Haochun Wang, Yuzheng Fan, Ting Liu