Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,257 stories · RSS feed

arXiv AI
Jul 7

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

arXiv:2607. 03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals.

By Nguyen Kim Hai Bui, Md. Easin Arafat, Tam\'as G\'abor Orosz, Mufti Mahmud
arXiv AI
Jul 7

Graph Representation Learning of Longitudinal Medical Imaging Trajectories for Treatment Response Prediction

arXiv:2607. 04912v1 Announce Type: cross Abstract: In patients with breast cancer, pathological complete response (pCR) has been established as a clinically meaningful surrogate marker for long-term outcomes.

By Johannes Kiechle, Richard Osuala, Daniel M. Lang, Stefan M. Fischer, Ivana Jan\'i\v{c}kov\'a, Karim Lekadir, Julia A. Schnabel, Jan C. Peeken
arXiv AI
Jul 7

CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs

arXiv:2607. 04854v1 Announce Type: new Abstract: Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that violate task constraints, undermining their reliability in real-world applications.

By Qiuyi Qi, Jinjian Zhang, Mutian Bao, Tian Liang, Guocong Li, Dongnan Liu, Wei Zhou, Jie Liu, Ming Kong, Linjian Mo, Feng Zhang, Qiang Zhu
arXiv AI
Jul 7

MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model

arXiv:2607. 03013v1 Announce Type: cross Abstract: Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to sensor limitations and imaging pipelines, which degrades visual quality and affects downstream vision tasks.

By Wanshu Fan, Xiangyu Li, Cong Wang, Kin-man Lam, Xin Yang, Haiyan Zhang, Dongsheng Zhou