Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,048 stories · RSS feed

arXiv AI
Aug 5

SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

arXiv:2608. 02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort.

By Abdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter
arXiv AI
Aug 5

WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks

arXiv:2608. 03499v1 Announce Type: new Abstract: Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, maintains state, and communicates with other agents through social and task relations.

By Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang, Xiyang Hu, Yue Zhao, Shuli Jiang
arXiv AI
Aug 5

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

arXiv:2607. 27670v2 Announce Type: replace-cross Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions.

By Shawn Li, Wei Yang, Jike Zhong, Jiate Li, Jiawei Yang, You Qin, Ryan Rossi, Franck Dernoncourt, Roger Zimmermann, Yue Wang, Zhengzhong Tu, Vicente Ordonez, Mohit Bansal, Yue Zhao