Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,220 stories · RSS feed

arXiv AI
Aug 5

ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs

arXiv:2608. 03006v1 Announce Type: new Abstract: Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs and to discourage contradictory reverse predictions.

By Xinghe Cheng, Jiapu Wang, Chaobo He, Ruihai Dong, Quanlong Guan
arXiv AI
Aug 5

Multi-Task Multi-Frame Visual Piano Transcription

arXiv:2608. 03419v1 Announce Type: cross Abstract: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release.

By Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch
arXiv AI
Aug 5

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

arXiv:2608. 03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate.

By Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo
arXiv AI
Aug 5

SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

arXiv:2608. 02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attention maps provide only local explanations: they indicate where a model focuses but not which histological features drive its predictions or how the model behaves across a patient cohort.

By Abdallah Lamane, Abdul Rahman Diab, Ren-Chin Wu, William Lotter