Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

14,048 stories · RSS feed

arXiv Machine Learning
Aug 5

Latent Reward Registers for Diffusion Preference Alignment

arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.

By Yuanshen Guan, Zipeng Feng, Zhiwei Xiong, Peiqin Sun
arXiv Machine Learning
Aug 5

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

arXiv:2608. 03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attributes even when they are never stated explicitly.

By Jong Wook Kim, Byoungjae Min, Kennedy Edemacu, Yoonhyuk Choi, Sae-Hong Cho, Beakcheol Jang
arXiv Machine Learning
Aug 5

Conformal risk control for model-form uncertainty in parametric non-intrusive reduced-order models

arXiv:2608. 03360v1 Announce Type: cross Abstract: Non-intrusive reduced-order models (NIROMs) have become a standard tool for approximating parametric partial differential equations from computer design of experiments while significantly reducing computational costs.

By Edgar Jaber (CB, ENS Paris Saclay), R\'emy Vallot (CB, Michelin), Thibault Dairay (CB, Michelin), Mathilde Mougeot (CB, ENSIIE, ENS Paris Saclay)
arXiv AI
Aug 5

ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs

arXiv:2608. 03006v1 Announce Type: new Abstract: Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction, limiting their ability to adaptively integrate complementary educational evidence for individual candidate pairs and to discourage contradictory reverse predictions.

By Xinghe Cheng, Jiapu Wang, Chaobo He, Ruihai Dong, Quanlong Guan
arXiv AI
Aug 5

Multi-Task Multi-Frame Visual Piano Transcription

arXiv:2608. 03419v1 Announce Type: cross Abstract: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release.

By Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch
arXiv AI
Aug 5

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

arXiv:2608. 03733v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate.

By Chunyang Jiang, Pingping Zhang, Yuzhi Zhao, Wenao Ma, Zhijian Hou, Mengyang Wu, Yiyang Cai, Senkang Hu, Sitong Cheng, Chi-Min Chan, Wei Xue, Yike Guo