arXiv Machine Learning

Multimodal Evaluator Preference Collapse: Cross-Modal Contagion in Self-Evolving Agents

arXiv:2606. 16682v1 Announce Type: new Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge.

arXiv Machine Learning
Aug 7

PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.

By Ayushi Agarwal
arXiv Machine Learning
Jun 25

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

arXiv:2606. 26079v1 Announce Type: cross Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines.

By Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli