arXiv Machine Learning By Zewen Liu

Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agents

Read the original on arXiv Machine Learning →

arXiv:2606. 16682v3 Announce Type: replace Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

When Seeing Overrides Knowing: Visual Dominance and Deferral-Based Method for Personalized Safety in VLMs

The paper introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile, to evaluate personalized safety in vision‑language models (VLMs). Eight leading VLMs were tested and found to almost always respond directly (86‑99%) without seeking missing context, scoring no higher than 2.6/5 on personalized safety. The authors identify a phenomenon called visual dominance, where visual information enters text representations early and suppresses textual risk signals, and propose PRISM, a lightweight input monitor that predicts when a query should be deferred, achieving 0.978 AUC and outperforming all tested models on the safety‑utility Pareto frontier.

By Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan
arXiv Computation and Language
Sep 22

Read-Best Is Not Steer-Best: A Probing--Steering Layer Dissociation in Omni-Modal Large Language Models

The paper investigates whether the layer that yields the highest probing accuracy in omni‑modal large language models is also the most effective for steering interventions. Across three independently developed models, the authors find that the best probing layers differ widely, whereas the most steerable layers consistently lie in a narrow mid‑to‑late range of the network. Using emotion as a testbed, they demonstrate a significant causal gap between probing and steering, and propose a two‑factor account linking readability and downstream plasticity to steering effectiveness.

By Yibo Wang, Jisheng Dang, Bimei Wang, Yitao Wu, Wencan Zhang, Hong Peng, Jizhao Liu, Bin Hu, Qi Tian, Tat-Seng Chua