arXiv Machine Learning By Hongbo Jiang, Jie Li, Yunhang Shen, Tianyu Xie, Pingyang Dai

Diagnosing and Mitigating Perception-Decision Misalignment in Omni-LLMs via Modality Subspace Activation

Read the original on arXiv Machine Learning →

arXiv:2608. 14655v1 Announce Type: new Abstract: Omni-Large Language Models (Omni-LLMs) power complex multi-modal reasoning in applications like World Action Models and autonomous agents.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 21

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

arXiv:2606. 25325v2 Announce Type: replace Abstract: We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities.

By Zhiyuan Han, Beier Zhu, Wenwen Tong, Pengyang Shao, Peipei Song, Xinyi Wang, Jiangnan Chen, Lewei Lu, Xun Yang