arXiv Machine Learning
Sep 11

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for omni models

The paper introduces the Platonic brain bridge hypothesis, asserting that omni models—capable of jointly processing video, audio, and text—naturally develop brain‑like representations, and that this relationship is bidirectional. Experiments show that seven omni models exhibit stable brain‑like representations across participants, and their internal states outperform others on the Algonauts 2025 out‑of‑distribution leaderboard. Building on this, the authors propose three brain‑to‑model methods: Brain‑MoE assigns a brain‑pretrained expert to each of seven cortical networks, improving benchmark accuracy; Brain‑AVQA generates video‑based questions using the most responsive brain network, outperforming shuffled mappings; and Brain‑Scope uses sparse autoencoders to pinpoint a small subset of networks whose removal weakens brain prediction, demonstrating that human brain networks can serve as a practical architectural prior for omni models.

By Pengfei Zhang, Biao Tian, Xiangang Li, Li Liu
arXiv AI
Jun 24

Mind the Heads: Topological Representation Alignment for Multimodal LLMs

arXiv:2606. 23885v1 Announce Type: cross Abstract: Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder.

By Davide Caffagni, Alberto Compagnoni, Federico Melis, Sara Sarto, Pier Luigi Dovesi, Mark Granroth-Wilding, Marcella Cornia, Lorenzo Baraldi
arXiv Computation and Language
Sep 11

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

OmniHallu is a unified framework for detecting hallucinations in multimodal large language models across both comprehension and generation tasks involving image, video, and audio modalities. It introduces OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations for six cross-modal tasks (I2T, V2T, A2T, T2I, T2V, T2A). The system uses a multi‑agent architecture that decomposes outputs into atomic claims, verifies them with modality‑specific experts, and aggregates evidence through structured reasoning, while a preference‑optimized verifier reduces expert calls by 66% with minimal performance loss.

By Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
arXiv Machine Learning
Jun 30

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

arXiv:2606. 30319v1 Announce Type: cross Abstract: Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience.

By Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun, Qihao Zheng, Mianxin Liu, Chi Zhang, Wanli Ouyang, Chunfeng Song, Changqing Zhang, Jiamin Wu
arXiv Computation and Language
Sep 11

Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accounts

The study examined whether brain-language model alignment reflects shared computational mechanisms or merely stable lexical‑semantic correspondences. Using whole‑brain encoding across Mandarin, English, and French, transformer representations predicted activity in a distributed network that overlapped across languages and remained stable across layers. Contextual embeddings and measures of prediction or compression did not outperform static lexical embeddings, suggesting that alignment is robust but not informative about shared computational processes.

By Ni Yang, Rui He, Philipp Homan, Iris Sommer, Davide Staub, Wolfram Hinzen