arXiv:2607. 16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues
arXiv:2607. 16292v1 Announce Type: cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues
arXiv:2606. 11930v1 Announce Type: cross Abstract: Predicting psychological traits from asynchronous video interviews (AVIs) is a challenging multimodal learning problem because labeled datasets are limited while each response contains high-dimensional visual, acoustic, and verbal signals.
By Kuo-En Hung, Hung-Yue Suen, Shih-Ching Yeh, Hsiang-Wen Wang
arXiv:2609.39080v1 Announce Type: cross
Abstract: Intracortical motor decoders degrade across sessions because the set of recorded units changes and persisting units can alter how their firing relate...
By Xinyuan Zhang, Handong Mo, Pengfei Wen, Shuang Liang, Jichang Yang, Yan Zeng, Zhongrui Wang, Han Wang
arXiv:2607. 25961v1 Announce Type: cross Abstract: Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change.
By Podakanti Satyajith Chary, Barath Parthiban, Pranesh Velmurugan, Adeeba Khan, Nagarajan Ganapathy
R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.
By Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao
arXiv:2602.10639v2 Announce Type: replace
Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what m...
By Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong
arXiv:2609.31654v2 Announce Type: replace-cross
Abstract: Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mec...
By Taewoo Ha, Shafayat Mowla Anik, Dae Yeol Lee, Byeong Kil Lee, Jeeho Ryoo
arXiv:2609.15128v1 Announce Type: new
Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...
By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou
The study investigates how internal representations of video diffusion models align with human visual cortex responses. It finds that representations used for future video generation in an autoregressive (AR) model better match cortical activity than those for observed video, with future‑generation alignment concentrated in higher‑order visual areas. A behavioral experiment further shows that humans prefer videos enhanced by layers that align more strongly with cortical responses.
By Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim
arXiv:2606. 00129v1 Announce Type: cross Abstract: Large language models (LLMs) have emerged as powerful representation learners whose internal features increasingly align with human cognition.
By Yousef A. Radwan, Xuhui Liu, Kilichbek Haydarov, Yuqian Fu, Mohamed Elhoseiny
arXiv:2606. 06345v1 Announce Type: cross Abstract: Brain decoding is limited by the availability of labeled neural data, and remains challenging in low-data regimes.
By Yohann Benchetrit, Marl\`ene Careil, Simon Dahan, Hubert Banville, St\'ephane d'Ascoli, Jean-R\'emi King