arXiv:2601. 01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos remains underexplored.
By Hyeonjeong Ha, Jinjin Ge, Bo Feng, Kaixin Ma, Gargi Chakraborty
As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the judgment crisis. While human perception naturally synthesizes the temporal and logical flow of a story, automated evaluation systems remain largely "blind" to sequential continuity, often failing to distinguish between a coherent narrative and a semantically shuffled or contradictory sequence.
arXiv:2608.29621v1 Announce Type: cross
Abstract: Long-horizon story-driven video generation requires a production agent to coordinate narrative decomposition, state tracking, shot design, prompt con...
By Junxiang Liu, Lin Wang, Haiyu Shi, Hongxu Ma, Xiaoyu Yang, Chunjie Chen, Xiaoxiao Xu, Kaiqiao Zhan, Boao Wang, Shuizhou Shi, Tianyun Zhu, Jie Li, Jiangtong Li
Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated...
arXiv:2608.28633v1 Announce Type: cross
Abstract: Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans,...
By Taaha Kazi, Vasu Sharma, Mohammad Saifullah, Abdur Rahman
arXiv:2408.11827v2 Announce Type: replace
Abstract: Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies...
By Nura Aljaafari, Danilo S. Carvalho, Andr\'e Freitas
arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.
By Danae S\'anchez Villegas, Samuel Lewis-Lim, Nikolaos Aletras, Desmond Elliott
arXiv:2608.22963v1 Announce Type: new
Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool...
By Yuchen Huang, Sijia Li, Jun Zhang, Yi R. Fung
arXiv:2609.02000v1 Announce Type: new
Abstract: Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess fina...
By Shuyao Xiao, Shengling Wang, Haoyu Niu, Ke Chao, Changwei Xu, Xinran Duan, Chaoyong Jiang
The paper investigates how listeners decide to preserve or revise their understanding when a speaker refers to a scene they cannot see. It introduces a data‑driven framework for turn‑by‑turn preserve/revise decisions and compares four theory‑driven revision strategies. The study finds that a mismatch‑driven policy destabilizes grounding, while an uncertainty‑sensitive policy balances preservation and revision, leading to coherent understanding that aligns with conceptual pact theory.
By Ziming Liu, Bhanu Chaitanya Jasti, Ziyang Xu, Hongyu Wu, Yi Wu, Jiqun Liu
Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how diff...
arXiv:2606. 27598v1 Announce Type: cross Abstract: Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long tail.
By Mreedul Gupta, Advait Deshmukh, Ashwin Umadi, Matt Pauk, Maria Leonor Pacheco