Causal-EVC: Breaking Emotional Spurious Causality via Spatiotemporal Grounding and Counterfactual Intervention
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 02963v1 Announce Type: cross Abstract: Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding and generation.
The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.
The paper introduces a steering‑vector‑based causal attribution framework to study how large vision‑language models (LVLMs) translate visual input into emotional narratives. By creating a specialized dataset, the authors uncover a functional decoupling in the LVLM’s three‑stage Adapt‑Aggregate‑Execute mechanism: visual emotional cues are first aggregated in middle layers via sentiment‑specific attention heads, then translated into narrative generation in deeper layers through emotion‑general pathways. Using these insights, they regulate emotional information routing to strengthen attention flow and amplify semantic activation, achieving significant performance gains on the MER‑UniBench and reducing emotional hallucinations through inference‑time intervention.
arXiv:2508. 21010v3 Announce Type: replace-cross Abstract: Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pipelines that entangle video understanding, causal inference, and answer generation.
arXiv:2606. 09181v1 Announce Type: cross Abstract: Recent advances in video multimodal models have significantly improved VideoQA performance.
arXiv:2608.23330v1 Announce Type: new Abstract: Video understanding requires intelligent agents to transcend mere recognition of visual facts and comprehend the underlying intents behind human action...