LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2608. 10448v1 Announce Type: new Abstract: Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues.
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.
arXiv:2607. 12787v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc.
AffectOmni is a reinforcement‑learning‑trained framework that enhances multimodal large language models for affective reasoning in social and art‑related scenes. It introduces People Focus and Temporal Order rewards to prioritize people‑centric cues and structured reasoning, and uses within‑group comparative scoring for more discriminative rewards. A Thinking Summarizer converts rationales into executable evidence instructions, which are grounded into pixel‑level regions via SAM3, enabling external auditability.
arXiv:2608.21022v1 Announce Type: new Abstract: Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little consciou...
arXiv:2606. 26348v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) can process diverse inputs, e.