arXiv Machine Learning By Tianlei Zhu, Zhiwei Liu, Yuyan Wang, Xiao-Yang Liu, Sophia Ananiadou

NTDH: Complex Reasoning for Comprehensive Affective Analysis

Read the original on arXiv Machine Learning →

arXiv:2608. 06425v1 Announce Type: cross Abstract: Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather than mapped directly to labels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 16

Faithful Action-unit Causal Reasoning for Counterfactually Faithful Emotion Explanations

arXiv:2606. 15779v1 Announce Type: cross Abstract: Multimodal models can name the action units (AUs) behind a facial emotion, but their AU->emotion rationales are typically plausible rather than faithful: nothing forces the AUs a model invokes to be the AUs that actually drive its prediction.

By Van Thong Huynh, Hong Hai Nguyen, Thuy Pham, Trong Nghia Nguyen, Soo-Hyung Kim
arXiv Computer Vision
4d ago

Decoding Affective Nuances: Enhancing MLLMs via Hierarchical Emotion Reasoning and Contrastive Discriminative Pruning

The paper introduces DAN, a training‑free inference‑time framework that improves affective reasoning in multimodal large language models. It combines a Hierarchical Emotional Reasoning Chain (HERC) to better capture fine‑grained visual cues and a Contrastive Discriminative Visual Pruning (CDVP) module to isolate discriminative tokens for semantically similar emotions. Experiments show significant gains, notably a +10.47% improvement on the WebEmo25 benchmark with Qwen3‑VL‑8B‑Instruct.

By Cheng Ye, Weidong Chen, Zhaobo Qi, Beier Zhu, Zhendong Mao
arXiv AI
Aug 28

AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes

AffectOmni is a reinforcement‑learning‑trained framework that enhances multimodal large language models for affective reasoning in social and art‑related scenes. It introduces People Focus and Temporal Order rewards to prioritize people‑centric cues and structured reasoning, and uses within‑group comparative scoring for more discriminative rewards. A Thinking Summarizer converts rationales into executable evidence instructions, which are grounded into pixel‑level regions via SAM3, enabling external auditability.

By Yibo Wang, Rui Yang, Jisheng Dang, Bimei Wang, Yitao Wu, Pengfei Cao, Wencan Zhang, Hong Peng, Bin Hu, Tat-Seng Chua
arXiv AI
Jul 15

Rethinking Reward Models for Multi-Domain Test-Time Scaling

arXiv:2510. 00492v3 Announce Type: replace Abstract: The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward models} that distinguish correct reasoning from flawed logic.

By Dong Bok Lee, Seanie Lee, Sangwoo Park, Minki Kang, Jinheon Baek, Dongki Kim, Dominik Wagner, Jiongdao Jin, Heejun Lee, Tobias Bocklet, Jinyu Wang, Jingjing Fu, Sung Ju Hwang, Jiang Bian, Lei Song
arXiv AI
Jul 7

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.

By Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna, Karim Bouyarmane
arXiv AI
2d ago

Dependency-Aware Reward Shaping for Agentic Reinforcement Learning

The paper introduces Dependency‑Aware Reward Shaping (DARS), a method that assigns step‑level credit in reinforcement learning by modeling task progress as a graph of predicates with prerequisite relations. Annotators mark each step’s effect on predicates, and DARS discounts verified predicates based on distance from broken prerequisites while preserving independent ones, converting these annotations into signed per‑step rewards. Experiments on five task families with models ranging from 1.5B to 8B show that DARS improves success rates by up to 10 points over GiGPO, boosts WebShop and Search‑R1 QA scores, complements AEPO on AIME24/25, and outperforms OmniOPD in tool‑free reasoning, with ablations confirming the contribution of step‑level credit, dependency attenuation, and graph topology.

By Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu