arXiv AI

YNU-HPCC at SemEval-2025 Task 11: Bridging the Gap in Text-Based Emotion Using Multiple Prediction Headers

The YNU-HPCC team participated in Subtask A of SemEval‑2025 Task 11, "Bridging the Gap in Text‑Based Emotion," using a RoBERTa model with a single prediction head to process one emotion at a time. Their system achieved an official ranking score of 0.44 across all languages after translating the dataset into English with Google Translate. Analysis showed that a single head outperformed six simultaneous heads and that training on the uniformly translated English data improved results.

arXiv AI
Aug 11

IndexTTS 2.5 Technical Report

arXiv:2601. 03888v4 Announce Type: replace-cross Abstract: In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-based Text-to-Semantic (T2S) module and a non-autoregressive Semantic-to-Mel (S2M) module, which together enable faithful emotion replication and establish the first autoregressive duration-controllable generative paradigm.

By Yunpei Li, Xun Zhou, Jinchao Wang, Lu Wang, Yong Wu, Siyi Zhou, Yiquan Zhou, Yining Wang, Yaogen Yang, Zhetao Hu, Shiyao Duan, Jiacheng Xu, Bin Xia, Jingchen Shu
arXiv AI
Sep 18

Reading Emotions in the Token Space: Discriminative Adaptation of SpeechLLMs for Emotion Recognition

The paper introduces a discriminative adaptation for SpeechLLMs that reads the hidden state of the final prompt token via a simple classification head, enabling emotion recognition in a single forward pass without altering the backbone. This approach replaces the generative decoder, which can produce out‑of‑set labels and favor frequent classes, with a controlled comparison between generative and discriminative inference. Experiments on IEMOCAP show improved Macro F1 scores, elimination of hallucinations, and larger gains on realistic ASR transcripts, while revealing that emotion directions encode indirect associations reflecting web‑scale text biases.

By Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv Computation and Language
Sep 3

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

TalkFa introduces a unified benchmark for Farsi dialogue generation and understanding, comprising three datasets: WIKI‑FADIAL (4.2K Wikipedia‑grounded dialogues), DAILYDIALOG‑FA (6.6K dialogues with dialogue‑act and emotion annotations), and PLAYDIAL‑FA (2.1K theatrical dialogues with sentiment labels). All dialogues are curated through multi‑stage review by native speakers, ensuring high quality. Experiments show that LoRA fine‑tuning improves generation performance with less data, while specific models excel on classification tasks, and human evaluation confirms the benchmark’s reliability.

By Neda Jamshidi, Kamyar Zeinalipour, Fahimeh Akbari, Monica Bianchini, Marco Maggini, Marco Gori
Hugging Face Trending Papers
Jul 9

Cross-seed explainability using Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoders

We present a Procrustes-conditioned Joint End-to-end Top-K Sparse Autoencoder (SAE) for extracting cross-seed universal features from independently trained BERT models. Cross-seed feature universality is a fundamental challenge in mechanistic interpretability: because dictionary learning is non-convex, independently trained networks learn misaligned feature spaces, so apparently identical features may differ by random initialization.

Hugging Face Trending Papers
Aug 11

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations.

arXiv Computation and Language
Sep 4

Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory

The paper introduces CHIARO, a 1,000-sentence benchmark for contrastive emotion inference grounded in appraisal theory, where each scenario elicits a positive emotion in one person and a negative emotion in another. The dataset covers ten emotion classes and is human‑annotated. Evaluation shows that the best large language model achieves 67.3 macro‑F1, below human agreement, while existing emotion classifiers perform near chance. When used as a training signal alongside an existing emotion corpus, models improve on CHIARO and on six of ten external emotion benchmarks, demonstrating its value as a complementary training resource.

By Divyesh Bommana, Mohammad Saim, Tianyu Jiang