arXiv:2606. 27717v1 Announce Type: cross Abstract: Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech.
By Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang, Haonan Chen, Zeyu Jin
The study examines how emotions are represented across layers of large language models (LLMs) by probing eight 1B–9B open‑weight models on three datasets (Twitter, Reddit, autobiographical narratives). It finds that the optimal probing layer varies systematically with the dataset, moving from near‑input layers to deeper layers, and that targeted forward‑pass interventions on these layers degrade performance more than random interventions. Additionally, the selected layers transfer across datasets and emotion categories, and early‑exit representations from these layers outperform full‑depth exits by an average of 6.9 percentage points.
By Tian Fang, Ga\"el Guibon, Davide Buscaldi
arXiv:2604. 15280v2 Announce Type: replace-cross Abstract: Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans.
By Madhav Agarwal, Sotirios A. Tsaftaris, Laura Sevilla-Lara, Steven McDonagh
arXiv:2605. 31220v2 Announce Type: replace-cross Abstract: Confidence estimation (CE), i.
By Athina Kyriakou, Dennis Ulmer, Ivan Titov
arXiv:2606. 15694v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content.
By Hangling Xie
arXiv:2602. 23197v2 Announce Type: replace-cross Abstract: Transformer-based large language models exhibit in-context learning, enabling adaptation to downstream tasks via few-shot prompting with demonstrations.
By Chungpa Lee, Jy-yong Sohn, Kangwook Lee
arXiv:2604. 03532v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult.
By Sing Hieng Wong, Hassan Sajjad, A. B. Siddique
The paper introduces a method that adapts Contextual Decomposition for Transformers (CD‑T) to discover neural circuits without relying on clean counterfactuals, using label‑balanced activation means and task‑directional relevance scoring. These circuits are then employed in Circuit‑Targeted Supervised Fine‑Tuning (CT‑SFT), which restricts parameter updates to task‑relevant heads and LayerNorm, leading to competitive performance on low‑resource language adaptation tasks such as NusaX cross‑lingual sentiment transfer. CT‑SFT consistently avoids catastrophic forgetting and preserves source‑language and related‑task performance, offering a more controlled alternative to global fine‑tuning, as further validated on the XNLI benchmark.
By Khumaisa Nur'aini, Ayu Purwarianti, Alham Fikri Aji, Derry Wijaya
arXiv:2608.30462v1 Announce Type: cross
Abstract: Large language models exhibit substantial performance variation across languages, even when solving semantically equivalent tasks. Existing analyses...
By Minju Song, Hyeon Hwang, Junhyun Lee, Jaewoo Kang
arXiv:2606. 07479v1 Announce Type: cross Abstract: Turkish idiomatic light verb constructions (LVCs) are challenging for multiword expression processing because they often share the same surface form as fully literal verb-object combinations while functioning as a single, partially idiomatic predicate.
By Sercan Karaka\c{s}, Yusuf \c{S}im\c{s}ek
The paper introduces PM4Bench, a multimodal, multilingual, multi-task benchmark built on a strictly parallel 10‑language corpus, allowing fair cross‑lingual comparison of Large Vision‑Language Models (LVLMs). It also proposes a vision setting that embeds textual inputs directly into images to better mimic real deployment scenarios. Experiments show OCR performance drives cross‑lingual gaps, leading to an OCR‑centric GRPO training strategy that improves multilingual VQA and reduces disparities without costly supervision.
By Junyuan Gao, Jiahe Song, Jiang Wu, Runchuan Zhu, Guanlin Shen, Shasha Wang, Xingjian Wei, Haote Yang, Weijia Li, Bin Wang, Lijun Wu, Conghui He
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (MER) and enabled interpretable description generation by jointly modeling video, audio, and language, etc. However, these performance improvements are often accompanied by an increase in model parameter size (e.