arXiv:2607. 05365v1 Announce Type: cross Abstract: Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech.
By Thomas Thebaud, Yuzhe Wang, Hao Zhang, Sathvik Manikantan Napa Ugandhar, Ashish Hallur, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text benchmarks do not capture whether these systems behave naturally in conversations, where timing, turn-taking, prosody, interpersonal stance, language and dialect consistency, and relationship-aware appropriateness jointly shape perceived quality.
The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.
By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
arXiv:2608. 19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged.
By Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions.
arXiv:2608.29206v1 Announce Type: new
Abstract: Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a wa...
By Qianqi Liu, Jin Huang, Fethiye Irmak Dogan, Hatice Gunes
The paper introduces UPHELD, a large benchmark of human-to-human dialogues written by professional script writers, featuring realistic turn densities and over 36,000 per-turn human annotations. It evaluates existing automatic metrics and LLM-as-a-judge methods, finding them unreliable against expert human judgment. Using UPHELD, the authors develop a Mixture-of-Judges framework that improves correlation with human assessments by about 30%.
By Ilija Subasic, Andrew Rabinovich, Zhao Chen
The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors compare seven LLMs to human annotations. They find that models agree more with each other than with humans, show systematic neutral bias in categorical predictions, and that alignment varies with explicit linguistic cues and rapport‑building strategies.
By Rong Wang, Kun Sun, Yadong Guo
The paper investigates whether large language models (LLMs) assess politeness in ways that match human judgments. Using two English datasets—one with continuous ratings and another with three‑way categorical labels—the authors find that LLMs agree more with each other than with humans. They observe that model–human alignment depends on explicit linguistic cues, while misaligned cases often involve rapport‑building strategies. Additionally, models tend to overproduce Neutral labels and underpredict Impolite labels, a pattern that persists even when expert consensus is used as a reference.
RoleBreak is an open benchmark designed to evaluate long‑horizon role‑playing robustness in spoken dialogue systems. It includes 310 character‑based and user‑centered roles, 6,688 human‑verified dialogue turns, and 11,743 fine‑grained evaluation criteria, with 1,856 turns specifically targeting expressive vocal emotion. The benchmark stresses role consistency, interaction quality, safety, and affect over extended conversations, and the authors evaluated nine system configurations across full‑duplex, omni‑modal, and cascaded ASR–LLM–TTS paradigms.
By Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu
arXiv:2603. 23841v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact their objectivity.
By Rohan Khetan, Ashna Khetan
arXiv:2608. 10810v1 Announce Type: cross Abstract: Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, implicit, polite, ironic, or deliberately mismatched expressions.
By Zhenyan Zheng, Yunyao Zhang, Junxi Sheng, Junqing Yu, Zikai Song