arXiv AI

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

arXiv:2606. 06168v1 Announce Type: new Abstract: We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline.

arXiv AI
Jul 28

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

arXiv:2607. 22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited.

By Yuzhe Wang (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Thomas Thebaud (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Jennifer Hu (Department of Cognitive Science, Johns Hopkins University, Baltimore, USA), Jes\'us Villalba-Lopez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Venkatesh Ravichandran (Amazon AGI, USA), Georgi Tinchev (Amazon Research, UK), Najim Dehak (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Laureano Moro-Vel\'azquez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA)