arXiv AI

Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue

arXiv:2608. 08210v1 Announce Type: new Abstract: Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating an \textbf{illusion of alignment (IoA)}.

arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv AI
Aug 28

When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.

By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
arXiv AI
Sep 4

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to give large language model (LLM) agents situational awareness in online meetings. CAPA uses a Perceiver to update meeting state, a Predictor to forecast conversation flow, a Controller to decide speaking actions, and a Generator to phrase contributions. Evaluated on 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination, demonstrating that structured state tracking is key to effective delegation.

By Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp
arXiv AI
4d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
Hugging Face Trending Papers
Sep 3

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to improve large language model (LLM) participation in online meetings. CAPA tracks meeting state with a Perceiver, predicts conversation flow, decides when and what to speak, and generates contributions in the participant’s style, all while being calibrated by judges. In experiments on 137 AMI meetings, CAPA cuts the LLM’s silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination.

arXiv Computation and Language
Sep 18

Towards Proactive Detection of User-Side Implicit Conflicts in Human-LLM Dialogue

The paper introduces UC-Bench, a human‑annotated benchmark for detecting user‑side implicit conflicts in Human‑LLM dialogue, a problem largely overlooked compared to LLM‑side conflicts. Experiments show current LLMs struggle with these conflicts, especially when they stem from implicit incompatibilities in dialogue history. To address this, the authors propose SynUC, a constraint‑guided data synthesis method that generates a new training set, UC‑Data, which improves performance of lightweight LLMs on UC‑Bench compared to larger general‑purpose models and existing synthesis approaches.

By Jinqiang Wang, Tao Zhu, Huansheng Ning