Don't Wait to Reply: Towards Responsive yet Thoughtful Dialogue through Proactive Thinking
arXiv:2607. 03093v1 Announce Type: cross Abstract: Thinking has emerged as a critical capability for Large Language Models (LLMs) tackling complex tasks.
arXiv:2607. 03093v1 Announce Type: cross Abstract: Thinking has emerged as a critical capability for Large Language Models (LLMs) tackling complex tasks.
arXiv:2510. 05150v3 Announce Type: replace-cross Abstract: Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, where the models continuously perceive user speech streams while generating responses.
RetroThinker is a multi-stage post‑training framework that enhances SpeechLLMs by enabling them to self‑verify and forward‑correct Chain‑of‑Thought reasoning steps during inference. It combines supervised fine‑tuning on curated retrospective thinking data with length‑based direct preference optimization to improve reasoning while the user speaks. On the GSM8K benchmark, RetroThinker achieves an 11% absolute accuracy gain over non‑retrospective baselines while maintaining comparable latency.
arXiv:2609.37818v1 Announce Type: cross Abstract: Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguis...
Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbalized mathematical expressions, which are harder to interpret than symbolic text.
arXiv:2606. 07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs.
arXiv:2608.20359v1 Announce Type: new Abstract: Large language models (LLMs) are deployed for increasingly complex tasks involving planning and multi-step decision making, but high-quality performanc...
AURAL is a speech language model that performs adaptive latent reasoning by modeling multiple plausible reasoning continuations in latent space and jointly predicting chunks of future states, thereby reducing sequential forward passes and latency. The authors introduce a large bilingual dataset, AuralReason-683K, containing concise chain‑of‑thought annotations for emotion recognition, empathetic dialogue, and general reasoning, and use reinforcement learning (AURAL‑RL) to reward concise, high‑quality reasoning that adapts to problem difficulty. Experiments on two backbones show that AURAL‑RL matches or exceeds chain‑of‑thought reinforcement learning while achieving significant latency reductions, such as an 11.8× speed‑up on Qwen2.5‑Omni. "whyItMatters":"The work demonstrates that latent reasoning can match the performance of explicit chain‑of‑thought methods while dramatically cutting response time, addressing the trade‑off between intelligence and speed in speech language models."
arXiv:2608. 19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged.
arXiv:2609.08977v3 Announce Type: replace-cross Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
arXiv:2606. 05121v1 Announce Type: cross Abstract: Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting.
Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged. Yet existing benchmarks typically evaluate prosodic perception, response appropriateness, and task-oriented dialogue in isolation, making it difficult to test whether prosodic evidence changes downstream decisions.