AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
Read the original on arXiv Machine Learning →AURAL is a speech language model that performs adaptive latent reasoning by modeling multiple plausible reasoning continuations in latent space and jointly predicting chunks of future states, thereby reducing sequential forward passes and latency. The authors introduce a large bilingual dataset, AuralReason-683K, containing concise chain‑of‑thought annotations for emotion recognition, empathetic dialogue, and general reasoning, and use reinforcement learning (AURAL‑RL) to reward concise, high‑quality reasoning that adapts to problem difficulty. Experiments on two backbones show that AURAL‑RL matches or exceeds chain‑of‑thought reinforcement learning while achieving significant latency reductions, such as an 11.8× speed‑up on Qwen2.5‑Omni. "whyItMatters":"The work demonstrates that latent reasoning can match the performance of explicit chain‑of‑thought methods while dramatically cutting response time, addressing the trade‑off between intelligence and speed in speech language models."
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.