Hugging Face Trending Papers

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models

Read the original on Hugging Face Trending Papers →

Spoken language models (SLMs) enable natural human-computer interaction, but their reasoning ability still lags behind that of text-based large language models, especially on spoken mathematical question answering tasks. One important reason is that SLMs reason over purely verbalized mathematical expressions, which are harder to interpret than symbolic text.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Jun 15

Towards Efficient Large Language Reasoning Models via Extreme-Ratio Chain-of-Thought Compression

arXiv:2602. 08324v5 Announce Type: replace Abstract: Chain-of-Thought (CoT) reasoning successfully enhances the reasoning capabilities of Large Language Models (LLMs), yet it incurs substantial computational overhead for inference.

By Yuntian Tang, Bohan Jia, Wenxuan Huang, Lianyue Zhang, Jiao Xie, Wenxi Li, Wei Li, Jie Hu, Xinghao Chen Rongrong Ji, Shaohui Lin
arXiv Computation and Language
Sep 11

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

RetroThinker is a multi-stage post‑training framework that enhances SpeechLLMs by enabling them to self‑verify and forward‑correct Chain‑of‑Thought reasoning steps during inference. It combines supervised fine‑tuning on curated retrospective thinking data with length‑based direct preference optimization to improve reasoning while the user speaks. On the GSM8K benchmark, RetroThinker achieves an 11% absolute accuracy gain over non‑retrospective baselines while maintaining comparable latency.

By Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath
arXiv AI
Jul 29

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.

By Yutong Chen, Shouqian Shi, Xinran Liu, Haochen Wang, Jiaying Wang, Tianxing Xu, Yuanxi Wang, Zirui Ding
arXiv Machine Learning
23h ago

AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models

AURAL is a speech language model that performs adaptive latent reasoning by modeling multiple plausible reasoning continuations in latent space and jointly predicting chunks of future states, thereby reducing sequential forward passes and latency. The authors introduce a large bilingual dataset, AuralReason-683K, containing concise chain‑of‑thought annotations for emotion recognition, empathetic dialogue, and general reasoning, and use reinforcement learning (AURAL‑RL) to reward concise, high‑quality reasoning that adapts to problem difficulty. Experiments on two backbones show that AURAL‑RL matches or exceeds chain‑of‑thought reinforcement learning while achieving significant latency reductions, such as an 11.8× speed‑up on Qwen2.5‑Omni. "whyItMatters":"The work demonstrates that latent reasoning can match the performance of explicit chain‑of‑thought methods while dramatically cutting response time, addressing the trade‑off between intelligence and speed in speech language models."

By Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu