RetroThinker is a multi-stage post‑training framework that enhances SpeechLLMs by enabling them to self‑verify and forward‑correct Chain‑of‑Thought reasoning steps during inference. It combines supervised fine‑tuning on curated retrospective thinking data with length‑based direct preference optimization to improve reasoning while the user speaks. On the GSM8K benchmark, RetroThinker achieves an 11% absolute accuracy gain over non‑retrospective baselines while maintaining comparable latency.
By Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed, David Harwath
arXiv:2609.22697v1 Announce Type: new
Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...
By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv:2609.20849v1 Announce Type: new
Abstract: Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) re...
By Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli
arXiv:2505.16782v3 Announce Type: replace
Abstract: Large Language Models (LLMs) have shown impressive performance on complex tasks through Chain-of-Thought (CoT) reasoning. However, conventional CoT...
By Xinghao Chen, Anhao Zhao, Heming Xia, Xuan Lu, Hanlin Wang, Yanjun Chen, Wei Zhang, Jian Wang, Wenjie Li, Xiaoyu Shen
arXiv:2607. 21550v1 Announce Type: new Abstract: While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data.
By Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin
arXiv:2606. 14591v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet they still struggle with complex audio reasoning.
By Hui Geng, Yi Su, Han Yin, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Zijian Gao, Hengzhu Liu, Xie Chen, Kele Xu