arXiv:2609.14896v1 Announce Type: cross
Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at infer...
By Jiayi Yuan, Hangoo Kang, James Jihao Liu, Yejin Choi, Vikram Iyer, Liwei Jiang, Natasha Jaques
arXiv:2608. 07460v1 Announce Type: cross Abstract: While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.
By Ananya Sahu, Mohit Bansal, Elias Stengel-Eskin
arXiv:2601. 03555v3 Announce Type: replace Abstract: Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning.
By Yuxuan Jiang, Francis Ferraro
arXiv:2607. 05861v1 Announce Type: cross Abstract: Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers.
By Kaishen Wang, Tong Zheng, Xuehao Cui, Ruibo Chen, Tianyi Xiong, Heng Huang
arXiv:2509. 25760v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have demonstrated strong performance on factoid question answering, they are still prone to hallucination and untruthful responses, particularly when tasks demand information outside their parametric knowledge.
By Zhepei Wei, Xiao Yang, Kai Sun, Jiaqi Wang, Rulin Shao, Jingxiang Chen, Mohammad Kachuee, Teja Gollapudi, Yiwei Liao, Nicolas Scheffer, Rakesh Wanga, Anuj Kumar, Yu Meng, Wen-tau Yih, Xin Luna Dong
The paper investigates how pretraining and midtraining enable reward-based learning by providing necessary information and computation. It analyzes sequential state computation and contextual memory, showing that task‑independent source observations resolve ambiguities in reward adaptation. Experiments on pretrained Qwen2.5 checkpoints across eight worlds demonstrate that correct source and first‑operation supervision significantly improve success rates, and that memory replay and independent confirmation further enhance performance.
By Chiwun Yang, Xiaoyu Li