arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.
By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
ZGCM-1 is a 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. It uses a core premise that compact models can overcome capacity limits by combining deliberate internal thinking with active external tool use, supported by a 256K context and an end‑to‑end high‑efficiency training recipe that includes interleaved gated sliding‑window and full attention, a stable FP8 Muon optimizer, progressive curriculum scaling, and reformulation of interaction traces into Markov Decision Processes. The model is competitive with much larger frontier models on challenging mathematical reasoning and agentic search tasks, offers a ~4.2× efficiency improvement in pre‑training time‑to‑loss, and its weights, checkpoints, training code, data recipes, and logs are fully open‑source to support community research.
By Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren
The paper introduces Guidance‑TTT, a method that separates strategic planning from execution in test‑time training for large language models. A small guidance model is trained at test time to propose high‑level changes, while a frozen, larger execution model implements these changes, reducing the cost of maintaining gradients and optimizer states. Guidance‑TTT achieves strong results across four domains—combinatorial optimization, heuristic programming, machine learning, and GPU kernel optimization—outperforming prior work and matching state‑of‑the‑art leaderboard scores.
arXiv:2607. 07646v1 Announce Type: new Abstract: Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies?
By Azwar Abdulsalam, Nishil Patel, Andrew Saxe
The paper introduces iCoder-27B, a 27‑billion‑parameter model for RTL design and GPU kernel optimization that is developed through a recursive AI‑led process with minimal human input. Human experts provide high‑level objectives and reusable research skills, while the agent autonomously selects experiments, diagnoses outcomes, and refines training strategies, coordinating SFT, self‑distillation, and reinforcement learning. iCoder outperforms GPT‑5.5 and Claude‑Opus‑4.8 on several benchmarks, demonstrating the feasibility of building frontier‑competitive models with largely automated development.
By Cheng Yang, Jiayang Lyu, Shangyuan Liu, Guibin Zhang, Jiong Lin, Xinlei Yu, Junchi Yan, Shuicheng Yan, Weinan E, Linfeng Zhang, Linfeng Zhang, Qibing Ren
Instella‑MoE is a fully open Mixture‑of‑Experts language model with 16 billion total parameters and 2.8 billion active parameters per token, trained from scratch on AMD Instinct GPUs. It incorporates a sparsely activated MoE design with Gated Multi‑head Latent Attention and FarSkip‑Collective connectivity, and follows a multi‑stage pipeline that includes pre‑training, long‑context extension, supervised fine‑tuning, direct preference optimization, and reinforcement learning with Multi‑Teacher On‑Policy Distillation. The model achieves an average score of 76.7 on pre‑training benchmarks and 73.2 on instruction‑following, reasoning, math, coding, and chat benchmarks, outperforming comparable fully open and open‑weight models, and its full training pipeline, weights, and code are released for reproducibility.
By Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum
arXiv:2609.22592v1 Announce Type: cross
Abstract: Training agents with reinforcement learning requires a gym, comprising a task, an executable environment in which the task can be attempted, and a ve...
By Aarati Andrea Noronha, Kavya Ravikumar, Carly Xiaoyu Lin
arXiv:2604. 28123v3 Announce Type: replace-cross Abstract: The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR).
By Sudong Wang, Weiquan Huang, Xiaomin Yu, Zuhao Yang, Hehai Lin, Keming Wu, Chaojun Xiao, Chen Chen, Wenxuan Wang, Beier Zhu, Yunjian Zhang, Chengwei Qin
arXiv:2605.10194v2 Announce Type: replace
Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own r...
By Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Lan-Zhe Guo
arXiv:2608.23830v1 Announce Type: cross
Abstract: RL has emerged as a powerful paradigm for enhancing the instruction following capabilities of LLMs. While existing training recipes achieve substanti...
By Mian Zhang, Yueqin Yin, Kaiyu He, Peilin Wu, Xinlu Zhang, Mingyuan Zhou, Zhiyu Zoey Chen
arXiv:2609.08368v1 Announce Type: new
Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...
By RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng
arXiv:2609.16059v1 Announce Type: cross
Abstract: Multimodal instruction following (MMIF) is crucial for building generalist agents. However, current training paradigms rely heavily on Supervised Fin...
By Yirong Zeng, Zhang Sai, Yuxian Wang, Yutai Hou, Yufei Liu, Xiao Ding, Bibo Cai