arXiv AI By Chia-Yuan Chang, Renyuan Cheng, Rui Feng, Xiaotian Han, Yuan He, Hongye Jin, Linwei Li, Shiyang Li, Fenglin Liu, Xin Liu, Priyanka Nigam, Haoyang Wen, Zhenghao Xu, Zhuocheng Xu, Bing Yin, Qingyu Yin, Chao Zhang, Rongzhi Zhang, Zhihan Zhang, Zixuan Zhang, Zixuan Zhang, Tuo Zhao

Rufus-Air: An Open LLM Post-Training Recipe

Read the original on arXiv AI →

Rufus-Air is an open, reproducible post‑training recipe for the GLM‑4.5‑Air‑Base model, structured as a serial pipeline of eight stages: SFT, Reasoning RL, Coding RL, Instruction‑Following RL, General Agent, Coding Agent, Search Agent, and RLHF. The paper documents the data, reward design, infrastructure, stage order, and stagewise results required to reproduce the recipe, emphasizing that stages progress from basic to advanced capabilities and from hard, verifiable rewards to softer judge‑based signals. Key findings highlight the importance of diverse, high‑quality SFT, difficulty filtering, reward reliability for stage ordering, and the role of infrastructure and engineering choices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 26

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.

By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
arXiv AI
Sep 15

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1 is a 7B dense foundation model trained from scratch with extreme data, system, and algorithmic efficiency. It uses a core premise that compact models can overcome capacity limits by combining deliberate internal thinking with active external tool use, supported by a 256K context and an end‑to‑end high‑efficiency training recipe that includes interleaved gated sliding‑window and full attention, a stable FP8 Muon optimizer, progressive curriculum scaling, and reformulation of interaction traces into Markov Decision Processes. The model is competitive with much larger frontier models on challenging mathematical reasoning and agentic search tasks, offers a ~4.2× efficiency improvement in pre‑training time‑to‑loss, and its weights, checkpoints, training code, data recipes, and logs are fully open‑source to support community research.

By Jiyan He, Guang Liang, Hao Liu, Haoxiang Guan, Jinbo Sun, Junyi Guo, Wenjun Feng, Yantai Xie, Yifei Shen, Bin Shao, Chuyang Wei, Kai Chen, Kexin Zhou, Minghang Zhu, Shuxin Zheng, Tie-Yan Liu, Taine Zhao, Wenhui Zhu, Xueyin Xu, Xiaoqing Zhang, Yatao Li, Yuxuan Ren
Hugging Face Trending Papers
2d ago

Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries

The paper introduces Guidance‑TTT, a method that separates strategic planning from execution in test‑time training for large language models. A small guidance model is trained at test time to propose high‑level changes, while a frozen, larger execution model implements these changes, reducing the cost of maintaining gradients and optimizer states. Guidance‑TTT achieves strong results across four domains—combinatorial optimization, heuristic programming, machine learning, and GPU kernel optimization—outperforming prior work and matching state‑of‑the‑art leaderboard scores.

arXiv AI
Sep 25

iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model

The paper introduces iCoder-27B, a 27‑billion‑parameter model for RTL design and GPU kernel optimization that is developed through a recursive AI‑led process with minimal human input. Human experts provide high‑level objectives and reusable research skills, while the agent autonomously selects experiments, diagnoses outcomes, and refines training strategies, coordinating SFT, self‑distillation, and reinforcement learning. iCoder outperforms GPT‑5.5 and Claude‑Opus‑4.8 on several benchmarks, demonstrating the feasibility of building frontier‑competitive models with largely automated development.

By Cheng Yang, Jiayang Lyu, Shangyuan Liu, Guibin Zhang, Jiong Lin, Xinlei Yu, Junchi Yan, Shuicheng Yan, Weinan E, Linfeng Zhang, Linfeng Zhang, Qibing Ren
arXiv AI
Sep 2

Instella-MoE Technical Report

Instella‑MoE is a fully open Mixture‑of‑Experts language model with 16 billion total parameters and 2.8 billion active parameters per token, trained from scratch on AMD Instinct GPUs. It incorporates a sparsely activated MoE design with Gated Multi‑head Latent Attention and FarSkip‑Collective connectivity, and follows a multi‑stage pipeline that includes pre‑training, long‑context extension, supervised fine‑tuning, direct preference optimization, and reinforcement learning with Multi‑Teacher On‑Policy Distillation. The model achieves an average score of 76.7 on pre‑training benchmarks and 73.2 on instruction‑following, reasoning, math, coding, and chat benchmarks, outperforming comparable fully open and open‑weight models, and its full training pipeline, weights, and code are released for reproducibility.

By Jiang Liu, Sudhanshu Ranjan, Prakamya Mishra, Yonatan Dukler, Gowtham Ramesh, Jialian Wu, Ximeng Sun, Wen Xie, Chaojun Hou, Vikram Appia, Zhenyu Gu, Zicheng Liu, Emad Barsoum