arXiv Computation and Language By Shuxing Yang, Kaihao Zhu, Junjie Yang, Rui Zhao, Junyao Wu, Yize Wang, Wenhao Li, Fujia Chen, Taowen Deng, Shenzhan Hong, Yaqi Li, Zichen Li, Jincheng Mi, Yuang Pan, Hongsheng Chen, Yihao Yang

Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

Read the original on arXiv Computation and Language →

The paper reports a data‑efficient language modeling study conducted by Qiushi Engine on the BabyLM 2026 Strict‑Small benchmark, using only 10 million corpus words and 100 million cumulative presentations. It describes a three‑stage research program: Stage I built a frontier model via compact restatements and incremental learning; Stage II identified that exact repetition versus aligned restatement affect context use and proposed a principle for organizing experience around contextual dependencies; Stage III applied selective supervision and preservation techniques, achieving a modest overall score increase from 42.02 to 42.25 and the highest public Strict‑Small result as of 8 September 2026. The work also discusses further studies on compression, relational anchors, shared representations, and measurement, and makes models and code publicly available.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
6d ago

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...

By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko
arXiv AI
Sep 3

DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models

The paper introduces DKL, a method for adding new knowledge to instruction‑tuned language models without compromising their instruction‑following abilities. DKL performs extended pre‑training on a base LLM to embed knowledge, then merges these weights into the instruction‑tuned model, avoiding costly instruction fine‑tuning. Experiments show DKL raises RAG accuracy from 54.17% to 79.26% on retrieval failure cases while using far less training data than previous approaches.

By Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi, Gaurav Pandey, Sonam Gupta, Vineet Kumar, Jaydeep Sen, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu