Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement
Read the original on arXiv Computation and Language →The paper reports a data‑efficient language modeling study conducted by Qiushi Engine on the BabyLM 2026 Strict‑Small benchmark, using only 10 million corpus words and 100 million cumulative presentations. It describes a three‑stage research program: Stage I built a frontier model via compact restatements and incremental learning; Stage II identified that exact repetition versus aligned restatement affect context use and proposed a principle for organizing experience around contextual dependencies; Stage III applied selective supervision and preservation techniques, achieving a modest overall score increase from 42.02 to 42.25 and the highest public Strict‑Small result as of 8 September 2026. The work also discusses further studies on compression, relational anchors, shared representations, and measurement, and makes models and code publicly available.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.