arXiv Machine Learning By Zetai Cen, Jin Zhu, Xinwei Shen, Chengchun Shi

Perturbation is All You Need for Extrapolating Language Models

Read the original on arXiv Machine Learning →

arXiv:2605. 04344v2 Announce Type: replace-cross Abstract: This paper develops a statistical theory of extrapolation for large language models, by reinterpreting them through pre-post-additive noise models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Structural priors for data-efficient language learning

The paper explores structural transfer, where models are first trained on non-language data such as music, probabilistic grammars, and cellular automata to induce priors for natural language tasks. This pretraining acts as a weight initialization for multilingual language modeling and leads to lower next-token prediction loss and smaller weight shifts during subsequent language training. However, the improved loss does not consistently translate into better downstream linguistic performance, and the efficiency of non-language data is lower than that of additional language data.

By Yana Veitsman, Jonas Mayer Martins, Jonathan Lautenschlager, Lisa Beinborn
arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post
arXiv AI
Jun 16

Data Augmentations for Data-Constrained Language Model Pretraining

arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.

By Michael K. Chen, Xikun Zhang, Zhen Wang
arXiv Computation and Language
Sep 1

Manac\'a-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation

Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.

By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
arXiv AI
6d ago

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

The paper introduces TokenAdapt, a model‑agnostic tokenizer transplantation method that uses a hybrid heuristic to initialize new token embeddings, and a novel pre‑tokenization learning approach for multi‑word Supertokens to improve compression. TokenAdapt combines local subword decomposition and global semantic similarity to preserve semantics while reducing retraining needs. Empirical results show that TokenAdapt outperforms existing baselines such as Transtokenizer and ReTok, achieving lower perplexity ratios and significant compression gains.

By Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath