arXiv AI

Pretraining Language Models on Historical Text

arXiv:2606. 02991v1 Announce Type: cross Abstract: We introduce TypewriterLM, a 7.

arXiv AI
Jul 15

Scaling Point-in-Time Language Models

arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.

By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv Machine Learning
Aug 27

Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training

arXiv:2608. 25826v1 Announce Type: cross Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched.

By Qiankai Xu, Qiguang Chen, Zixin Su, Wenhao Huang, Yue Gao, Jiaheng Liu, Ge Zhang
arXiv AI
3d ago

Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning

The paper introduces the problem of cross‑lingual loopholes in large language model (LLM) unlearning, where forgetting a fact in one language can leave it accessible in others. It presents a new 174‑language benchmark, the Cross‑Lingual Unlearning Tensor, and proposes COVER, a method that selects a subset of source languages to maximize unlearning coverage under a language budget. Experiments show COVER reduces residual knowledge by 7.8–27.3% compared to uniform selection and works on both synthetic and real low‑resource news data.

By Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav
arXiv AI
Aug 26

Relative Time Intervals Representation for Word-level Timestamping with Masked Training

The paper introduces a method for improving fine-grained, temporally aligned outputs in Speech Large Language Models (SpeechLLMs) by replacing absolute timestamps with relative timestamps, which reduces vocabulary size and enhances generalization. It proposes a hybrid fine‑tuning strategy that fully fine‑tunes the timestamp‑augmented embedding layer and language model head while applying LoRA to decoder layers, and introduces a masked timestamp training objective to prevent over‑reliance on ground‑truth timestamps. Experiments show significant gains in timestamp prediction accuracy without compromising transcription quality.

By Quanwei Tang, Zhiyu Tang, Xu Li, Dong Zhang, Shoushan, Guodong Zhou