arXiv AI By Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels

Read the original on arXiv AI →

arXiv:2608. 03507v1 Announce Type: cross Abstract: Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 11

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models

CHRONOBERG is a temporally structured corpus of English book texts covering 250 years, curated from Project Gutenberg and enriched with temporal annotations. It enables quantification of lexical semantic change via time‑sensitive Valence‑Arousal‑Dominance analysis and the creation of historically calibrated affective lexicons. Experiments show that language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts, highlighting the need for temporally aware training and evaluation pipelines.

By Niharika Hegde, Subarnaduti Paul, Lars Joel-Frey, Manuel Brack, Kristian Kersting, Martin Mundt, Patrick Schramowski
arXiv Computation and Language
Sep 3

A Universal Vibe? Finding and Controlling Language-Agnostic Informal Register with SAEs

The study probes Gemma‑2‑9B‑IT with Sparse Autoencoders across English, Hebrew, and Russian to examine how multilingual LLMs handle informal register. By using a dataset of polysemous terms that appear in literal and informal contexts, the authors isolate pragmatic register processing from lexical cues. They discover a small, robust cross‑linguistic core that forms an informal register subspace, which becomes clearer in deeper layers and can causally shift output formality across all tested languages, even transferring zero‑shot to six unseen languages.

By Uri Z. Kialy, Avi Shtarkberg, Ayal Klein
arXiv AI
Sep 2

Lingua Franca or Probing Artifact? Rethinking Latent Language in Multilingual LLMs

The paper investigates whether different latent language probes—GMM-based representation probes and decoding-based probes—measure the same phenomenon in multilingual language models. Across various model families, training regimes, domains, tasks, checkpoints, and up to 27 languages, the authors find systematic disagreement: representation probes indicate earlier cross‑lingual mixing, while decoding probes reveal sharper, English‑biased language signals. These differences correlate with model multilinguality and training progression but remain relatively stable across domains, suggesting that current probes capture distinct aspects of multilingual processing rather than a single internal lingua franca.

By Deniz Bayazit, Badr AlKhamissi, Antoine Bosselut
arXiv Computation and Language
Aug 28

Double Trouble: Bilingual Pretraining Leaves Language-Conditioned Effects in Shared-Language Representations

The study compares an English-only and a bilingual decoder-only model, each 310 M parameters, trained on eight diverse languages while controlling for English exposure, compute, and document overlap. After aligning on shared English vocabulary, the authors find that token embeddings appear similar, but the deeper hidden states used for prediction diverge across models. This hidden‑state mismatch grows through middle transformer layers and persists despite controls, indicating that contextual processing differs between the models. "whyItMatters":"The findings show that embedding alignment can conceal significant internal representation differences, which is crucial for any downstream work that assumes aligned multilingual models are interchangeable."

By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos