A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicly funded 9B-parameter model for European Portuguese, appears competitive on agreement alone: asked to code the moral foundation of authority, it agrees with trained human coders to within six F1 points of open models eight to thirteen times its size.
arXiv:2607. 08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value.
By Manuel Pita
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.
arXiv:2607. 14399v1 Announce Type: new Abstract: Evaluations of language-model honesty read the model's verdicts as evidence about the model.
By Justin Bronder (Corabo Inc.)
The paper introduces the concept of "linguistic illegibility," describing how a large language model’s (LLM) language outputs and extracted linguistic features may not accurately reflect its internal computations. It argues that because LLMs compute primarily in activation spaces, any reliance on linguistic self‑reporting for security—such as chain‑of‑thought monitoring or constitutional self‑critique—cannot be fully reliable. The authors propose taint tracking and other sandboxing techniques that do not depend on the model’s linguistic state as a more robust security foundation.
By James Mickens
arXiv:2609.39001v1 Announce Type: cross
Abstract: LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across lang...
By Thomas Serval
arXiv:2607. 05031v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to produce test oracles, the part of a test that decides whether observed behavior is correct.
By Ali Hassaan Mughal, Muhammad Bilal
The paper argues that as AI systems increasingly generate code, the bottleneck has shifted to supervising these systems, revealing a vocabulary gap between cybernetic coordination (actions aligning with the world) and epistemic coordination (understanding that can be verified). It critiques current oversight that merely approves outputs, proposing instead that every consequential choice by an agent must include a retrievable condition explaining why it was made, enabling third‑party verification. The authors illustrate this with three delegation episodes, introduce a two‑part reconstruction test, and propose the ORRCF convention to embed such conditions in all recorded decisions.
By J\'er\'emie Lumbroso
arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.
By Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
arXiv:2609.14754v1 Announce Type: cross
Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
By Orion Reblitz-Richardson
arXiv:2608. 09028v1 Announce Type: new Abstract: Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints.
By Ponkrit Kaewsawee, Chaklam Silpasuwanchai, Chutiporn Anutariya
The paper argues that modern inference pipelines add an unseen layer of control between a language model’s frozen weights and its output, altering probability distributions before token selection. It introduces the concepts of the Inference Attribution Problem, Probability Placement, and Inference Policy Transparency to describe how such interventions can bias generated language toward specific frames and how these biases cannot be traced solely to model weights. The authors discuss the governance, security, and economic implications of these undisclosed inference policies, referencing EU AI Act, Digital Services Act, and FTC doctrines.
By Augusto Camargo