arXiv AI By Augusto Camargo

The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models

Read the original on arXiv AI →

The paper argues that modern inference pipelines add an unseen layer of control between a language model’s frozen weights and its output, altering probability distributions before token selection. It introduces the concepts of the Inference Attribution Problem, Probability Placement, and Inference Policy Transparency to describe how such interventions can bias generated language toward specific frames and how these biases cannot be traced solely to model weights. The authors discuss the governance, security, and economic implications of these undisclosed inference policies, referencing EU AI Act, Digital Services Act, and FTC doctrines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 28

Safety from Honesty in a Disinterested AI Predictor

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified. We present a formal safety argument for the Scientist AI (SAI) Predictor, trained to approximate the Bayesian posterior conditioned on a dataset of "epistemically contextualized" natural-language statements.

arXiv AI
Jun 30

Safety from Honesty in a Disinterested AI Predictor

arXiv:2606. 29657v1 Announce Type: new Abstract: As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified.

By Yoshua Bengio, Oliver Richardson, Tom\'a\v{s} Gaven\v{c}iak, Michael Cohen, Rory Svarc, Damiano Fornasiere, Gael Gendron, David Hyland, Aton Kamanda, Adam Oberman, Francis Rhys Ward, Anna Gaven\v{c}iak, Jacob Livingston Slosser, Vincent Mai, Iulian Serban, Joumana Ghosn
arXiv AI
Aug 11

InfoOps Bench: A live information operations safety benchmark

arXiv:2607. 28503v3 Announce Type: replace Abstract: In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for use by authoritarian state "information operations": intentional, coordinated activities by one state to influence public opinion and information ecosystems in another state.

By Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright, John Gallacher
arXiv AI
4d ago

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

arXiv:2602. 02304v3 Announce Type: replace Abstract: Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context learning.

By Martino Ciaperoni, Marzio Di Vece, Roberto Pellungrini, Luca Pappalardo, Fosca Giannotti, Francesco Giannini