arXiv Machine Learning

The Attribution Contract: Feature Attribution for Generative Language Models

arXiv:2605. 23080v2 Announce Type: replace Abstract: Feature attribution methods promise to identify which input features matter for a model output.

arXiv Machine Learning
Aug 28

The Attribution Contract for Generative Language Models

The paper argues that feature attribution scores for generative language models lack a fixed meaning because each generated token is both output and input, leading to multiple distinct explanatory questions. It introduces the Attribution Contract framework, which explicitly defines the model score, fixed variables, target output, generation process, and eligible features, showing how these choices affect attribution outcomes. Experiments demonstrate that different contracts (e.g., local next-token vs. prompt-level) and model architectures (mixture-of-experts vs. masked-diffusion) yield markedly different attribution distributions, highlighting the need for careful contract specification.

By Giang Nguyen
arXiv Machine Learning
Jul 7

How Much is Left? LLMs Linearly Encode Their Remaining Output Length

arXiv:2607. 05316v1 Announce Type: cross Abstract: Large language models generate one token at a time, yet their responses show remarkably consistent length structure: step-by-step solutions converge in predictable token counts, retrievals stop after a few sentences, retractions extend responses by measurable amounts.

By Mohamed Amine Merzouk, Dmitri Carpov, Mirko Bronzi, Damiano Fornasiere, Adam Oberman
arXiv AI
2d ago

Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models

The paper introduces Diffusion Layer Integrated Gradients (DLIG), a token attribution technique for diffusion language models that extends Integrated Gradients to any layer and denoising step. DLIG tracks how a model progressively commits to a self-generated or fixed completion for a prompt, and it satisfies the IG axioms of completeness, implementation invariance, linearity, and symmetry preservation. The method is lightweight and complements interventional analysis, and the authors apply it to tasks such as word-sense disambiguation, multi-hop graph reasoning, and sentence infilling to show how diffusion language models utilize inputs across positions, layers, and denoising steps.

By Darpan Aswal, C\'eline Hudelot