arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2609.37824v1 Announce Type: cross
Abstract: Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outco...
By Timur Mudarisov, Mikhail Burtsev, Radu State
arXiv:2607. 18348v1 Announce Type: cross Abstract: We propose a transition-centred geometric analysis of transformer residual streams.
By Sunit Bhattacharya, Ravi Shankar Kolli
arXiv:2609. 12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities.
By Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte, Holger Giese
arXiv:2607. 06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized.
By Truong Xuan Khanh
The paper investigates how many transformer components influence a token prediction by measuring the absolute contribution of each unit and channel to the logit. It finds that thousands of components contribute to a single prediction, yet a small subset—often just dozens—carries the majority of the predictive mass. Across models ranging from 124 M to 7 B parameters, the proportion of the model involved in a prediction remains around one to three percent, independent of size, and the study demonstrates that specific components can be directly read and written to modify model behavior without additional training.
By Mark Oskin