arXiv:2607. 00479v1 Announce Type: new Abstract: Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning.
By Peilin Liu, Ding-Xuan Zhou
arXiv:2608.22368v1 Announce Type: new
Abstract: While linear attention is a compelling mechanism for high-resolution object detection due to its reduced cost for global token mixing, converting the S...
By Huaiyuan Qin, Gabriel James Goenawan, Zihang Lin, Muli Yang, Hongyuan Zhu
arXiv:2606. 27748v1 Announce Type: cross Abstract: Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences.
By Haoran Zhang, Feng Zhou
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
arXiv:2606. 06458v1 Announce Type: new Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied in fields ranging from computational pathology to satellite imagery.
By Alexander M\"ollers, Marvin Sextro, Julius Hense, Gabriel Dernbach, Klaus-Robert M\"uller
The paper introduces Neural Low-Degree Filtering (Neural LoFi), a stylized limit of gradient-based training that turns hierarchical feature learning into an explicit iterative spectral procedure. In this framework, each layer independently selects directions with maximal low-degree correlation to the label, providing a tractable surrogate for deep learning and a kernel-space interpretation. Experiments on fully connected and convolutional networks show that Neural LoFi outperforms lazy random-feature baselines, recovers meaningful structured filters, and aligns with early gradient-descent feature discovery on real datasets.
By Yatin Dandi, Matteo Vilucchio, Luca Arnaboldi, Hugo Tabanelli, Florent Krzakala
arXiv:2607. 11875v1 Announce Type: cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models.
By Tiberiu Musat, Tiago Pimentel, Nicholas Zucchet, Thomas Hofmann
arXiv:2510. 25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits.
By Rabin Adhikari
arXiv:2509. 05130v2 Announce Type: replace Abstract: In classification problems, models are trained to predict a class label based on the input data features.
By Davide Pirovano, Federico Milanesio, Michele Caselle, Piero Fariselli, Matteo Osella
The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.
By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
arXiv:2410. 24050v3 Announce Type: replace Abstract: Large-scale pretraining of transformers has been central to the success of foundation models.
By Ambroise Odonnat, Wassim Bouaziz, Vivien Cabannes
The paper introduces Cubit, a Transformer‑style architecture that replaces the standard attention mechanism with Kernel Ridge Regression (KRR). By interpreting attention as Nadaraya‑Watson regression, Cubit incorporates the closed‑form KRR solution, combining kernel‑based value aggregation with normalization via the inverse kernel matrix. The authors also propose a Limited‑Range Rescale (LRR) to stabilize training and report that Cubit shows improved long‑sequence modeling, with gains increasing as training sequence length grows.
By Chuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang, Liangchen Tan, Mac Schwager, Anderson Schneider, Yuriy Nevmyvaka, Xiaodong Liu