arXiv:2607. 04926v1 Announce Type: cross Abstract: How does the way information reaches a transformer -- as symbolic tokens, a clean per-factor "oracle" code, or an entangled perceptual vector -- shape whether it binds that information compositionally?
By Yoshiyuki Ootani
arXiv:2609.39445v1 Announce Type: cross
Abstract: Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup...
By Hung Phan, Thuy T. Nguyen, Minh Ngoc Dinh, Nhat-Quang Tran
The paper investigates test‑time adaptation for medical image segmentation, showing that a fixed adaptation horizon can harm many individual cases. It introduces prediction fragmentation—a measure of disagreement between the source model and the adapted mask—to predict harmful adaptation without extra labels or backward passes. Using a case‑level router based on this metric, the authors reduce harmful adaptation on cardiac MRI from 58.7% to 20% while maintaining accuracy.
arXiv:2608. 01575v1 Announce Type: new Abstract: Whether large language models perform genuine algorithmic reasoning or mere pattern completion is hard to test, because most benchmarks lack a ground truth for correct inductive inference.
By Hector Zenil, Luan Ozelim
arXiv:2608.20647v1 Announce Type: new
Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (...
By Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
arXiv:2608.21098v1 Announce Type: new
Abstract: Fusing prior knowledge with data-driven learning is attractive where data is scarce, yet no controlled account says when it helps, is redundant, or har...
By Ahmad AlMughrabi, Albert Clop, Benjamin Busam, Ricardo Marques, Petia Radeva
arXiv:2607. 24797v1 Announce Type: cross Abstract: In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and a fronto-parietal encoding route (impaired in pure agraphia), sharing a partial orthographic core.
By Diego Salda\~na Ulloa
The paper investigates why differentiable causal discovery methods that encode expert priors as forbidden-edge constraints via an Augmented Lagrangian (ALM) penalty—termed the "guide, not bind" approach—often fail. It identifies two key failures: (1) the sequential penalty‑ramping ALM suppresses a true edge before counterfactual checks can detect it, and the proposed adaptive relaxation rule DADU violates necessary conditions for safe relaxation, leading to a high failure rate across thousands of training runs; (2) the standard correlation‑matching objective inherently ties a true edge and its reverse to the same cost, whereas covariance matching can separate them by a provable margin. The authors provide theoretical propositions, corollaries, and empirical evidence to support these claims.
By Sairam Sundararaman, Sara Girdhar, Manit Narasimha Murthy, Samrudh N, Bhaskarjyoti Das
arXiv:2605. 04893v2 Announce Type: replace Abstract: When a language model processes a hallucinated response, its attention routing tends to fail in one of two shapes: over-concentrating on a narrow set of positions, or spreading so diffusely that relevance is diluted, and the shape of the failure carries diagnostic signal.
By Dominik Dahlem, Diego Maniloff, Mac Misiura
The paper introduces a method for deciding whether to adapt a frozen segmentation model at test time, arguing that a fixed adaptation horizon conflates two distinct decisions: how far to adapt and whether to adapt at all. By measuring disagreement geometry—called prediction fragmentation—between the source model and the adapted mask, the authors predict harmful accepted area (HA) without extra labels or backward passes, achieving strong correlation across three medical benchmarks. A case‑level router built on this metric reduces HA significantly while maintaining or improving Dice scores, and the approach generalizes across architectures and domains.
By Lili Wang, Jing Li, Xiaowen Sun, Xiangyu Hu, Zhuangzhuang Gu, Jian Liu, Srihari Nelakuditi, Yan Tong
arXiv:2608. 10441v1 Announce Type: new Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using.
By Ying Yuan
arXiv:2607. 12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior.
By Gunner Levi Howe