arXiv:2605. 11644v2 Announce Type: replace-cross Abstract: We study positive-data learning of languages admitting reduced working binary linear nondeleting multiple context-free grammar presentations of bounded fan-out.
By Takayuki Kuriyama
arXiv:2605. 11644v3 Announce Type: replace-cross Abstract: Positive data can show that two tuple occurrences share a successful sentence context without certifying that they are safely interchangeable.
By Takayuki Kuriyama
arXiv:2607. 29496v1 Announce Type: new Abstract: We study transcript management for fixed, finite-precision causal Transformers.
By Sergey Salishev
arXiv:2607. 26988v1 Announce Type: cross Abstract: What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length?
By Franz Nowak, Ryan Cotterell, Reda Boumasmoud
arXiv:2607. 15107v1 Announce Type: new Abstract: This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting.
By Sridhar Mahadevan
arXiv:2608. 13433v1 Announce Type: cross Abstract: Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise characterization of which tasks admit length generalization.
By Andy Yang, Blerta Veseli, Corentin Barloy, Micha\"el Cadilhac, Andreas Krebs, Charles Paperman, Howard Straubing, Michael Hahn
arXiv:2603. 25414v4 Announce Type: replace-cross Abstract: A prevailing assumption in machine learning is that model correctness must be enforced after the fact.
By Houston Haynes
arXiv:2607. 11897v1 Announce Type: new Abstract: Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory.
By Tiantian Zhang
arXiv:2608. 10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws.
By Burc Gokden
arXiv:2607. 20586v1 Announce Type: new Abstract: We study vector-valued affine refinement operators of the form [ (W\gamma)(t)=\sum_{j\in\mathbb{Z}} A_j\gamma(Mt-j)+B(t), ] with finitely supported matrix mask and compactly supported continuous piecewise linear input and forcing data.
By Boldsaikhan Bolorkhuu, Tsogtgerel Gantumur
arXiv:2601.05280v4 Announce Type: replace-cross
Abstract: On the one hand, the question of whether Large Language Models (LLMs) are Solomonoff induction estimators has become an explicit question at...
By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim
arXiv:2606. 16541v1 Announce Type: new Abstract: Autoformalization, translating natural-language mathematics into formal proof assistants, is bottlenecked not by translation fluency but by \emph{faithfulness}: a formal statement can typecheck and be provable, yet still encode a different theorem than the source intended.
By Noor Islam S. Mohammad, Tamim Sheikh