arXiv:2605. 11644v3 Announce Type: replace-cross Abstract: Positive data can show that two tuple occurrences share a successful sentence context without certifying that they are safely interchangeable.
By Takayuki Kuriyama
arXiv:2605. 21379v3 Announce Type: replace-cross Abstract: In The Algebraic Mind (2001), Marcus held that any adequate cognitive architecture needs operations over variables, recursively structured representations, and an individual/kind distinction, and that multilayer perceptrons support none of them; he left a register-and-treelet implementation as a conjecture.
By Hiroyuki Chuma, Kanji Otsuk, Yoichi Sato
arXiv:2608. 14004v1 Announce Type: new Abstract: In-context learning is commonly formalized as inference from examples of a function.
By Faizanuddin Ansari, Debanjan Dutta, Swagatam Das
arXiv:2606. 24418v1 Announce Type: new Abstract: Data augmentation is a simple and model-agnostic approach for exploiting known invariances in learning problems.
By Behrooz Tahmasebi, Melanie Weber, Stefanie Jegelka
arXiv:2605. 11644v2 Announce Type: replace-cross Abstract: We study positive-data learning of languages admitting reduced working binary linear nondeleting multiple context-free grammar presentations of bounded fan-out.
By Takayuki Kuriyama
arXiv:2607. 15107v1 Announce Type: new Abstract: This paper develops a categorical framework -- Learning in Infinitesimal Non-Compositional Sketches (LINCS) -- as the repair of non-compositionality: failures of diagrams to factor through quotient sketches lifted to the tangent category setting.
By Sridhar Mahadevan
arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.
By Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet
Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately. What happens at the extreme of this spectrum, when the architecture's expressible function class collapses to a finite-dimensional algebraic variety?
arXiv:2609. 04046v1 Announce Type: cross Abstract: What can a single layer of self-attention compute?
By Rajmohan Rajaraman, Ravi Sundaram, Amanuel Tesfaye
arXiv:2409. 15600v3 Announce Type: replace Abstract: A representation of a molecule or material should be invariant to the symmetries of physics, unique, continuous, efficient and general.
By Rahul Khorana, Marcus Noack, Jin Qian
arXiv:2608. 11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.
By Zijian Zhao, Sen Li
arXiv:2609.08961v1 Announce Type: cross
Abstract: For a finite set $O$ of Boolean functions, we consider the class of propositional formulas built using the functions in $O$ as connectives. We determ...
By Balder ten Cate