The paper introduces recurrent Graph Neural Networks (GNNs) that use set-based aggregation and establishes conditions that can be verified directly from the network weights. It proves a two‑directional equivalence between these networks and the Boolean closure of reachability and safety properties, corresponding to the fragment BΣ◦₁ of the modal μ‑calculus. This equivalence allows for verifiable symbolic explanations of networks that satisfy the identified conditions, without relying on counting logic or external halting signals.
By Blai Bonet
arXiv:2606. 20325v1 Announce Type: new Abstract: Classical approximation theorems ask for a new neural network whenever the target accuracy is improved.
By Valentin Abadie, Clemens Hutter, Helmut B\"olcskei
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers. In language modeling, \textbf{transformers} have emerged as the dominant architecture, with early layers capturing local syntactic patterns and later layers encoding more complex clause-level dependencies.
arXiv:2511. 14953v2 Announce Type: replace-cross Abstract: Discrete structures are currently second-class in differentiable programming.
By Joey Velez-Ginorio, Nada Amin, Konrad Kording, Steve Zdancewic
arXiv:2603. 03612v3 Announce Type: replace Abstract: The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability.
By William Merrill, Hongjian Jiang, Yanhong Li, Anthony Lin, Ashish Sabharwal
arXiv:2606. 17522v1 Announce Type: cross Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers.
By Vinoth Nandakumar, Qiang Qu, Pramod Thebe, Sakshi Khachariya, Tongliang Liu
arXiv:2603. 05573v2 Announce Type: replace Abstract: Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallelism, which enables efficient training.
By Gyuryang Heo, Timothy Ngotiaoco, Kazuki Irie, Samuel J. Gershman, Bernardo L. Sabatini
arXiv:2606. 03645v1 Announce Type: cross Abstract: Large Language Models exhibit paradoxical fragility in fundamental arithmetic, implying a disconnect between internal computation and discrete output.
By Liuyuan Wen, Xun Zhu, Lihao Huang, Wenbin Li, Yang Gao
arXiv:2606. 01372v1 Announce Type: cross Abstract: Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns?
By Divyansh Jha, Yuanfang Xie, Varan Mehra, Brennen Yu
arXiv:2607. 26988v1 Announce Type: cross Abstract: What types of decision problems can a causally masked, finite-precision transformer solve for inputs of arbitrary length?
By Franz Nowak, Ryan Cotterell, Reda Boumasmoud
The paper revisits the debate between Pinker & Prince (1988) and Rumelhart & McClelland (1986) regarding neural network models of English past tense. It argues that modern Encoder-Decoder architectures in NLP address the empirical shortcomings highlighted by Pinker and Prince, eliminating the need to simplify the past tense mapping problem. The authors suggest that the strong performance of these contemporary models merits a reassessment of their role in linguistic and cognitive modeling.
By Christo Kirov, Ryan Cotterell