arXiv:2607. 14144v1 Announce Type: new Abstract: The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality.
By Wenhui Chen, Jianlin Chen, Ziyao Lin, Chi Man Vong
arXiv:2608. 01548v2 Announce Type: replace Abstract: Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--environment complex can execute, and which possibilities can be realized with finite resources.
By Yi Liu
arXiv:2607. 18305v1 Announce Type: cross Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text.
By Priyansh Srivastava, Romit Chatterjee
arXiv:2605. 25889v4 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models reach high success rates on clean inputs but collapse under small adversarial perturbations: a $16/255$ PGD attack drops OpenVLA-7B's LIBERO success from $95\%$ to under $5\%$.
By Jianwei Tai
arXiv:2608. 05085v1 Announce Type: cross Abstract: Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, and when to stop.
By Ahmed Hassoon, Mark Dredze
arXiv:2608. 01548v1 Announce Type: cross Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition.
By Yi Liu
The paper introduces Coupled Scaling, a framework that links neural scaling laws to the relationship between task structure and the geometry that an architecture‑optimization system can access. It shows that finite‑budget scaling depends on how well the system’s representational support aligns with the task’s energy distribution, deriving residual exponents that vary with architectural coverage and tail decay. The authors propose tests to verify whether static task‑relevant geometry tracks loss and whether multiscale geometry follows coupling‑specific exponent ordering, suggesting a factorial audit of emergence trajectories to isolate geometry from scaling fits.
By Jie Wang
arXiv:2609.21523v1 Announce Type: new
Abstract: A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by lim...
By Ronald Katende
arXiv:2608.22347v1 Announce Type: new
Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal b...
By Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fern\'andez
arXiv:2607. 22361v1 Announce Type: new Abstract: We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive.
By Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately. What happens at the extreme of this spectrum, when the architecture's expressible function class collapses to a finite-dimensional algebraic variety?
arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.
By Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet