arXiv AI

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale

arXiv:2607. 14144v2 Announce Type: replace Abstract: The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of reality.

arXiv Machine Learning
Sep 4

Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws

The paper introduces Coupled Scaling, a framework that links neural scaling laws to the relationship between task structure and the geometry that an architecture‑optimization system can access. It shows that finite‑budget scaling depends on how well the system’s representational support aligns with the task’s energy distribution, deriving residual exponents that vary with architectural coverage and tail decay. The authors propose tests to verify whether static task‑relevant geometry tracks loss and whether multiscale geometry follows coupling‑specific exponent ordering, suggesting a factorial audit of emergence trajectories to isolate geometry from scaling fits.

By Jie Wang
arXiv Machine Learning
Jul 27

Indexing: the Beginning and the End

arXiv:2607. 22361v1 Announce Type: new Abstract: We study information bottlenecks in modern deep-learning architectures -- RNNs, softmax transformers, linear-attention transformers and state-space models -- through the lens of the indexing primitive.

By Alexander Kozachinskiy, Vicente Opazo, Felipe Urrutia
Hugging Face Trending Papers
Jul 15

Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately. What happens at the extreme of this spectrum, when the architecture's expressible function class collapses to a finite-dimensional algebraic variety?

arXiv Machine Learning
Jul 16

Algebraic Representability as the Limiting Regime of Grokking: An Exactly Solvable Model with Holomorphic Activations

arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.

By Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet