arXiv Machine Learning By Burc Gokden

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

Read the original on arXiv Machine Learning →

arXiv:2608. 10288v1 Announce Type: new Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bilinear form of scaled dot-product attention (SDPA) with a learned, input-generated bilinear operator $G_{LM}$, built from a positive tensor $A_{LM}$ by elementwise power laws.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 21

Large Language Models As Shannon Lossy Compressors Not Solomonoff Induction Estimators: The Singularity Is Not Near Without Symbolic Model Synthesis

The paper argues that Large Language Models (LLMs) do not function as Solomonoff induction estimators because their training objectives—cross‑entropy, negative log‑likelihood, and next‑token prediction—optimize fit to a supplied conditional distribution rather than a program‑weighted universal mixture. It further contends that additional computation alone does not transform these models into optimal predictors without external hyper‑parameter or architectural changes. The authors suggest that neurosymbolic machine learning, exemplified by models such as Fable and Astra, represents a shift toward symbolic model synthesis, moving beyond purely statistical LLMs.

By Hector Zenil, Abicumaran Uthamacumaran, Luan Ozelim
arXiv AI
Aug 24

From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics

The paper introduces OAttention, a token‑level attention mechanism that assigns each token a presence coefficient based on its hidden representation. This coefficient both gates the token’s output and weights its contribution to other tokens, making zero‑vector tokens behave as true zeros and enabling exact null‑receiver, null‑source, and empty‑support properties. The authors extend this idea to local O‑components and an O‑Transformer, and demonstrate small performance changes when retrofitting a pretrained TabPFN model.

By Heyang Gong
arXiv Machine Learning
Sep 4

Coupled Scaling: A Representational Accessibility Framework for Neural Scaling Laws

The paper introduces Coupled Scaling, a framework that links neural scaling laws to the relationship between task structure and the geometry that an architecture‑optimization system can access. It shows that finite‑budget scaling depends on how well the system’s representational support aligns with the task’s energy distribution, deriving residual exponents that vary with architectural coverage and tail decay. The authors propose tests to verify whether static task‑relevant geometry tracks loss and whether multiscale geometry follows coupling‑specific exponent ordering, suggesting a factorial audit of emergence trajectories to isolate geometry from scaling fits.

By Jie Wang