arXiv AI By Divyansh Jha, Yuanfang Xie, Varan Mehra, Brennen Yu

BRo-JEPA: Learning Modular Arithmetic in Latent Space

Read the original on arXiv AI →

arXiv:2606. 01372v1 Announce Type: cross Abstract: Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

BRo-JEPA: Learning Modular Transformations in Latent Space

The paper introduces BRo-JEPA, a world model that learns modular arithmetic operations as rotations in latent space. Using MNIST and EMNIST datasets, BRo-JEPA achieves near-perfect zero‑shot generalization to unseen operations, outperforming standard supervised and JEPA baselines by a large margin. The model demonstrates that latent transformations can encode underlying algebraic structures, enabling strict zero‑shot operation generalization.

By Divyansh Jha, Yuanfang Xie, Brennen Yu, Varan Mehra
arXiv AI
Sep 21

Understanding In-context Learning of Addition via Activation Subspaces

The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.

By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen
arXiv AI
4d ago

In-Context Learning Amplifies a Latent Symbolic Circuit

The paper investigates how large language models activate a latent symbolic reasoning circuit—comprising abstraction, induction, and retrieval—when presented with in-context examples. By tracking this circuit across different shot counts and model families, the authors show that its components become detectable and functional long before the model reaches high accuracy. They further demonstrate that per-head causal contributions can increase eightfold from 1- to 10-shot, and that interventions such as cross-shot activation patching or function vector injection can dramatically improve accuracy, even at 0-shot, by leveraging the pre‑existing circuit in the model weights.

By Melissa Wessel
arXiv Machine Learning
Jun 11

Composing Linear Layers from Irreducibles

arXiv:2507. 11688v4 Announce Type: replace Abstract: Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood.

By Travis Pence, Daisuke Yamada, Vikas Singh