arXiv Machine Learning By R\'emi Bourgerie, \v{S}ar\={u}nas Girdzijauskas, Viktoria Fodor

Fixed Points Without Fixed Diffusion: Implicit Neural Sheaves for Convergent Test-Time Computation

Read the original on arXiv Machine Learning →

The paper introduces SheafDEQ, a subhomogeneous deep-equilibrium architecture that uses adaptive neural-sheaf propagation to allow richer, edge-dependent transformations in implicit graph neural networks while guaranteeing a unique equilibrium. The authors prove that SheafDEQ’s equilibrium is globally reachable from any positive initialization and remains contractive even with bounded communication staleness. Experiments demonstrate that SheafDEQ outperforms fixed-propagation implicit baselines on tasks such as Sums, MNIST Terrain, Coordinates, and community detection, especially as graph connectivity becomes increasingly heterophilic.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 13

The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op.

arXiv Machine Learning
Sep 14

Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models

The paper introduces a feedforward graph architecture that uses several frozen large language models as computational nodes connected through a shared continuous latent space via learned linear projections. By jointly optimizing projection matrices through backpropagation, the system combines the representations of three small frozen models with two larger ones, culminating in a lightweight cross‑attention output node. With only 17.6 M trainable parameters, the architecture attains state‑of‑the‑art results on ARC‑Challenge, OpenBookQA, and MMLU, surpassing both individual constituent models and parameter‑matched learned classifiers.

By Marcus Armstrong, Navid Ayoobi, Arjun Mukherjee