Lock-in EP: An In-Situ Training Algorithm for Oscillatory Hardware
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates how temperature affects analog deep neural network (DNN) inference, focusing on both stochastic and systematic non‑idealities in analog hardware. Experiments show that temperature‑induced performance loss is mainly driven by systematic errors rather than random noise. The study evaluates various mitigation techniques, finding that noise‑aware training and temperature‑aware calibration—especially hardware‑in‑the‑loop training—best preserve inference accuracy across different thermal conditions.
The paper introduces a communication‑efficient method for adapting large language models on decentralized GPU meshes. It proposes an asynchronous two‑circuit system that uses fast compressed training with activation masking for pipeline‑parallel transfer and compressed data‑parallel synchronization, while a slower anchor circuit performs occasional unmasked passes. A spectral correction optimizer then denoises the masked gradients using these anchor priors, enabling high compression rates and achieving up to 40× throughput gains over internet‑grade connections while matching dense uncompressed performance.
arXiv:2505. 20137v5 Announce Type: replace-cross Abstract: Predictive Coding (PC) offers a brain-inspired alternative to backpropagation for neural network training, described as a physical system minimizing its internal energy.
arXiv:2605. 11855v2 Announce Type: replace-cross Abstract: Sequence learning is dominated by Transformers and parallelizable recurrent neural networks (RNNs) such as state-space models, yet learning long-term dependencies remains challenging, and state-of-the-art designs trade power consumption for performance.
arXiv:2608. 01997v1 Announce Type: new Abstract: Single-optimizer training is a poor fit for the distinct phases of deep network optimization: adaptive methods handle noisy early gradients well but overshoot flat minima, while SGD with momentum generalizes better in the late phase but converges slowly early on.
arXiv:2609.36584v1 Announce Type: new Abstract: Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, s...