Nonlinear computation in deep linear networks
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
arXiv:2501. 02436v5 Announce Type: replace Abstract: Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning.
arXiv:2511. 02003v2 Announce Type: replace Abstract: We present the bulk--boundary decomposition as a new framework for understanding the training dynamics of deep neural networks.
arXiv:2602.08515v3 Announce Type: replace-cross Abstract: This work investigates shallow physics-informed neural networks (PINNs) for solving forward and inverse problems governed by nonlinear partia...
The paper demonstrates that a resistor‑diode network’s port behavior solves a ReLU monotone operator equilibrium network, effectively realizing a neural network in analog hardware. It introduces hardware linearization to compute gradients directly in the circuit, enabling in‑hardware training demonstrated via device‑level simulation. The study also extends to cascaded networks for feedforward architectures and shows how different nonlinear elements yield distinct activation functions, including a novel diode ReLU from a non‑ideal diode model.
arXiv:2501. 07400v2 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations.
arXiv:2402. 00152v5 Announce Type: replace Abstract: Constructing the architecture of a neural network is a challenging pursuit for the machine learning community, and the dilemma of whether to go deeper or wider remains a persistent question.