arXiv:2609.15439v1 Announce Type: cross
Abstract: Generative thermodynamic computers turn thermal noise into structured data through Langevin dynamics. We train these systems with a local update at e...
By Huilin Wang, Weibing Deng
arXiv:2609.37043v1 Announce Type: new
Abstract: Sampling from unnormalized distributions over large discrete state spaces becomes difficult when a multimodal target is far from a tractable reference....
By Yuwen Qian, Yidong Ouyang, Zhengyan Wan, Hongyuan Zha
arXiv:2605. 29547v2 Announce Type: replace-cross Abstract: Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern architectures due to non-smooth components such as ReLU activations and quantization operators.
By Ruoran Xu, Borong She, Xiaobo Jin, Qiufeng Wang
arXiv:2608. 05025v2 Announce Type: replace Abstract: Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection.
By Dmytro Knopov
Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers. We study short-h...
arXiv:2606. 17120v1 Announce Type: new Abstract: Deep neural networks (DNNs) exhibit first order phase transitions under variations of the L2 regularization strength, with each transition marking the onset of a new learnable feature.
By Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.
By Irina Piontkovskaia, Sergey Nikolenko
arXiv:2608. 15483v1 Announce Type: new Abstract: Modern deep networks are trained through long update trajectories, yet their temporal organization remains less systematically characterized than architectures, losses, or optimizers.
By Fanqi Wang, Weisheng Tang, Hairong Qi
arXiv:2609.08618v1 Announce Type: new
Abstract: Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this mis...
By Zhongxuan Liu, Sicheng Zhou, Hongzhi Wang
arXiv:2608. 15373v1 Announce Type: new Abstract: Inverse physics-informed neural networks (PINNs) can reconstruct a field accurately while returning an incorrect physical parameter.
By Yifan Zhang, Qian Tao
arXiv:2608. 05025v1 Announce Type: new Abstract: Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection.
By Dmytro Knopov
arXiv:2606. 17572v1 Announce Type: new Abstract: Learned dynamics models often answer global physical questions, such as fault severity or impact stiffness, by pooling a per-step feature sequence into one readout vector.
By Yifan Wang