When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2605.06240v2 Announce Type: replace-cross Abstract: Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inheri...
arXiv:2605.11608v2 Announce Type: replace-cross Abstract: A single base LLM now comes with dozens of post-training variants, quantized, LoRA-adapted, or distilled, and each has to be checked before r...
arXiv:2605. 05209v2 Announce Type: replace-cross Abstract: Flat minima are an account of why deep networks generalise.
Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quantitative mechanism underlying its implicit bias toward flat minima remains unclear. In particular, the perturbation radius $ρ$ is typically treated as an isolated tuning parameter, despite defining the neighborhood in which SAM measures sharpness.
arXiv:2606. 28654v1 Announce Type: cross Abstract: Deep Neural Network (DNN) classifiers suffer from poor calibration when their softmax outputs (predictive confidence) deviate from the empirical likelihoods.
arXiv:2604. 21395v3 Announce Type: replace-cross Abstract: Ordinary supervised training minimises the task loss and then stops.