arXiv Machine Learning

Dynamic Parameterization Is Not Dynamic Inference

arXiv:2607. 26192v1 Announce Type: new Abstract: Input-dependent controller coefficients are often treated as evidence of dynamic inference or computational savings.

arXiv Machine Learning
5d ago

Deterministic Regime Switching and Feasibility Inversion in Dynamic Tensor Rematerialization

The paper reports fine‑grained, deterministic instability in Dynamic Tensor Rematerialization (DTR), an online eviction policy for memory‑constrained DNN training. On an LSTM trace, tiny changes in memory budget (0.10% of peak) switch the system between fast and slow execution regimes with up to 7.3× overhead differences, driven by repeated re‑eviction of the same storages. On a ResNet‑32 trace, a deterministic feasibility inversion is observed: the run is feasible at a 0.101 budget ratio, infeasible (OOM) between 0.102–0.106, and feasible again from 0.107, caused by a fully pinned recursive rematerialization frontier exceeding the budget after all evictable tensors are removed. The authors attribute the LSTM instability to the joint size‑staleness scoring term and argue that these represent two distinct budget‑sensitive pathologies rather than a single mechanism.

By Mahesh Reddy Pagadala
arXiv AI
Jul 21

First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers

arXiv:2607. 16821v1 Announce Type: cross Abstract: Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbations around an already trained checkpoint, and they rely on different local approximations: individual perturbations should be first-order predictable, task updates should compose with controlled interference, useful tangent structure should be stable and possible to estimate, and weight edits should have counterparts in representation space.

By Irina Piontkovskaia, Sergey Nikolenko
arXiv Machine Learning
Sep 23

Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference

The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead. "whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."

By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
Hugging Face Trending Papers
Jul 5

Asymptotic-Preserving A Posteriori Analysis of Diffusion and Flow-Matching Samplers

Diffusion and flow-matching samplers integrate a learned probability-flow ODE from a large noise scale down to a small terminal floor $σ_{\min}$, at which the score is stiff and the flow develops a boundary layer. We treat $σ_{\min}$ as a singular-perturbation parameter and determine which fixed-step samplers are asymptotic-preserving (AP), that is, stable and uniformly accurate as $σ_{\min}\to0$, casting the criteria as an a posteriori audit: residual functionals with $σ_{\min}$-uniform coefficients, computable on a pretrained checkpoint without ground-truth scores or exact trajectories.