arXiv:2605. 31191v2 Announce Type: replace Abstract: We investigate how teacher-student capacity relationships modulate knowledge distillation (KD) effectiveness in ResNet-based image classification on CIFAR-10.
By Umut Onur Yasar
arXiv:2608. 12652v1 Announce Type: cross Abstract: Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release.
By Florian Braun
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.
By Zachary Speck, Asa Shepard
The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.
By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran
arXiv:2603. 28921v3 Announce Type: replace-cross Abstract: The critical damping condition of the damped harmonic oscillator model of SGD with momentum (Qian, 1999) yields a momentum schedule with no tuned hyperparameters: mu(t) = 1 - 2*sqrt(alpha(t)).
By Ivan Pasichnyk
arXiv:2608. 05025v1 Announce Type: new Abstract: Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection.
By Dmytro Knopov
arXiv:2609.38956v1 Announce Type: new
Abstract: Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict wheth...
By Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen
arXiv:2605. 18838v3 Announce Type: replace-cross Abstract: Scaling laws predict loss from compute but not how capabilities interact.
By Adil Amin
arXiv:2606. 24903v1 Announce Type: new Abstract: Deciding when to stop collecting labeled examples is a fundamental but undertheorized problem in applied machine learning.
By Arnav Gupta
DeepGOF-1 introduces a pretrained convolutional network as a goodness‑of‑fit test for logistic regression, where the network reads a grid of standardized residuals as an image and outputs a test statistic. The test is fully calibrated via the analyst’s own bootstrap, ensuring the nominal level is maintained regardless of the network’s training. The authors prove exactness under pivotality, asymptotic exactness without it, and provide a computable consistency certificate from the frozen weights, demonstrating superior stability and power across multiple benchmarks and sample sizes.
By Ebrahim Khaled Ebrahim
The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead.
"whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."
By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li