arXiv:2608. 05025v2 Announce Type: replace Abstract: Joint Energy-Based Models (JEM) unify classification and generation within a single network and support out-of-distribution (OOD) detection.
By Dmytro Knopov
arXiv:2608. 13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable.
By Teng-Ruei Chen
FORGE is a forward‑only test‑time adaptation technique designed for integer‑only vision models running on microcontrollers. It restores batch‑normalization statistics after BN folding by re‑normalizing each convolution’s per‑channel output using only forward‑pass estimates, enabling adaptation on deployed, folded integer models. The method achieves accuracy gains comparable to gradient‑based TENT, requires adapting only a few layers, works with single‑sample streaming, and has been validated on an ESP32‑S3 with minimal energy and latency overhead.
By Muhammad Rehan, Haider Ali, Muhammad Ali Munir, Moaz Amjad
arXiv:2607. 19393v1 Announce Type: cross Abstract: While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.
By Vishnu Bindu Balachandran
This paper presents a reproducibility audit of frozen‑encoder anomaly detection experiments originally reported on arXiv. The authors confirm that the numerical discrimination results can be reproduced from the preserved artifacts, but they find that the claimed causal link to interferometric pretraining is unsupported. They show that near‑zero embeddings and architectural choices, rather than a morphological prior from gravitational‑wave instrumentation, explain the observed anomaly‑detection performance.
By Jose S\'anchez Andreu
The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.
By Zachary Speck, Asa Shepard
The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.
By Arunan J
The paper demonstrates that greedy decoding from large language models is not precision‑invariant: the same model, prompt, and decoding algorithm can produce different outputs when run in BF16 versus FP16 on identical hardware. Across six models (1.1B–7B parameters, four families, and 12B) and three benchmarks, 49–100 % of prompts diverge, with a single token flip often cascading into trajectory‑level divergence. The authors develop an empirical error‑propagation analysis that identifies the top‑two logit margin at the LM head as the key factor, and they propose a low‑overhead intervention—selective FP32 LM head recomputation—that improves exact agreement by 22–36 percentage points with less than 4 % latency overhead.
"whyItMatters":"The findings reveal that precision choices can fundamentally alter model outputs, challenging the assumption of deterministic greedy decoding and highlighting the need for precision‑aware inference strategies."
By Gaoyuan Du, Anam Nawaz Khan, Rex Zhou, Xiaoyang Liu, Deepayan Chakrabarti, Fnu Suya, Xueping Li
The paper introduces a schema‑adaptive action‑conditioned Joint‑Embedding Predictive Architecture (SAAC‑JEPA) for cross‑machine CNC transfer when only a subset of sensors overlap between source and target machines. Experiments show that pretraining does not improve source‑only forecasting, but a carefully selected action‑conditioned JEPA model achieves a zero‑shot RMSE of 0.546 on the target, outperforming persistence but falling short of certain baseline models. Ablation studies reveal that adding RevIN improves RMSE but harms calibration, and limited post‑lock adaptation can further reduce error.
By Ayoub Louaye Bouaziz, Matthieu Ostertag, Anton Demasles
arXiv:2606. 09682v1 Announce Type: new Abstract: AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA.
By Jaber Jaber, Osama Jaber
arXiv:2606. 20502v1 Announce Type: cross Abstract: Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains unresolved.
By Arastoo Zibaeirad, Marco Vieira
arXiv:2606. 22054v2 Announce Type: replace-cross Abstract: Detectors for GNSS radio-frequency impairments (jamming, spoofing, multipath) are usually reported with a single AUC measured on the distribution they were tuned on.
By Chakshu Baweja