arXiv:2609.15825v1 Announce Type: new
Abstract: A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's que...
By Dier Tang, Jing Yee Tan, Guangyue Han
arXiv:2607. 11950v1 Announce Type: cross Abstract: Brain field potentials are scale-free: their power spectra follow a $1/f^{\beta}$ law whose aperiodic exponent $\beta$ tracks cortical state, and sleep depth in particular is a shift in $\beta$.
By Dibakar Sigdel
arXiv:2606. 03003v1 Announce Type: cross Abstract: A latent world model built from an equivariant encoder $E$ and an equivariant predictor $f$ inherits a provable symmetry of its training loss: when the world's dynamics genuinely carries a group $G$ acting on latents by an orthogonal representation $\rho(g)$, the one-step prediction relMSE is exactly invariant across the whole group, so fitting the dynamics on a restricted slice of orientations mathematically determines it on the entire orbit (j\v{u} y\=i f\v{a}n s\=an).
By Hongbo Wang (Stony Brook University)
arXiv:2607. 13006v1 Announce Type: new Abstract: A growing family of indices scores how predictable a series is from its spectrum.
By Mert Onur Cakiroglu, Mehmet Dalkilic, Hasan Kurban
arXiv:2607. 12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior.
By Gunner Levi Howe
The study investigates whether the number of discrete class‑separability jumps (phase transitions) observed during ResNet fine‑tuning can predict final test accuracy. Across 75 experiments on four benchmarks (CIFAR‑10, CIFAR‑100, TinyImageNet, CIFAR‑10‑C) and three ResNet variants, a strong negative correlation is found on standard i.i.d. datasets (r = −0.84 on CIFAR‑10, r = −0.87 on CIFAR‑100), while the correlation weakens under distributional stress. Additional analyses show that the transition count retains predictive power after controlling for architecture depth and outperforms other training‑curve signals on in‑distribution benchmarks, though it is dominated by other signals on stressed datasets.
By Arunan J
arXiv:2608.20647v1 Announce Type: new
Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (...
By Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
arXiv:2606. 02670v1 Announce Type: cross Abstract: Many recent multivariate time series anomaly detection (MT-SAD) models incorporate cross-channel modeling, under the implicit assumption that the structure of anomalies may be spread across multiple channels.
By Marc Pinet (LIG), Julien Cumin (LIG), Samuel Berlemont (LIG), Dominique Vaufreydaz (LIG)
arXiv:2606. 24903v1 Announce Type: new Abstract: Deciding when to stop collecting labeled examples is a fundamental but undertheorized problem in applied machine learning.
By Arnav Gupta
arXiv:2608. 10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute $S$ requires a representation $Z$ that is statistically independent of $S$.
By Yijin Ni, Xiaoming Huo
The paper proposes a flexible four‑coefficient parameterization for per‑token gating in on‑policy knowledge distillation, unifying existing methods such as EOPD and ToDi as special cases. Experiments on TweetEval with a Qwen3 teacher‑student pair show that the full family of gating configurations outperforms single‑channel baselines in most cells, and dynamic gating beats static baselines in a majority of isolated comparisons. The authors present the framework mainly as a shared coordinate system for comparing gating designs rather than as definitive statistical evidence.
By Suwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang
The paper introduces the Intersection Euler Characteristic Profile, a topological metric for measuring class overlap in neural representations, and provides exact permutation and sign‑flip tests to assess disentanglement across layers. Using this statistic, the authors analyze 111 networks and 52,650 measurements, finding that disentanglement is depth‑graded, occurs early, and is influenced by training choices such as augmentation and weight decay. The study also demonstrates that the unnormalized mass of the profile predicts test accuracy, while the dimensionless quotient does not outperform simple linear probes.
By Sushovan Majhi