arXiv:2608. 01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows.
By Xizhe Zhang
arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.
By Vasileios Sevetlidis
arXiv:2606. 12138v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs.
By Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov
The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.
By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv:2604. 27967v2 Announce Type: replace Abstract: Background: We introduce StructGP, a continuous-time multi-task Gaussian process that couples process convolutions with differentiable structure learning to uncover a sparse, ordered directed acyclic graph (DAG) of inter-variable dependencies while preserving principled uncertainty.
By Ivan Lerner, Jean Feydy, Alexandre Kalimouttou, Anita Burgun, Francis Bach
arXiv:2510. 12744v2 Announce Type: replace-cross Abstract: We develop a unified statistical framework for softmax-gated Gaussian mixture of experts (SGMoE) that addresses three long-standing obstacles in parameter estimation and model selection: (i) non-identifiability of gating parameters up to common translations, (ii) intrinsic gate-expert interactions that induce coupled differential relations in the likelihood, and (iii) the tight numerator-denominator coupling in the softmax-induced conditional density.
By Do Tien Hai, Trung Nguyen Mai, TrungTin Nguyen, Nhat Ho, Binh T. Nguyen, Christopher Drovandi
arXiv:2502.04131v2 Announce Type: replace
Abstract: The successful application of modern machine learning for time series classification is often hampered by limitations in quality and quantity of av...
By Janis Norden, Elisa Oostwal, Michael Chappell, Peter Tino, Kerstin Bunte
The paper introduces ePID, an embedding-based approach that scales partial information decomposition (PID) to large symptom networks by compressing non‑focal symptoms into a low‑cardinality discrete embedding. Using a supervised Agglomerative Conditional Information Bottleneck (ACIB) embedding, ePID accurately recovers source‑unique, remainder‑unique, redundant, and synergistic components for each ordered source‑target pair across 83 real‑world datasets, outperforming 12 other embeddings. Applied to PHQ‑9 and the Interpersonal Reactivity Index, ePID reveals distinct patterns of redundancy and synergy that align with each instrument’s construction, demonstrating its ability to separate overlapping from interaction‑dependent information in symptom networks.
By Cillian Hourican, Eric Dignum, Rick Quax, Debraj Roy
arXiv:2607. 26528v1 Announce Type: cross Abstract: Symbolic regression provides analytical expressions, but it is usually applied one output at a time.
By Manuel Rodriguez
arXiv:2609.23937v1 Announce Type: cross
Abstract: Robust linear fits can resist response contamination yet remain too dense or unstable for useful global explanations. We propose penalized distillati...
By Wooyoung Shin, Seunghwan Park
arXiv:2607. 22313v1 Announce Type: cross Abstract: Estimating contemporaneous bidirectional interactions from observational data is difficult because each outcome is endogenous to the other, while flexible regressions may capture only reduced-form dependence.
By Masahiro Tanaka
arXiv:2603. 25414v4 Announce Type: replace-cross Abstract: A prevailing assumption in machine learning is that model correctness must be enforced after the fact.
By Houston Haynes