arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin
arXiv:2508. 14268v2 Announce Type: replace-cross Abstract: Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2609.07031v1 Announce Type: new
Abstract: Learning with group invariances is central to many scientific and geometric learning problems, yet its computational foundations remain poorly understo...
By Ashkan Soleymani, Behrooz Tahmasebi, Patrick Jaillet, Stefanie Jegelka
Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.
arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2608. 12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
By Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2604. 03146v2 Announce Type: replace-cross Abstract: We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs.
By Chiheb Yaakoubi, Cosme Louart, Malik Tiomoko, Zhenyu Liao
The paper proposes three information‑theoretic criteria for selecting the most relevant basis functions in sparse Gaussian process regression, tailored to different levels of prior knowledge. Experiments on six UCI regression datasets and three basis families (HSGP, VFF, VISH) show that the no‑data criterion is a robust default, often outperforming simple truncation, while the data‑aware criteria yield significant improvements for HSGP. The study demonstrates that careful basis‑function selection can lead to better performance without increasing computational cost.
By Marnix Van Soom, Ivan De Boi
arXiv:2603. 27631v2 Announce Type: replace Abstract: Self-supervised pre-training, where large corpora of unlabeled data are used to learn representations for downstream fine-tuning, has become a cornerstone of modern machine learning.
By Mohammad Tinati, Stephen Tu
arXiv:2607. 24145v1 Announce Type: new Abstract: Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task.
By Muhammad Rajabinasab, Arthur Zimek