arXiv Machine Learning

Renewable Lasso without Batch-Number Constraints: A Gradient-Enhanced Approach

arXiv:2606. 11738v1 Announce Type: cross Abstract: We study online estimation for high-dimensional generalized linear models with streaming data.

arXiv Machine Learning
Sep 14

High-Probability Convergence of SGD via Batched Updates

The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.

By Feng Zhu, Robert W. Heath Jr., Aritra Mitra
arXiv Machine Learning
Jun 18

BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training

arXiv:2606. 18650v1 Announce Type: new Abstract: As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories.

By Jiaxing Wang, Deping Xiang, Jin Xu, Zirui Liu, Zicheng Zhang, Guoqiang Gong, Jun Fang, Chao Liu, Pengzhang Liu, Tongxuan Liu, Ke Zhang, Qixia Jiang
arXiv Machine Learning
Jun 10

FedSLoP: Memory-Efficient Federated Learning with Low-Rank Gradient Projection

arXiv:2604. 24012v3 Announce Type: replace Abstract: Federated learning enables a population of clients to collaboratively train machine learning models without exchanging their raw data, but standard algorithms such as FedAvg suffer from slow convergence and high communication and memory costs in heterogeneous, resource-constrained environments.

By Yutong He, Zhengyang Huang, Jiahe Geng, Kun Yuan
arXiv Machine Learning
Aug 19

Online Generalized Sparse Regression: How Does Overparametrization Help?

The paper introduces an online generalized-sparsity-constrained regression framework that addresses key challenges in online sparse regression, such as dynamic regularization, memory usage, and real-time computation. It proposes an efficient online hard‑thresholding algorithm that performs closed‑form updates using only summary statistics, achieving global convergence at optimal statistical rates when the projection set is overparameterized. Numerical experiments show the method consistently outperforms existing alternatives in online cardinality‑constrained linear regression and low‑rank matrix sensing.

By Shuoguang Yang, Qiang Sun
arXiv Machine Learning
Jul 7

A Gradient Flow Perspective on Minimum MMD Estimation

arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.

By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv AI
Jul 3

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

arXiv:2602. 03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune.

By Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi
arXiv Machine Learning
Sep 18

FedFIbOS: Fisher Importance based Optimal Submodelling for Heterogeneous Federated Learning

FedFIbOS introduces a Fisher‑importance based criterion for selecting submodel parameters in heterogeneous federated learning, addressing the lack of theoretical justification in prior heuristic methods. By deriving a Fisher‑weighted quadratic masking surrogate and showing that the raw Fisher top‑k rule satisfies this surrogate under a Fisher‑dominant ranking condition, the method preserves convergence guarantees while efficiently estimating Fisher scores from squared gradients. Experiments on CIFAR‑10, CIFAR‑100, and AGNews demonstrate that FedFIbOS outperforms state‑of‑the‑art approaches by roughly 10% in accuracy, especially under strong non‑IID heterogeneity.

By Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang