arXiv:2610. 02191v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions.
By Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu
The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.
By Robert Hu
EchoPress is a training‑free method for pruning key‑value caches in large language models. It approximates the reconstruction attention used by KVzip by leveraging queries and keys from standard prefill, reconstructing only the first chunk to calibrate importance scores for the rest of the context. Experiments on LongBench and RULER with Qwen3‑8B and Llama‑3.1‑8B‑Instruct show that EchoPress matches KVzip’s task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by 1.7–19.6× and total prefill time by up to 2.9×.
By Jiawei Lin, Saibo Geng, Thomas Bourgeat
MoRA is a framework for pruning Mixture-of-Experts (MoE) models by learning a router bias for each expert and optimizing it with a language‑modeling loss and a routing‑diversity regularizer. The learned biases sharpen routing distributions to identify critical experts and encourage diverse routing preferences. After pruning, MoRA uses an expert approximation mechanism that approximates the outputs of pruned experts with affine transformations of remaining experts, further improving performance.
By Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei
The paper presents a method for distilling tabular foundation models (TFMs) into lightweight, dataset‑specific students. By using the full labeled training set as teacher context and training students on both observed and synthetic queries, the authors achieve significant performance gains over traditional supervised models on TabArena and TALENT benchmarks. The distilled students also provide substantial inference speedups, reducing the cost of repeated inference.
By Minho Jeong, Dooho Lee, Jinmo Lee, Jaemin Yoo
FedSAP is a federated learning framework that addresses heterogeneous edge devices by using structured pruning as a budget-constrained tri-state channel allocation. It partitions model channels into a Global pool, pseudo-domain-specific Private pools, and a Dropped state, allowing broadly useful features to be shared while isolating domain-sensitive updates. Experiments on Digits and Office-Caltech datasets show FedSAP achieving higher mean global accuracy than the strongest baseline while supporting up to 80% client pruning ratios.
By Wentao Yue, Tianyou Lai, Hongji Li, Qingyu Mao, Qilei Li
BranchIP introduces a single-model framework that learns adaptive tensor product computation for equivariant machine learning interatomic potentials (MLIPs). Using a novel distillation loss, it achieves up to 2.4× speed‑up and 2.6× memory reduction across model sizes while preserving physical fidelity. The adaptive computation also offers interpretability by indicating which interactions require deeper processing and how depth correlates with chemical complexity and dynamics.
By Laura Zichi, Gil Harari, Chuin Wei Tan, Marc L. Descoteaux, Albert Zhu, Menghang Wang, Yoel Zimmermann, H. T. Kung, Boris Kozinsky
The paper introduces LoRA‑Norm, a post‑training normalization technique for Low‑Rank Adaptation (LoRA) that rebalances the gains of learned singular directions without altering the directions themselves. LoRA‑Norm uses spectral rebalancing and nuclear‑norm restoration to preserve total spectral mass, requiring no calibration data or extra training and adding no inference overhead. Experiments on two backbones and three adaptation tasks show that LoRA‑Norm improves both specialization and capability retention, outperforming other post‑hoc spectral pruning and gradient‑guided editing methods.
By Zailong Tian, Yanzhe Chen, Zhuoheng Han, Houfeng Wang, Lizi Liao
AIR-LLM is an edge inference architecture that broadcasts large language model (LLM) weights over radio, allowing edge devices to perform matrix-vector multiplications directly in the RF domain without storing or loading the weights. The system uses MIMO spatial multiplexing and an energy‑efficient precoder‑postcoder pair to reduce airtime and calibrate the wireless channel, enabling a single broadcast to serve unlimited users. Experiments on real urban channel models show that AIR-LLM achieves only a 4.0% perplexity loss on LLaMA‑3.1‑8B while saving energy by up to 157.7× compared to FP16 and reducing airtime by over 100× for 20 users.
By Zhihui Gao, Tingjun Chen, Dirk Englund
The paper investigates how reducing numerical precision through model quantization impacts the vulnerability of neural networks to model inversion attacks. It provides theoretical bounds on mutual information changes and identifies data-dependent effects, especially at 4‑bit precision. Based on these findings, the authors propose a privacy‑aware post‑training quantization strategy that allocates bits adaptively, calibrates activation ranges, and jointly optimizes weight and activation scaling to improve inversion resistance while preserving model utility.
By Rongke Liu, Youwen Zhu
The study evaluates 4‑bit quantization and low‑rank adapter fine‑tuning (QLoRA) on several large protein language models, finding that many model‑task pairs retain over 90% of full fine‑tuning performance while achieving up to 90% GPU memory savings. QLoRA preserves early‑layer representations and induces task‑specific changes in later layers, closely resembling full fine‑tuning with smaller representational shifts. For generative models, 4‑bit quantization largely maintains structural and sequence‑level properties, though token‑level analysis reveals model‑dependent changes in autoregressive output distributions.
By Ilan Yaniv Zeisler, Sebastian Clancy, Pouriya Bayat, Saaim Raad, Ivan Kraskov, Matthew Xie, Vivian White, Spencer Perkins, Serena Singh, Sepehr Bayat, Keith Pardee
The paper introduces entry-state sharpening, a data‑free pre‑training step that prepares a language model’s checkpoint in a sharper, lower‑entropy state before test‑time reinforcement learning (TTRL). By reducing policy entropy, the model can more efficiently use its limited adaptation budget, leading to higher endpoint conversion efficiency across tasks such as MATH, GPQA, and AMC. Experiments with different data‑free objectives (e.g., R‑Zero vs. SPIRAL) demonstrate that the choice of pre‑training objective strongly influences the checkpoint’s readiness for TTRL, and a label‑free self‑distillation intervention can further sharpen the entry state.
By Zhanming Zhang, Vinoth Selvendran
arXiv:2610. 01088v1 Announce Type: cross Abstract: The nonparametric maximum likelihood estimator (NPMLE) of a Gaussian location mixture maximizes the likelihood over the infinite-dimensional space of mixing distributions.
By Hansheng Jiang
arXiv:2605. 19145v3 Announce Type: replace Abstract: In the literature, many continual learning (CL) algorithms have been proposed to address the issue of catastrophic forgetting in ML models (i.
By Srijith Nair, Atilla Eryilmaz, Jia Liu
JET (Justification Evaluation in Transformer) leverages pretrained language and vision‑language models to choose among a limited set of answers without extra training. It directly evaluates candidate likelihoods, reuses computation across candidates, and runs experiments on desktop CPUs and consumer GPUs to measure decision accuracy and execution cost. Results show high accuracy on the MMLU test set, significant speedups from prefix reuse and cache management, and a 30.8% reduction in process time through input preparation optimizations, all while maintaining unchanged outputs.
By Shenghao Ding
The paper proposes ANCHOR, a composition policy that decouples crystal structure prediction (CSP) from de novo crystal generation (DNG) by using a frozen CSP model as a fixed evaluator and training the policy with multi‑objective rewards, including adaptive novelty. ANCHOR significantly improves metrics such as MSUN and SUN compared to traditional DNG approaches, and demonstrates transferability across different CSP backbones and distillation into Crystalite‑CSP. The study shows that fine‑tuning DNG models directly on these rewards yields limited gains, whereas the ANCHOR framework better leverages the CSP prior for more efficient crystal discovery.
By Emma Lei Hovmand, Jonas Elsborg, Melih Kandemir, Arghya Bhowmik
The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.
By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
REALM is a retrospective knowledge distillation framework that enables causal decoding of behavior from local field potentials (LFPs). It trains a bidirectional Mamba‑2 teacher on multi‑session data using continuous masked autoencoding, then distills its representations into a compact causal student model. The resulting LFP‑only decoder achieves the highest mean accuracy among compared methods, surpassing state‑of‑the‑art baselines while using fewer parameters and less pretraining time.
By Peicheng Wu, Zhenyu Bu, Runze Ma, Lin Du
TopK-Guided is a training‑free method that improves activation sparsity for large language model inference by combining token‑level sparsity adaptation with block‑level budget allocation that accounts for block sensitivity. It addresses limitations of existing methods like TEAL, which adapts sparsity per token but lacks tight control, and WINA, which enforces a fixed sparsity across all tokens and blocks. Experiments on Llama‑2 and Llama‑3 show that TopK‑Guided consistently yields better perplexity and downstream accuracy while maintaining similar compute costs to WINA, especially at high sparsity levels.
By Mukund Agarwalla, Chih-Jen Lin
The paper introduces Quasar, a model‑free Q‑learning algorithm that guarantees asymptotic convergence for reachability objectives in Markov Decision Processes that are free of non‑terminal maximal end components (MECs). Unlike prior model‑based methods, Quasar does not estimate transition probabilities, reducing memory usage from O(|S|²|A|) to O(|S||A|). Experiments on the Quantitative Verification Benchmark Set show that Quasar converges to optimal policies with far fewer samples than existing state‑of‑the‑art model‑based approaches.
By Lu-Chin Chang, Suguman Bansal