Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,731 stories · RSS feed

arXiv AI
Sep 24

Bridging LLM Serving and CXL-SSDs with Chunk-Aware KV Cache Management

The paper introduces LM‑CXD, a CXL‑SSD design tailored for large language model (LLM) prefix caching. By aligning KV chunk management between the serving engine and the storage device, exposing NAND-to‑DRAM progress, and using device DRAM as a GPU‑accessible buffer, LM‑CXD reduces time‑to‑first‑token (TTFT) by up to 4× compared to a stock CXL‑SSD and brings performance within 1.5× of local DRAM across five LLM models. The approach also incorporates windowed prefetching and layer‑wise KV movement to hide NAND latency under limited device DRAM.

By Hyunsun Chung, Taewan Noh, Minji Kim, Joo-Young Hwang, Hong-Yeon Kim, Youngjae Kim
arXiv AI
Sep 24

Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Crossflow introduces an elastic boundary for prefilling and decoding in large language model serving, allowing decode nodes to publish short‑lived leases that limit prefilling resources and output projections. By adapting to dynamic phase demand, Crossflow improves token throughput by 16.2‑17.4% on average and up to 43.4% under high load, while consistently reducing mean time‑to‑first‑token. The approach eliminates the inefficiencies of static partitioning, which can leave 17% of cluster capacity idle or cause queueing and lost throughput.

By Yi Xu, Ehsan K. Ardestani, Wenyin Fu, Martin Schatz, Krishna Malladi, Zhan Shu, Adnan Aziz, Shobhit Kanaujia, Ajit Mathews, Chunqiang Tang
arXiv AI
Sep 24

Combining LLMs and Genetic Search for ARC-AGI-2

The paper presents a hybrid approach that combines large language models (LLMs) with genetic algorithms to solve ARC-AGI-2 tasks. An LLM (Qwen3.5-4B) first generates a small set of programs, which seed a genetic algorithm that evolves these programs within a domain‑specific language ensuring validity. This method yields 6 correct solutions out of 60 tasks (10%), outperforming either technique alone.

By Val Dyachenko
arXiv AI
Sep 24

Evolving Inspectable O-RAN Slicing xApps with LLMs

The paper presents a method for creating Open RAN (O‑RAN) slicing xApps that are both adaptive and fully inspectable. By using a large language model (LLM) to evolve compact Python programs, the authors generate slicing controllers whose decision logic remains readable and editable after optimization. In experiments on the NSF POWDER 5G testbed, the evolved controllers improved best‑effort throughput by 44.5% and reduced SLA misses from 79.9% to 2.2%, while evolutionary search outperformed independent prompting in trace‑driven simulations.

By Faezeh Dehghan Tarzjani, Bhaskar Krishnamachari
arXiv AI
Sep 24

Geometry-Conditioned Visual Place Recognition in Natural Environments

The paper introduces Depth‑Aware Distillation (DAD), a method that conditions a pretrained Vision Foundation Model’s token representations on geometry inferred by a Geometric Foundation Model, without using a depth sensor. DAD projects image‑aligned depth into the VFM token space and selectively modulates visual representations through channel‑wise geometric conditioning. On the WildCross benchmark, DAD raises average inter‑sequence Recall@1 from 61.41% to 66.37% and Recall@5 from 65.86% to 72.49%, especially improving performance under reverse traversal and long‑term appearance variation.

By Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya, David Hall, Kaushik Roy, Peyman Moghadam
arXiv AI
Sep 24

What Looks Like a Capability Limit in Vision-Language Models Is a Readout Limit

The paper argues that apparent capability limits in vision‑language benchmarks often stem from the way answers are presented rather than from the models themselves. By comparing performance on COCO images with answer choices given as English names versus pixel coordinates, the authors show that models like Qwen3‑VL‑4B perform far better when answers are in natural language, and that the choice of answer format can swing model rankings by dozens of points. The study also demonstrates that different conventions (e.g., hue angles vs. pixel coordinates) reveal which formats a model can actually interpret, highlighting that a fixed answer vocabulary is not neutral across models.

By Alfredo F. Frontera Del Valle
arXiv AI
Sep 24

MemBodied: Recurrent Associative Memory for Vision-Language-Action Models

MemBodied introduces a fixed‑size episodic memory for Vision‑Language‑Action models, comprising an associative state that tracks interactions across policy calls and an episode anchor that stores a compact representation of the initial scene. By conditioning action generation on these memory components instead of raw past observations, MemBodied reduces context bloat and inference latency. In five memory‑dependent RMBench tasks, it outperforms stateless and vanilla recurrent policies by significant margins, and achieves a 90.6% success rate on the LIBERO‑Long suite, improving over the baseline by 5.4%.

By Tej Deep Pala, Navonil Majumder, Bryce Goh, Raphael Yee, Jianfei Yang, Liming Chen, Soujanya Poria
arXiv AI
Sep 24

When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment

The paper introduces UECR-GRPO, a method that unifies on‑policy distillation and verifier‑based reinforcement learning for mathematical reasoning. It combines verifier rewards and teacher‑derived log‑ratios into a single KL‑regularized objective (Path‑Utility Unification) and then redistributes credit at the token level using entropy‑calibrated redistribution, preserving total task credit. Experiments on five benchmarks show that UECR‑GRPO improves average accuracy by up to 0.89 percentage points over the best baseline for both Qwen3‑1.7B and Qwen3‑4B students.

By Jie Zhang, Jingxiao Yang, Zhehao Huang, Yuhang Liu, Xiaolin Huang
arXiv Machine Learning
Sep 24

RL Starts before RL: On Policy Distillation for Better Reinforcement Learning

The paper investigates on‑policy distillation (OPD) as a preparatory step for reinforcement learning (RL). It shows that students initialized with OPD achieve higher final RL performance than those trained directly with RL or with supervised fine‑tuning followed by RL, even when OPD offers little immediate accuracy gain. The study also finds that the choice of distillation objective (reverse‑KL vs forward‑KL) and the source of trajectories influence OPD’s effectiveness at different stages of RL training.

By Shuai Dong, Yongfu Zhu, Yuqi Xu, Weichu Xie, Liuwenpu, Ziyue Wang, Kaiwen Tuo, Congcong Wang, Siyuan Wang, Wenqi Shao, Shuai Yang, Ji Zhao, Caoyuan Ma, Wenzheng Chang, Taiqiang Wu, Xinlei Yu, Hongrui Wu, Xiaoxuan He, Fangke Chen, Dianyi Wang, Kanghui Tian, Sirry Chen, Xingyu Liu, Xiangnan Wu, Jiawei Guo, Haowen Hou, LingHan Chen, Zhongyu Wei, Jiaqi Wang
arXiv AI
Sep 24

AnchorReasoning: A Visual Grounding and Causal Reasoning Dataset in Long-Tail Autonomous Driving Scenarios

arXiv:2609. 28366v1 Announce Type: cross Abstract: Vision-language models (VLMs) offer a promising approach to long-tail autonomous driving, but existing driving datasets provide limited supervision for connecting decision-critical visual evidence with reasoning and planning.

By Zhipeng Bao, Wenjie Zhao, Tianle Zhu, Haohua Que, Chence Yang, Geng Yuan, Qianwen Li
arXiv Machine Learning
Sep 24

Confidence Falls Short: Asymmetric Certainty Gains from Optimization Hinder Multimodal Classification

The paper identifies that in multimodal learning, optimization often produces asymmetric certainty gains, with the stronger modality becoming more confident than the weaker one, which leads to imbalanced contributions and suboptimal performance. The authors attribute this issue to unimodal characteristics and propose a Max Confidence Regularization (MaxCR) method that tracks each modality’s semantic confidence via a nonlinear sparsity measure and applies max suppression and excitation to balance confidence levels. Experiments on standard datasets demonstrate that MaxCR improves overall performance compared to state‑of‑the‑art multimodal baselines.

By Longfei Huang, Xiangyu Wu, Yang Yang
arXiv Machine Learning
Sep 24

NeuroRule: Making Black-Box Neural Networks Explainable through Rule-set Evolution

NeuroRule is a knowledge distillation framework that transforms high‑capacity neural networks into explainable rule‑sets. It adapts the EVOTER rule‑set evolution infrastructure to evolve propositional logic expressions that capture the neural network’s performance. The approach includes a conciseness objective to enhance explainability and demonstrates viability even without access to the original training data.

By Tapaswini Kodavanti, Hormoz Shahrzad, Risto Miikkulainen
arXiv AI
Sep 24

DRSR: Learning Set-Level Deletion Risk for Efficient Long-Horizon Agents

The paper introduces Direct Relational Set‑Risk Pruning (DRSR), a method for compressing the history of long‑horizon language‑model agents by selecting deletion sets based on risk constraints rather than independent unit scores. DRSR builds counterfactual supervision offline, then uses a lightweight scorer to predict set‑level harm during deployment, removing the largest safe set while respecting recency, protocol, and budget limits. Experiments on WorkBuddyBench Full260 and Eval40 show that DRSR improves mean reward from 0.699 to 0.802 and reduces token usage by over 20%, with further analyses highlighting the importance of decision‑conditioned relations, retained context, pair interactions, and abstention.

By Mingxuan Wang, Bo Wang, Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han
arXiv AI
Sep 24

Sparse-Observation Atmospheric Thermal Forecasting with Physics-Informed Neural Networks for Climate-Aware Digital Twins

The paper evaluates a physics-informed neural network (PINN) for short‑horizon atmospheric temperature forecasting where observations are sparse. Using ERA5 data at three pressure levels, the PINN outperforms persistence, local‑trend, and two neural‑network baselines, with mean RMSE improvements ranging from 8.1 % at one hour to 23.8 % at three hours. The advantage persists under severe observation sparsity and transfers across regions, though it degrades in complex terrain, highlighting limits of a fixed vertical‑coordinate representation.

By Tannaz Goodarzvand Chegini, Elyas Shivanian, Behzad Karimi, Faraz Dadgostari
arXiv AI
Sep 24

Learn How to Act from Your Own Interactions: On-Policy Self-Distillation for GUI Agents

The paper introduces GUI‑SD‑v2, an on‑policy self‑distillation framework that extends previous methods from GUI grounding to multi‑turn GUI interaction. It employs a two‑stage training process that first improves privilege following by jointly optimizing rollouts with and without privileged guidance, then selectively distills step‑specific reasoning and memory guidance via a privilege‑conditioned self‑teacher. Experiments on AndroidWorld and MobileWorld benchmarks demonstrate that GUI‑SD‑v2 outperforms existing OPSD baselines and state‑of‑the‑art methods in Pass@1 and Pass@3 success rates.

By Yan Zhang, Daiqing Wu, Huawen Shen, Liang Li, Gang Cao, Zhi Gong, Wei Dai, Xiaode Zhang, Can Ma, Yu Zhou
arXiv AI
Sep 24

SMDDFNet: State-space Modeling and Dynamic Dual Fusion Network for Traffic Sign Detection

SMDDFNet is a deep learning detector designed for traffic sign images, addressing challenges such as small objects, scale variation, and occlusion. It combines a Dynamic Dual Fusion (DDF) module—integrating multi-scale attention and frequency‑domain dynamic filtering—with a state‑space modeling backbone that captures long‑range dependencies efficiently. A multi‑scale feature fusion neck further aggregates pyramid features, enabling robust localization of small signs while maintaining real‑time throughput on datasets like TT100K, GTSDB, PASCAL VOC, and Roboflow.

By TianYi Yu, DaJian Zhong, Lilin Wang
arXiv Machine Learning
Sep 24

EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

EBRL is an asynchronous embodied reinforcement learning training system that overlaps rollout and training stages, pipelines simulation and generation across environment groups, and eliminates synchronization stalls. It employs a fine‑grained resource manager that pools CPU cores and GPU streaming multiprocessors, dynamically adjusting resource quotas and batch sizes based on stage profiles and runtime feedback. Experiments on RLinf with four policies and four simulation benchmarks show EBRL improves rollout throughput by 1.30–3.47× and training convergence by 2.5× over state‑of‑the‑art embodied RL systems.

By Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao
arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu
arXiv Machine Learning
Sep 24

Reliable Federated TinyML Deployment for IoT Security

The paper explores how to combine Federated Learning with TinyML model compression techniques—such as knowledge distillation, structured pruning, and quantization—to create lightweight, privacy‑preserving intrusion detection systems for IoT devices. It evaluates these strategies within a federated training pipeline and finds that training stability is crucial; a server‑coordinated cosine learning‑rate schedule boosts Attack Recall from 46.7% to 93.85% while still allowing significant model compression and efficient edge deployment.

By Younsoo Park, Seokhyoen Bae, Shasi Kumar Ramachandran Prabhu, Suman Saha, Peilong Li
arXiv Machine Learning
Sep 24

Dirichlet Process Mixtures of Trees with Gaussian Process Splits: A Bayesian Nonparametric Framework with Posterior Contraction Rate

arXiv:2609. 27930v1 Announce Type: cross Abstract: We propose a Bayesian nonparametric mixture of regression trees with a Dirichlet process prior over tree-parameter pairs, enabling data-driven selection of ensemble size and unifying CART, BART, random forests, and boosting.

By Subhasish Basak, Anik Roy, Sourabh Bhattacharya