Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Machine Learning
Sep 22

A discrete generative model of neuronal spiking activity on microelectrode arrays

The paper presents a discrete generative model for neuronal spiking activity recorded on microelectrode arrays. It uses a shared vocabulary of spatiotemporal motifs learned by a residual vector‑quantized autoencoder and predicts motif occurrence with a factorized masked transformer. Evaluated on 31 assays from human brain organoids and ex vivo hippocampal tissue, the model achieves superior reconstruction and generation performance compared to baselines and shows that motifs are largely reused across assays.

By Md Sayed Tanveer, Mohammed A. Mostajo-Radji, Ge Wang
arXiv Machine Learning
Sep 22

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

NAVIR is an end‑to‑end audio‑visual speech recognition system designed for the BrainChip Akida neuromorphic processor, which only supports sequential 2‑D convolutions. The architecture separates spatial and temporal encoding into three AkidaNet modules—per‑frame visual, temporal video, and spectrogram audio encoders—fused by a lightweight predictor and decoded with constrained beam search. Trained with CTC on noise‑augmented audio and fine‑tuned via quantization‑aware training, the quantized model achieves 14.0% WER on GRID’s unseen‑speaker split and 3.3% on overlapped‑speaker split, outperforming audio‑only baselines, and delivers 98.6% command accuracy at 1.5% WER on an industrial‑command corpus, while offering a 13‑fold energy advantage over conventional ANNs and roughly 5‑fold lower energy per inference than a Raspberry Pi CPU.

By Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis
arXiv Machine Learning
Sep 22

Anatomy of a Closed-Loop Collapse: A Causal Case Study of a Compressed VLA Policy

The paper presents a causal analysis of a compressed VLA policy that performs well in offline tests but fails in closed‑loop execution on a simulated pick‑and‑place task. An 8‑layer distillation of Octo‑Base retains most parameters and passes all offline metrics, yet collapses during deployment, with early stages degrading gradually and final transport failing entirely. The failure is traced to a negative, late‑heavy residual in the action trace, and standard remedies (continued training, offline data, command‑level compensation, clamping) do not restore performance; only a minimal‑pair intervention that mixes deployment‑distribution rollouts with teacher data restores parity with the teacher. whyItMatters":"The study demonstrates that offline validation metrics alone are insufficient to guarantee closed‑loop success for compressed policies, highlighting the need for targeted deployment‑time testing and interventions."

By Fengze Jia (The Ohio State University)
arXiv Machine Learning
Sep 22

Real-Time Plasma State Prediction via FPGA-Accelerated Quantized Recurrent Probabilistic Neural Networks

The paper presents an end‑to‑end workflow for deploying a recurrent probabilistic neural network (RPNN) on FPGA hardware to enable real‑time plasma state estimation for Tokamak control. By reducing architecture size and applying quantization‑aware training with QKeras, the model is synthesized with hls4ml for a Xilinx Alveo U50 device, achieving deterministic sub‑10 µs single‑timestep latency while staying within all four resource budgets (DSP, LUT, FF, BRAM).

By Daniel Gaytan-Villarreal, Aiken Xie, Tu Pham, Rohit Sonker, Chiara Amendola, Matteo Cremonesi, Cong Hao, Jeff Schneider
arXiv Computer Vision
Sep 22

GraphSVR: q-Space--Aware Graph-Based Slice-to-Volume Registration for Diffusion MRI

GraphSVR is a q‑space‑aware graph‑based framework for 4D slice‑to‑volume registration in diffusion MRI. It models slice groups as nodes in a graph whose edges encode temporal, spatial, and diffusion‑encoding relationships, and uses a graph neural network to predict globally consistent stack‑wise rigid motion in a self‑supervised, zero‑shot manner. In synthetic and realistic simulations, GraphSVR reduces grid and rotation errors by up to 73% compared to the standard FSL eddy method, especially under severe motion and sparse‑direction regimes.

By Noga Kertes, Daphna Link Sourani, Alex M. Bronstein, Moti Freiman
arXiv Machine Learning
Sep 22

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation explores how on‑policy distillation (OPD) can be improved by addressing teacher uncertainty that contracts as the teacher continues from a student‑generated prefix. The authors identify Teacher Uncertainty Contraction (TUC) and theoretically analyze its variance‑bias trade‑off, leading to the proposal of Adaptive‑Continuations On‑Policy Distillation (AC‑OPD). Experiments on mathematical reasoning and code generation show that AC‑OPD consistently outperforms standard OPD, with controlled‑continuation and matched‑budget analyses supporting the adaptive‑continuation design.

By Jingang Zhou, Yuyi Zhou, Haiyang Guo, Xukai Wang, Shuai Feng, Sirui Gao, Jian Xu, Qingpei Guo, Xu-Yao Zhang
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv Machine Learning
Sep 22

ValueDiff: Value-Geometric KV Cache Eviction for Sink-Suppressed LLMs

ValueDiff introduces a value‑geometric KV cache eviction strategy for large language models that suppress attention sinks. It ranks tokens by the L2 deviation of their value vectors from the cache mean, a score that aligns with minimal‑disturbance eviction under a max‑entropy assumption. Across several benchmarks—RULER, LongBench, and MATH‑500—ValueDiff consistently retains a higher proportion of useful tokens than prior methods, especially under tight cache budgets.

By Junyoung Park, Jungwook Choi, Mingu Lee
arXiv Computation and Language
Sep 22

RPMem: Learning Long-Term Recurrent Parametric Memory Across Sessions for LLM Agents

RPMem introduces a two‑stage architecture that compiles each session into a model‑independent latent memory and then consolidates it with retained memory via a task‑trained recurrent gate. The consolidated memory is mapped to backbone‑specific low‑rank adaptation (LoRA) parameters, enabling the memory to transfer when the backbone is replaced. Across three long‑term memory benchmarks and five diverse backbones, RPMem achieves broad generalization with near‑constant update cost and memory footprint, outperforming existing parametric and text‑based baselines on the PERMA benchmark.

By Fanyu Zhao, Ruike Cao, Liang Dong, Fugen Yao, Jian Xu, Guanjun Jiang, Han Zhang, Yifei Zhao, Yinsheng Li