Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,725 stories · RSS feed

arXiv AI
5d ago

DualCast: A Dual-Path Language Model for Bimodal Financial Time-Series Forecasting

DualCast is a dual‑path language model that forecasts financial time‑series by combining a fast numerical forecaster with an optional text‑conditioned revision mechanism. The fast path trains only new financial‑token embeddings and output heads on a frozen Qwen3‑8B backbone, while the slow path uses a LoRA adapter to incorporate news and refine predictions. In zero‑shot tests across equities and energy prices at multiple time resolutions, the slow path achieves the lowest mean absolute percentage error in most settings, especially for longer horizons, and news ablations show additional gains in many markets.

By Wentao Zhao, Hongqiang Wu, Shanghang Liu, Zhaochen Zan, Yu Zhang, Biqing Huang
arXiv AI
5d ago

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Aegis is a client‑side defense for medical federated learning that protects against model inversion attacks by adding a masking gradient derived from locally synthesized data. The method exploits the fact that attacks fail when the effective batch size exceeds the model’s leakage capacity, turning this bottleneck into a privacy guarantee. Experiments on MNIST, CIFAR‑10, and MedMNIST datasets show that Aegis neutralizes state‑of‑the‑art attacks while preserving model accuracy and adding only modest overhead.

By Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
arXiv AI
5d ago

ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights

ShamAN-Q is a sub‑1‑bit post‑training quantization technique that builds on NanoQuant by replacing its diagonal reconstruction geometry with a dense curvature metric inspired by the Shampoo optimizer. For each linear weight, it fits a Kronecker product to the empirical Fisher information matrix of a small calibration set via Kullback–Leibler minimization, yielding a Mahalanobis reconstruction loss. The method updates continuous ADMM steps to Sylvester equations while keeping the discrete projection and deployment format unchanged, and it redistributes uniform rank across layers, achieving lower perplexity on Qwen3‑Base at roughly 1 bpw and matching or improving zero‑shot accuracy on the Eleuther LM Evaluation Harness.

By Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler
arXiv AI
5d ago

dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale

The paper introduces dattri-LLM, a library designed to make training data attribution (TDA) practical for large language models. It achieves efficiency by using compact gradient representations and a cost‑based routing system, while maintaining compatibility by capturing per‑example gradients from existing training loops without modifications, even in distributed settings. The library also offers extensibility through reusable gradient operations and callbacks, supporting various attribution methods and applications such as online data selection, and demonstrates significant performance gains and scalability up to 110B‑parameter models.

By Shixuan Liu, Tongli Zhou, Junwei Deng, Pingbang Hu, Jiaqi W. Ma
arXiv AI
5d ago

Audio Token Attention Is Predictable Before the Language Model Runs

The paper introduces Triage, a method that predicts the attention distribution of audio tokens before a language model processes them, enabling early pruning of less important tokens. By fitting a linear map to encoder outputs, Triage achieves high correlation (ρ ≥ 0.69) with full-model attention across eleven of thirteen large audio language models. Using this prediction, Triage compresses audio inputs while maintaining near‑full performance, outperforming baselines in transcription accuracy and significantly increasing the amount of audio that fits within a model’s context window.

By Kyoungjun Park, Yunzhe Li, Lili Qiu
arXiv AI
5d ago

DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?

DrivingBench is the first benchmark that tests general‑purpose vision‑language models on the task of driving a real Toyota Corolla around a parking‑lot cone course. The models receive live camera frames and issue steering and velocity commands, with inference latency counted as part of the challenge. In tests, only GPT‑6 Astra completed the course, while other models showed limited progress or failed to pass half the course.

By Aditya Ramabadran, Simon Mahns, Tobias Gessler
arXiv AI
5d ago

WinoTS: Wavelet-based Self-Distillation for Time Series Models

WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.

By Noam Major, Kathy Razmadze, Yoli Shavit
arXiv AI
5d ago

From Imitation to Reward Discovery: On-Policy Warmup for Agentic RL

The paper introduces On‑Policy Warmup (OPW), a teacher‑guided training stage where a student agent learns from a teacher on its own interaction trajectories before switching to reinforcement learning with verifiable rewards (RLVR). OPW differs from traditional imitation by focusing on states generated by the student’s own decisions, including imperfect actions and recovery situations. The authors provide a theoretical link between on‑policy reverse‑KL distillation and trajectory‑level distribution matching, showing that, under a competent teacher and low distillation loss, OPW can lower bound initial verifier success and reduce reward‑discovery complexity, thereby accelerating RLVR performance.

By Yitong Qiao, Tiantian He, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
arXiv AI
5d ago

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

OmniReasoning introduces a new benchmark, OmniReasoningBench, that requires both audio and visual evidence for answering 1,150 multiple-choice and open-ended questions across two tasks. The authors also develop OmniQA, a data engine that automatically generates evidence‑grounded QA pairs with time‑stamped clue chains, producing training datasets OmniReasoning‑SFT‑112K and OmniReasoning‑RL‑19K. Finally, they propose Modality‑Factored Self‑Distillation (MFSD), an on‑policy self‑distillation method that assigns token‑level credit by evaluating responses under modality‑specific clue contexts, enabling the OmniReasoning‑30B‑A3B model to achieve significant performance gains on both the new benchmark and existing video benchmarks.

By Junming Lin, Yuxuan Wang, Zhenxin Lei, Yuxin Liu, Ruixun Liu, Yinsong Yan, Ling Wang, Minghao Han, Yunfei Chu, Shun Lei, Xueyao Zhang, Qize Yang, Jin Xu, Yiwu Zhong
arXiv AI
5d ago

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

GroundingPI is a 4‑billion‑parameter grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It is trained with multimodal and spatial pretraining, supervised fine‑tuning, and reinforcement learning, achieving a new state‑of‑the‑art average of 73.68% across 34 grounding benchmarks. As a visual backbone, GroundingPI improves performance in robotic manipulation and autonomous driving, outperforming larger models and mainstream backbones in several out‑of‑distribution settings.

By Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang
arXiv AI
5d ago

TRACE: Trajectory Selection for Parallel Scaling of Search Agents

TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.

By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan
arXiv Computation and Language
5d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata
arXiv Computer Vision
5d ago

From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning

arXiv:2609.38924v1 Announce Type: new Abstract: Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical d...

By Jialu Pi, Yanan Ma, Weijie Chen, Owen Crystal, Shubham Trivedi, Stephen Xie, Anna Silverman, Matthew Stib, Chadi Ayoub, Reza Arsanjani, Imon Banerjee