Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv AI
Sep 23

The Moral Check: Strategic AI Governance for the Pacing Problem

The paper titled "The Moral Check: Strategic AI Governance for the Pacing Problem" argues that technology cannot self‑steer and that strategy must guide AI development by ensuring purpose and judgment precede compute. It presents a dual contribution: a PRISMA 2020 review of 130 empirical studies and the Strategic AI Governance Ex‑Ante Framework (SAGE‑X), which operationalizes four strategic mindset pillars to mitigate velocity myopia, moral hazard, empirical hazard endpoints, and guardrail decay. The framework includes a calculable Moral Check Index and an Enterprise Lifecycle Audit Instrument to enforce that AI scaling does not outpace deliberative moral judgment, human agency, and societal trust.

By Zaid Amin, Rahma Santhi Zinaida, Nazlena Mohamad Ali
arXiv Machine Learning
Sep 23

The Virtue of Sparsity in Complexity

The paper investigates the trade‑off between sparsity and complexity in high‑dimensional asset pricing models. By separating capacity sparsity (restrictions on effective model capacity) from factor sparsity (parsimonious structure of priced risks), the authors use nonlinear feature expansions, basis pursuit, column generation, and GPU acceleration to estimate models with up to 432 million candidate factors. Their empirical results show that while sparse portfolios underperform dense ridgeless benchmarks at lower complexity, they achieve higher Sharpe ratios and lower pricing errors when the candidate set is large, indicating that capacity expansion and factor sparsity can complement each other.

By Nima Afsharhajari, Jonathan Yu-Meng Li
arXiv Computation and Language
Sep 23

Efficient Iterative Retrieval with Heterogeneous Batching

Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.

By Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park
arXiv AI
Sep 23

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

The paper evaluates seven graph database engines, including Corvic AI, on a synthetic biomedical property graph with 1.02 million nodes and 5.34 million rows. It benchmarks query latency, bulk‑ingest throughput, point‑update latency, and correctness across a twenty‑query workload that covers neighborhood lookups, bounded paths, set intersections, anti‑joins, aggregation, ranking, temporal filters, full scans, and relational joins. The study finds that no single engine is universally fastest; performance depends on query shape, and the largest cost difference arises from bulk‑ingest throughput, which varies by three orders of magnitude and dominates total cost for workloads with fewer than about 10⁵ queries per data refresh.

By Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach
arXiv AI
Sep 23

Recidivism Prediction, Peer Effect Estimation, and Prediction-Powered Inference with LLM Text Measures

The paper introduces a framework for estimating peer effects using multivariate behavioral measures extracted from written text via a large language model (LLM). It demonstrates that LLM embeddings enhance out‑of‑sample recidivism prediction by up to 30% compared to pre‑entry covariates alone, and presents a novel instrumental variable estimator that is √N‑consistent for sparse networks with multidimensional latent homophily. By combining limited human annotations with LLM zero‑shot vectors, the authors develop a prediction‑powered peer inference method that yields de‑biased estimates and reveals significant peer effects in behavioral profiles.

By Shanjukta Nath, Jiwon Hong, Jae Ho Chang, Keith Warren, Subhadeep Paul
arXiv AI
Sep 23

HybridFlow: A 2-NFE Generative Policy for Real-Time Robotic Manipulation

HybridFlow is a generative policy for robotic manipulation that uses a three‑stage inference procedure requiring only two network function evaluations (2‑NFE). The policy first generates a coarse action trajectory with a Global Jump based on MeanFlow, then refines the state using a parameter‑free ReNoise interpolation, and finally performs a Local Refine to query the instantaneous‑velocity limit. Experiments on RoboMimic and five real‑robot settings show that HybridFlow achieves high success rates and improves task performance over a 16‑step Diffusion Policy while reducing action‑generation latency by roughly eightfold.

By Zhenchen Dong, Fulin Chen, Jinna Fu, Jiaming Wu, Qingran Wu, Shengyuan Yu, Hongyu Yu, Yide Liu
arXiv Computer Vision
Sep 23

Real-Time Atomic-Resolution Electron Phase Imaging without Probe Calibration via Ptychography-Supervised Learning

The paper introduces a ptychography‑supervised learning framework that transforms 4D‑STEM data into real‑time atomic‑resolution phase images. By training a compact model on physics‑constrained reference phase maps from a single AuPd dataset, the method predicts local phase patches directly from diffraction patterns without probe calibration or iterative optimization. The resulting workflow achieves an online latency of ~0.27 ms per probe position, a 1,000‑fold speed‑up over GPU‑accelerated ePIE, and maintains atomic‑scale lattice contrast while generalizing across materials, defocus conditions, and instruments.

By H. Yue, C. -C. Chen, C. -N. Hsiao, J. Cheng, Y. Liu, X. Z. Liao, Steve F. Shu
arXiv Computer Vision
Sep 23

MIAR: Medical Image Super-Resolution With Autoregressive Modeling

MIAR introduces a multi‑scale autoregressive framework for medical image super‑resolution, treating the task as a conditional, progressive next‑scale prediction. It incorporates a Scale‑Adaptive Structural Decoder to preserve structural fidelity and uses a hierarchical beam search during inference to reduce recursive error accumulation. Experiments show MIAR outperforms existing methods, achieving a 7.86% MUSIQ improvement and a 2.02× speedup over diffusion‑based approaches.

By Fang Li, Yinglong Li, Hongyu Wu, Yang Gao, Minwei Zhao, Aimin Hao
arXiv Computer Vision
Sep 23

StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training

StableVQ introduces practical guidelines to improve training stability for vector‑quantized tokenizers used in image generation models. It addresses instability caused by the entanglement of encoder–decoder and codebook training by proposing three techniques: Dynamic STE for the encoder, Region VQ Loss for the codebook, and a Decoupled Schedule for independent learning rates. Experiments on ImageNet show consistent gains in stability, codebook utilization, and reconstruction quality across various settings.

By Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu, Kun Gai, Xinggang Wang