Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,633 stories · RSS feed

arXiv AI
2d ago

Four Ways to Grow a Classifier and Why One of Them Cannot Learn

The paper investigates four ways to grow a classifier—adding a tree level, a hidden unit, a leaf split, and a statistically significant split—under a fixed protocol for tree‑structured and constructive models. It shows that the most natural method of deepening a soft decision tree by duplicating a leaf’s class distribution leaves the gradient of new gates identically zero, preventing learning, and proposes a small random perturbation as a fix. The other three growth decisions each provide a distinct benefit: fitting a new hidden unit to residual error yields a smaller network, splitting the leaf with the largest expected error adds sparsity, and requiring statistical significance before splitting adds no value and reduces accuracy.

By Cagri Temel
arXiv AI
2d ago

DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies

DriftOPD is a teacher‑free, rollout‑free framework that performs sequence‑level on‑policy distillation of continuous Vision‑Language‑Action (VLA) action experts. It decomposes the sequence‑level reverse‑KL divergence into a chunk‑level reverse‑KL term and a future‑potential term, optimizing them with a one‑step drifting objective and a Q‑function critic learned from offline demonstrations. Experiments on multiple VLA architectures in simulation and real‑world manipulation show that DriftOPD outperforms existing one‑step distillation baselines while matching the task success of multi‑step teacher policies.

By Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye
arXiv Computer Vision
2d ago

MEGA: Object-Level Mesh Extraction from 3D Gaussian Splatting via Spatial Visual Distillation

MEGA is a new framework that extracts object-level, watertight meshes from 3D Gaussian Splatting (3DGS) scenes. It uses a segment-then-mesh approach, leveraging Spatial Visual Distillation (SVD) to sample diverse camera views of each segmented object and train a mesh reconstruction model with photometric supervision. Experiments on popular benchmarks show that MEGA outperforms existing methods in accurately recovering object-level 3D occupancy and supports complex physical interactions by combining high-quality meshes with photorealistic 3DGS rendering.

By Liwei Liao, Yingkui Zhang, Qianqian Tong, Ronggang Wang
arXiv Computer Vision
2d ago

Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation

Omni-Embed-Mini is a 0.9B‑parameter model that embeds text, speech, audio, images, video, and visually‑rich documents into a single shared cosine space without updating any text‑side parameters. It uses a dense cascaded caption as a teacher signal, allowing the teacher and student to share identical backbone weights and requiring only lightweight projectors and phased LoRA adapters for alignment. The model achieves strong text retrieval performance (49.57 nDCG@10 on MTEB‑v2 BEIR‑8) while extending to five additional modalities and is significantly smaller than other open omni‑modal embedders.

By Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
arXiv AI
2d ago

Spatial Strategies, Not Actions: Vector-Quantized Geodesics as Tools for LLM-Driven Agents

The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.

By Gabriel Turinici
arXiv AI
2d ago

Fold'EM: Direct atomic structure inference from Cryo-EM particles

Fold'EM is an inference-time framework that directly combines protein generative model priors with cryo‑EM particle images to recover atomic structures, bypassing the traditional density reconstruction step. It can determine structures from a small number of particles, jointly infer orientations in an ab‑initio setting, and resolve distinct conformational states in heterogeneous samples without separate map reconstruction. The method demonstrates accurate atomic models on both synthetic and experimental datasets, showing promise for low‑sample and low‑population conformational analysis.

By Advaith Maddipatla, M\"art-Erik M\"aeots, Marco Pegoraro, Nikolaus Dr\"ager, Roberto Covino, Sanketh Vedula, Martin Pacesa, Alex M. Bronstein
arXiv Machine Learning
2d ago

Characterizing High Bandwidth Flash for LLM Serving

arXiv:2609.39131v1 Announce Type: new Abstract: Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer,...

By Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
arXiv Machine Learning
2d ago

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...

By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
arXiv Machine Learning
2d ago

No Task Vector Is an Island: A Comprehensive Study on the Composability of Task Vectors from On-Policy Distillation

arXiv:2609.39405v1 Announce Type: new Abstract: Task vectors provide a simple mechanism for composing learned capabilities through model merging. However, the composability of task vectors produced b...

By Jingang Zhou, Feiyu Han, Han Zhu, Yuyi Zhou, Ruiyang Zhang, Jian Xu, Sirui Gao, Qingpei Guo, Xu-Yao Zhang