The paper introduces SwitchSD, an adaptive framework that controls speculative decoding by distinguishing genuine copy intent from accidental repetitions using lightweight probes on a model’s internal representations. SwitchSD dynamically switches between neural drafting and context-based copying, achieving up to 15% throughput gains over existing baselines such as EAGLE3. The approach turns copying from a noisy heuristic into a principled, model-aware decoding regime.
By Roy Eisenstadt, Ido Cohen, Edo Cohen-Karlik, Lior Wolf, Itamar Zimerman
The paper investigates how attention sparsity behaves in autoregressive image generation, finding a distinct diagonal sparsity pattern due to spatial locality of visual tokens. It introduces a diagonal‑aware sparse attention mechanism that skips KV entries along the diagonal within a recent window, achieving up to 3.1× higher throughput and 1.19× lower latency with less than 2% quality loss compared to dense inference.
By Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park
GS-PI introduces an optimization‑decoupled framework that transforms Gaussian Splatting (GS) assets into physically based rendering (PBR) compatible Gaussian assets. By treating PBR material generation as a geometry‑conditioned diffusion process on 3D point clouds, it achieves multi‑view consistency and avoids the pixel‑correspondence problems of 2D diffusion. The method employs a multi‑scale cross‑view conditioning mechanism—combining global semantic priors, photometric cues, and spatial view‑direction signals—to prevent specular highlights from baking into intrinsic colors, and then distills the predicted attributes back into a fully relightable PBR‑GS asset without requiring proxy meshes.
By Jieting Xu, Rengan Xie, Zijian Huang, Zehui Jin, Rui Wang, Yuchi Huo
The paper introduces GRF-Recon, a framework for stable and scalable feed-forward 3D reconstruction from long monocular image sequences. It combines coarse-to-fine trajectory alignment, lightweight geometric prior injection via LoRA adaptation, and a hybrid-weight sparse ray-field optimization to refine local point clouds while enforcing cross-frame consistency. An efficient trajectory stitching strategy with joint ray-error optimization further reduces accumulated drift, achieving competitive trajectory accuracy compared to SLAM systems while maintaining globally consistent reconstructions in large-scale scenarios.
By Enpeng Li, Yunzhou Zhang, Zhiyao Zhang, Dexuan Lyu, Chenyu Wang, Chiyuan Cui, Cheng Cheng
The paper introduces MiX, a micro‑inverted‑scaling format that replaces shared exponents with shared mantissas to avoid microscaling collapse in low‑bit vision‑language models. An adaptive dual‑format inference framework (MiX‑MX) maps this format to a custom accelerator, replacing multipliers with shifters. Experiments show 4.5‑bit MiX matches or outperforms NVFP4 accuracy while improving area efficiency by 25 % and delivering 2.3–4.5× speedup with 1.4–2.9× energy savings over the Focus accelerator.
By Yuan Liao, Jae-sun Seo
The paper introduces a training‑adaptive convolutional sparse coding (CSC) framework that learns the sparsity coefficient jointly with network parameters using an unfolded FISTA optimization. By treating the coefficient as a differentiable variable, the method balances information retention and compression through an information bottleneck perspective, promoting compact yet task‑relevant representations. A label‑free post‑training strategy further adjusts compression for corrupted inputs, yielding competitive accuracy on clean data and enhanced robustness to perturbations on CIFAR and ImageNet.
By Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing
The paper proposes a new unsupervised safety detection method for large language models that relies on anomaly detection rather than supervised training on unsafe data. By leveraging local sparsity in a linear representation space obtained via a sparse autoencoder, the authors develop a framework for locally masked SAE-based anomaly detection, providing theoretical support and empirical validation across multiple architectures and datasets. When calibrated with only 1% out-of-distribution data, the method achieves near‑optimal performance while using just 1–2% of SAE neurons for computation.
By Xin Chen, Gil Kur, Alexander Shevchenko, Andreas Krause
QVAC Genesis III is a 191.43 B‑token synthetic STEM corpus covering 19 domains and multiple difficulty levels, created through a dual generation strategy that uses a weak edge‑scale student model to generate corrective explanations and contrastive reasoning. The authors evaluate the corpus with an LLM‑as‑a‑parser protocol and demonstrate that 1.7 B‑parameter models trained on QVAC Genesis III outperform those trained on Cosmopedia‑v2 and the Cosmo‑1B model on ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% improvement on ARC‑E and a 99.45% valid answer rate.
By Davide Vitabile, N. Ranjan, Akshay Nambiar, Kamal K. Gupta, Amril Nazir
The paper introduces DART, a training‑free technique that combines low‑rank coordinate transport with target‑schedule response calibration to improve the reuse of LoRA adapters in few‑step video diffusion models. By using forward evaluations without source training videos, DART enhances joint quality scores and functional retention on a four‑step Wan2.2 target, with calibration contributing most of the gains. Component analysis shows complementary benefits from coordinate transport, and adapter‑level results indicate both positive functional effects and reduced negative transfer across different adapters.
By Shihong Li, Juntao Xu, JinCao, Maowen Tang, Jun Huang, Jintao Li
The paper introduces PreDE, a policy‑calibrated framework that predicts how post‑training quantization will degrade task performance in world action models (WAMs) before deployment. By calibrating two thresholds on a small development set, PreDE can accept, reject, or defer new quantization configurations based on offline action deviations, achieving 75% coverage of decisions that match closed‑loop outcomes. Experiments on five WAMs and real‑robot trials show that PreDE accurately identifies high‑deviation configurations and enables significant speedups and memory reductions without compromising performance.
By Jiuyi Xu, Jinjia Guo, Meida Chen, Jing Du, Yangming Shi
The paper demonstrates that the number of candidates generated during test-time scaling of large language models does not fully capture the system cost. By comparing different generation schedules (e.g., one batched call versus multiple serial calls) while keeping the total candidate count fixed, the authors show that serial calls consume significantly more GPU energy and latency. The study suggests that reporting candidate count alone is insufficient; evaluations should also include generation schedule and GPU-level metrics.
By Mobina Kashaniyan, Ali Jannesari
The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.
By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao
The paper investigates why on‑policy distillation (OPD) can produce excessively long student responses, attributing this to a termination‑token mismatch between base students and post‑trained teachers. Experiments on Qwen3, Llama, and Gemma show that differing stopping probabilities for identical EOS token sets suppress the student’s preferred termination and fail to transfer the teacher’s alternative. Aligning decoding stopping sets alone is insufficient; treating functionally equivalent EOS tokens as a shared semantic stopping action reduces length inflation, though late‑stage inflation persists beyond termination alignment.
By Yuxiao Yang, Tianrun Yu, Shangzhe Li, Kaixiang Zhao, Xuchao Zhang, Chetan Bansal, Huaxiu Yao, Taylor W. Killian, Weitong Zhang
PrefixBench-H100 is a reproducible benchmark that evaluates how reusing prompt prefixes affects LLM serving performance on NVIDIA H100 GPUs. It tests two popular runtimes (vLLM and TensorRT-LLM) across varied workloads, measuring metrics such as time-to-first-token, latency, throughput, cache hits, and GPU memory usage. The study identifies when prefix reuse significantly reduces first‑token latency and when cache pressure diminishes those gains, noting that cache effectiveness is largely unaffected by concurrency or output length, while differences arise mainly in scheduling.
By Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav
NS3Learn is a closed‑form model that captures realistic 5G NR sidelink Mode‑2 reception losses—such as half‑duplex loss, scheduling collisions, receiver capture, and decoding—by fitting 10.5 million labeled outcomes from ns‑3 5G‑LENA traces. The model achieves a mean absolute deviation of 0.06 in per‑instant delivery compared to ns‑3, outperforming alternative models, and its parameters transfer with minimal error to new intersections. Using NS3Learn in traffic‑network simulations reverses traffic speed trends and more than doubles predicted hard‑braking events, demonstrating its impact on safety assessments.
By Rasheed Bello, Arthur Mukwaya, Gurcan Comert, Varghese Vaidyan, Vijay Bendigeri, Anthony Dontoh, Jagruti Sahoo, Judith Mwakalonge
QCPruner is a training‑free visual token pruning method that conditions both token selection and representation on the query via bilateral utility weighting. It fuses keyword‑matched query anchors with cross‑modal cues to compute a nonnegative facility‑location objective that is monotone and submodular, guaranteeing a (1‑1/e) greedy approximation. Across multiple multimodal large language models, QCPruner consistently outperforms existing pruning methods, achieving over 96% of unpruned performance even with very few tokens retained.
By Shengli He (Guizhou University), Yongchao Liang (Guizhou University), Roumeng He (Shanghai Ocean University), Junjie Zeng (Guizhou University), Jiyuan He (Guizhou University), Can Wu (Guizhou University), Li Zheng (Guizhou University)
The paper introduces PhGS, a post‑hoc pruning and refinement pipeline for single‑view feed‑forward 3D Gaussian Splatting models. It keeps the base network frozen and applies importance‑score‑based pruning followed by a lightweight recurrent refinement module to reduce spatial redundancy while maintaining rendering quality. The method is backbone‑agnostic, integrates seamlessly with existing baselines, and allows flexible inference‑time keep ratios for different application needs.
By Rinto Yagawa, Han Cheng, Dieter Schmalstieg, Hideo Saito, Shohei Mori
The paper introduces Latent-Centroid Steering (LCS), a single-pass classifier-free guidance method for vision‑language autonomous driving models. LCS replaces instance‑level residuals with class‑level latent shifts, projecting conditional representations toward precomputed command‑specific centroids to enhance command adherence. Experiments on Bench2Drive and nuScenes show that LCS cuts inference latency by about 50% while improving driving performance.
By Meibo Hu, Jiamian Wang, Pichao Wang, Zhiqiang Tao
REACT is a fully spiking state‑space model that processes raw event‑camera data one event at a time, avoiding temporal accumulation and its associated delay. It employs a complex‑valued spiking neuron (C‑SiLIF) whose dynamics are driven by the inter‑event interval, enabling continuous‑time state updates at microsecond resolution. Evaluated on gesture recognition and time‑to‑collision estimation, REACT achieves low latency (4.6 ms) and high accuracy, supports anytime prediction, zero‑shot transfer, and INT8 quantization, dramatically reducing energy consumption.
By Geoffroy Keime, Nicolas Cuperlier, Benoit R. Cottereau
The paper introduces truncated automatic sparse differentiation (ASD) to efficiently compute higher‑order derivatives, such as Hessians, for machine learning interatomic potentials (MLIPs). By exploiting the locality of atomic interactions, ASD identifies a sparsity pattern that allows exact Hessian calculation for large porous materials, while truncated ASD discards distant, small Hessian entries to achieve order‑of‑magnitude speedups with minimal loss in predictive accuracy. The authors demonstrate these methods on several foundational MLIPs, showing modest speedups for full ASD and significant gains for the truncated approach.
By Marcel F. Langer, Adrian Hill, Michele Ceriotti