arXiv:2602.15396v2 Announce Type: replace
Abstract: Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces in...
By Jeongwoo Shin, Jinhwan Sul, Joonseok Lee, Jaewong Choi, Jaemoo Choi
arXiv:2606.21562v2 Announce Type: replace
Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
By Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz, Guillaume Bono, Gianluca Monaci
arXiv:2512.04705v3 Announce Type: replace-cross
Abstract: The deployment of Early-Exiting Neural Networks (EENNs) on edge accelerators requires optimizing not only the network architecture but also i...
By Alaa Zniber, Arne Symons, Ouassim Karrakchou, Marian Verhelst, Mounir Ghogho
DecomVoxel introduces a guided in‑situ denoising optimization that fuses 3D‑native priors with neural scene reconstruction to improve decompositional scene reconstruction. The method employs an epsilon‑based distillation loss for stable latent refinement and adaptive spatial guidance using occupied and vacant anchors with temporal annealing to reduce hallucinations and spatial drift. Experiments on Replica and ScanNet++ demonstrate that DecomVoxel outperforms state‑of‑the‑art approaches while preserving spatial layout, structural fidelity, and style‑consistent texture, yielding high‑quality textured meshes with clean topology.
By Junfeng Ni, Zirui Zhou, Yixin Chen, Yu Liu, Nan Jiang, Zhifei Yang, Song-Chun Zhu, Siyuan Huang
Dyna‑DINO introduces a curriculum for Vision Transformer (ViT) knowledge distillation that uses the teacher’s intermediate feature maps as progressively harder targets, enabling a student to build foundational representations before tackling higher‑level abstractions. The approach accelerates convergence and improves performance across multiple tasks: on ImageNet‑100 the distilled ViT‑S reaches 90.1% accuracy (+12.24% over baseline), while on ImageNet‑1K it yields +3.9% and +6.09% gains on Oxford and Paris retrieval, +1.93% on semantic segmentation, and notable classification improvements. Additionally, the curriculum reduces training FLOPs by 25.1% and training time by 21% on ImageNet‑100 through early‑stopping of teacher inference.
By Jiaqi Zhang, Ashton Lee, Anthony Wong, John Zou, Sami BuGhanem, Randall Balestriero
RISED introduces a framework that uses rubric-based textual feedback to improve training of a single large language model (LLM) agent across multiple interactive environments. By having an LLM judge tag rollouts with a shared rubric vocabulary, the system guides both online data selection and policy supervision, enabling richer cross‑environment relationships and within‑group reward contrast. Experiments show that RISED achieves the highest mean pass rate and ranks first or second in every individual environment, with rubric analysis revealing behavioural changes behind these gains.
By Jingtan Wang, Sirajul Salekin, Young mok Jung, Javier Movellan, Bryan Kian Hsiang Low, Manjot Bilkhu
ShatterQuant is a hardware-software co-designed framework that enables mixed-precision quantization within individual tensors by assigning different bit-widths to blocks of a weight projection. It couples precision granularity with processing element configuration, allowing each precision to determine an effective block height. The framework includes a hardware-aware post-training method based on block-level standard deviation and weight sensitivity, a ShatterQuant Transformer Accelerator supporting 1/2/4/8-bit weight precision, precision-dependent PE configuration, block rescaling, and integrated softmax and piecewise-linear nonlinearities, and an evaluation showing 1.5 TOPS, 760 GOPS/$mm^2$ area efficiency, and 2.8 TOPS/W energy efficiency on a TSMC 16nm PDK implementation.
By Mikolaj Walczak, Edward Humes, Chao Fang, Marian Verhelst, Tinoosh Mohsenin
The paper introduces a new approach to distill reasoning abilities from large language models (LLMs) into smaller student models by framing the task as a constrained reinforcement learning problem. It enforces a worst‑case constraint on the teacher’s log‑likelihood for every prefix of the reasoning chain, avoiding reward hacking and excessive teacher regularization. Experiments on mathematical reasoning and code generation show that this method improves the balance between accuracy and fidelity, achieving the highest rigorous reasoning success rate among evaluated settings.
By Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar
IrekoGPT is a post‑hoc technique that transforms pretrained large language models into slimmable versions, enabling dynamic width adjustment during inference. It builds on SliceGPT by keeping the original projection matrices intact, thereby exposing nested subnetworks at various widths. The method enhances robustness through layer‑wise calibration across multiple compression ratios and refines downstream linear layers using gradient‑free ridge regression, yielding better performance than naive PCA‑based slimming on Llama and Qwen models, especially at high compression levels.
By Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara
arXiv:2610.00432v1 Announce Type: cross
Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponential...
By Xiaofan Que, Nir Elkayam, Spandan Pyakurel, Shuokai Pan, Dibakar Gope
arXiv:2610.00838v1 Announce Type: cross
Abstract: Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error...
By Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard
arXiv:2610.00888v1 Announce Type: cross
Abstract: Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per fo...
By Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song
The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.
By Jundong Hu, Shekar Ramachandran
arXiv:2610.00812v1 Announce Type: cross
Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamic...
By Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli
arXiv:2610.00899v1 Announce Type: cross
Abstract: Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with st...
By Keisuke Shirai, Tomohiro Motoda, Hanbit Oh, Ryoichi Nakajo, Roman Mykhailyshyn, Ryo Hanai, Shotaro Miwa, Yukiyasu Domae
arXiv:2610.00997v1 Announce Type: cross
Abstract: Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may...
By Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim
arXiv:2610.01408v1 Announce Type: new
Abstract: Trajectory crossing remains a critical bottleneck in Flow Matching (FM), and previous works typically view these crossings from a theoretical optimizat...
By Ziqi Jiang, Zhenqi He, Long Chen
arXiv:2610.02117v1 Announce Type: cross
Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a froze...
By Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc, Nicolas Thome, Spyros Gidaris
arXiv:2506.11030v2 Announce Type: replace-cross
Abstract: Training neural networks has traditionally relied on backpropagation (BP), a gradient-based algorithm that, despite its widespread success, s...
By Nazmus Saadat As-Saquib, A N M Nafiz Abeer, Hung-Ta Chien, Byung-Jun Yoon, Suhas Kumar, Su-in Yi
arXiv:2604.14908v2 Announce Type: replace-cross
Abstract: We study downlink beam and rate adaptation in a multi-user mmWave MISO system where multiple base stations (BSs), each using analog beamformi...
By Emre \"Ozy{\i}ld{\i}r{\i}m, Bar{\i}\c{s} Yayc{\i}, Umut Eren Akturk, Cem Tekin