The paper introduces TISD, a trajectory-intervention self-distillation method that forces a teacher-selected branch action and then lets the student generate the suffix, distilling the full trajectory under a privileged-context-conditioned teacher. This approach addresses a data-collection bottleneck in on‑policy self‑distillation by exposing successor contexts that the student would otherwise miss. Experiments on coding and science domains show modest but consistent improvements in average performance metrics compared to baseline methods.
By Taeckyung Lee, Rinat Amankos, Jeonghye Kim, Hyungjun Yoon, Woogyeol Jin, Sung-Ju Lee
The paper introduces a momentum-guided federated split distillation framework for personalized temporal edge intelligence. It presents TeRR-SAtt, a temporal reservoir student attention design that uses fixed reservoir representations, a lightweight temporal student, and personalized output modules. Additionally, it proposes AMGF, an anticipatory momentum-guided fusion mechanism that clusters clients via learning momentum and generates specialized teacher updates. Experiments on real-world smart‑building data show that TeRR-SAtt cuts edge training latency by 65.50%, inference latency by 44.70%, training memory usage by 18.40%, and inference CPU usage by 33.10% compared to baselines, while AMGF improves local learning by up to 35.31% in RMSE over global updates.
By Ahmed-Rafik Baahmed (LINEACT), Jean-Fran\c{c}ois Dollinger (LINEACT), Amine Brahmia (LINEACT), Mourad Zghal (LINEACT)
G2MAF is a test‑time refinement framework for offline multi‑agent reinforcement learning that applies a single globally normalized, projected critic gradient to adjust all agents’ actions while keeping them close to a frozen policy proposal. The method improves performance on 24 Multi‑Party Environment (MPE) and StarCraft Multi‑Agent Challenge (SMAC) benchmarks, achieving mean relative gains of 9.2% on MPE and 8.9% on SMAC, with only a 6% increase in inference latency.
By Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu
The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.
By Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst
The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.
By Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)
The paper investigates prompt minimization, aiming to reduce prompts to their smallest, most information-dense form without losing output fidelity. It argues that shorter prompts lower computational overhead and inference latency, especially when large contexts are unnecessarily included, and that longer prompts can harm LLM reasoning and accuracy. The authors propose three frameworks to identify minimal prompts and show that these often produce outputs comparable to longer versions, highlighting redundancy in the input space and opening new avenues for efficient prompt engineering.
By Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu
DeepEdu‑v1 is an AI‑tutoring system tailored for Vietnamese education that addresses data‑sovereignty and local curriculum alignment issues. It uses a long‑context inference engine to reduce retrieval calls and prefill latency by about 35%, and a self‑improving agentic layer that curates verified local knowledge without fine‑tuning. In deployment, DeepEdu achieves nearly twice the speed of standard vLLM serving and raises agentic accuracy from 70.0% to 79.5% on complex tasks, especially in financial reasoning and interactive‑agent benchmarks.
By Quang Nguyen, Hieu Nguyen, Hien Hoang, Toan Pham, Cong Tran, Nam Vu
The paper introduces an adaptive multi‑resolution Gaussian process framework that achieves scalable, exact inference by constructing a naturally data‑sparse covariance matrix using basis functions anchored directly to samples. By shrinking the support domains of these basis functions, the resulting matrix has limited block sizes, ensuring sparsity and enabling efficient computation of its inverse via a sparse Cholesky algorithm. The authors demonstrate that this approach yields exact inference with training cost ≠≠ O(n log^2 n) and prediction cost ≠≠ O(log^d n), while also improving predictive uncertainties through an augmented basis function.
By Yanchuang Cao, Jun Liu, Tengchao Yu, Heng Yong
The paper investigates whether causal softmax attention can realize policy mirror descent (PMD) as a repeated controller rather than a one‑step algebraic identity. It constructs a fixed causal‑softmax actor–environment–one‑step‑critic protocol, detailing actor, routing, sampling, and normalization residuals, and shows that a frozen one‑step audit model closely approximates PMD. Empirical results demonstrate that the learned actor with an exact one‑step critic achieves median policy loss only about 5% higher than the exact PMD oracle across multiple control settings.
By Yuhe Sui, Yingzhi Tang, Shufang Chen
MedTokenBudget introduces a supervised token routing framework for Vision Transformers applied to dermoscopic image classification. Its Lesion-Aware Token Scoring (LATS) module combines attention entropy, feature norm, and local feature contrast to select the top‑K patches under a target budget, trained with curriculum learning, diversity regularization, attention distillation, and lesion‑mask supervision. On the ISIC 2019 dataset, mask‑supervised LATS outperforms Random and ToMe at headline budgets while retaining more lesion patches.
By Zhexiang Li
MOPD‑Router rethinks teacher routing in multi‑teacher on‑policy distillation by routing supervision over the full teacher pool at each token, eliminating the need for prompt‑level domain labels or a separate routing model. The framework offers a plug‑in interface for various metrics, and introduces ExpertAlign, which scores teachers based on how well their corrections reflect their specialized post‑training knowledge. Experiments on both unlabeled and domain‑labeled mixtures show that ExpertAlign outperforms existing methods, improving overall scores by up to 12.3% on unlabeled data and 7.8% on domain‑labeled data.
whyItMatters":"Token‑level routing enables the use of complementary supervision across domains without relying on domain labels, leading to significant performance gains in multi‑teacher distillation settings."
By Tianze Xu, Yanzhao Zheng, Zhentao Zhang, Yuanqiang Yu, Chao Ma, Jihuai Zhu, Lelun Wu, Lyumanshan Ye, Pengfei Liu, Baohua Dong, Hangcheng Zhu, Ruohui Huang, Gang Yu
The paper introduces Persistent Negative Adversarial Distillation, a method that improves black-box on-policy distillation by maintaining a live pool of historical teacher–student comparisons to stabilize the discriminator’s negative distribution. By anchoring the discriminator with these persistent negatives, the approach reduces reward-estimation error and yields smoother, higher-performing student policies across multiple benchmarks. The study demonstrates that the choice of negative samples is a critical design factor in effective black-box distillation.
By Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang
G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.
By Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou
The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.
By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
The paper introduces acoustic-to-text KV compression for full‑duplex speech models, converting acoustic key‑value states into compact textual memory during listening‑time slack. When the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and recent acoustic context are retained. Experiments on ten‑minute LongSpeech sessions show a 64.6% reduction in peak streaming KV‑cache size and improved transcription, temporal question answering, and summarization, with comparable pause‑handling, turn‑taking, and interruption performance in Full‑Duplex‑Bench.
By Yejin Lee, Seungbeom Kim, Yongha Lee, Kyuhong Shim
The paper introduces a post‑training softmax reparameterization technique that selects a functionally equivalent output head before quantization. By subtracting a scalar multiple of the vocabulary‑row mean from each output row and tuning this coefficient via validation KL, the method preserves the full‑precision softmax distribution while enabling efficient W4 quantization. Experiments on seven heads show significant error reductions and latency improvements, with the approach remaining complementary to other quantization strategies and transferable across datasets.
By Asim Kadav, Christian Flores, Chirag Arora, Varun Kotte, Hongbo Zheng, Lan Yan, Priya Shanmugasundaram, Tracy Holloway King
UniAR is a unified framework that improves autism spectrum disorder (ASD) recognition by using multi-granularity prompt learning and a large multimodal model to generate diagnostic descriptions at word, phrase, and sentence levels. It aligns these semantic representations with visual evidence through a Mixture-of-Experts-based Multi-Scale Alignment Module, enabling robust ASD detection across heterogeneous data types. Experiments on four brain MRI and facial expression benchmarks show that UniAR outperforms state‑of‑the‑art methods, achieving 75.9% accuracy on MRI and 91.6% on facial benchmarks, with gains of 1.5 and 1.2 percentage points respectively.
By Lei Xin, Zeheng Wang, Jiayin Zhu, Shihong Huang, Fanhu Zeng, Changjiang Jiang, Dengbo He, Yutao Yue, Zhenglun Kong
DyMD introduces a Distribution Matching Distillation framework that adapts teacher supervision and critic fitting to preserve interaction dynamics in few-step video generation. By employing temporal affinity–conditioned re‑noise sampling and dynamics‑guided fake‑score tracking, DyMD balances motion recovery with visual quality. The method distills a 14B teacher into a 1.3B student that achieves significant gains on embodied‑video benchmarks and downstream action planning tasks.
By Haojun Xu, Jie Huang, Xin Lu, Mingchen Zhong, Zihao Fan, Linjiang Huang, Si Liu
Intent2Tc is a closed‑loop, language‑model‑driven framework that translates high‑level business traffic‑shaping intents into executable Linux traffic‑control (tc) configurations. It uses an AQM‑based digital twin semantic model, automated metadata extraction, critique‑driven refinement, and Retrieval‑Augmented Generation to improve semantic consistency and configuration reliability. Evaluation on 100 RFC 9315‑compliant intents shows high semantic fidelity and deployment readiness, with Claude Sonnet‑4.6 achieving 0.98 semantic similarity and 0.045 normalized edit distance, while RAG reduces token consumption and latency for compact models.
By Andrea Masini, Sudipta Acharya, Paolo Bellavista, Luca Foschini, Burak Kantarci
The paper introduces Stepwise Marginal Information Gain (MIG), an intrinsic process reward that evaluates how each reasoning step of a large language model (LLM) or vision-language model (VLM) improves the likelihood of the reference answer. MIG rewards only new likelihood maxima, preventing duplicate credit, and is combined with outcome, format, and self‑distillation objectives to guide training. Experiments on eight benchmarks show that this method outperforms outcome‑only reinforcement learning and improves accuracy by up to 4.8 points over binary‑reward training, including a 12.6‑point gain on MathVerse and a 12.9‑point advantage on vision‑language transfer at 7B parameters.
By Xiangwei Wang, Wei Wang, Ken Chen, Nanduni Nimalsiri, Sachith Seneviratne, Saman Halgamuge