arXiv:2605.10194v2 Announce Type: replace
Abstract: On-policy self-distillation (OPSD) uses a model as its own teacher under privileged context, providing token-level supervision on the model's own r...
By Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen, Yulan Hu, Lan-Zhe Guo
arXiv:2609.32259v2 Announce Type: replace
Abstract: Recent multi-agent LLM systems increasingly combine heterogeneous models for specialized agent roles. However, text-based communication requires ea...
By Vincent-Daniel Yun, Woosang Lim, Haneul Yoo, Sungjoo Yoo, Murali Annavaram, Sai Praneeth Karimireddy
arXiv:2609.32540v2 Announce Type: replace-cross
Abstract: Few-step autoregressive video diffusion generates a long video by splitting the video into temporal chunks and generating chunk-by-chunk, eac...
By Yikai Wang, Xiao Han, Mengmeng Xu, Juan Camilo Perez, Yiannis Douratsos, Sen He, Zijian Zhou, Fei Zhang, Zhaochong An, Juan-Manuel Perez-Rua, Chen Change Loy, Tao Xiang
arXiv:2609.38630v1 Announce Type: new
Abstract: Privacy redaction must remove personal information while preserving relationships expressed in text. We develop a multilingual named-entity tagger with...
By Jonathan Graehl
arXiv:2609.39111v1 Announce Type: new
Abstract: Human intelligence relies heavily on learned intuition: recognising patterns and judging situations without explicitly unfolding every intermediate ste...
By Li Ding, Haidi Jin, Chen Ji
The paper investigates on‑policy self‑distillation (OPSD) as a method to enhance reasoning in language models, focusing on mathematical reasoning across models from 0.6B to 8B parameters. Through controlled experiments and token‑level analysis, the authors find that OPSD’s effectiveness depends on alignment between the teacher’s reasoning mode and the full teacher prefix, rather than on privileged semantics alone. They observe that OPSD only improves reasoning in limited compatibility regimes, while often causing length growth, degradation, or behavioral collapse, and that the teacher’s signal is unstable and not predictive of downstream performance.
By Yang Li, Gongle Xue, Yuheng Yuan, Yijia Guo, Shizhe Zhang, Liwen Hu, Lei Ma
arXiv:2609.39346v1 Announce Type: new
Abstract: Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (S...
By Bohan Zhang (Southeast University), Linan Yue (Southeast University), Weibo Gao (Hong Kong Polytechnic University), Pengyu Chen (Southeast University), Hong Guo (Southeast University), Yanqi Hao (ZTE Corporation)
arXiv:2609.38658v1 Announce Type: cross
Abstract: TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but the...
By Jian Chen, You Zhang, Mark Vinton
arXiv:2609.39334v1 Announce Type: cross
Abstract: Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, subs...
By Jinwoo Jeong (Korea University), Woohyung Choi (Korea University), Myeongjae Jeon (POSTECH), Jeongseob Ahn (Korea University)
arXiv:2609.39920v1 Announce Type: cross
Abstract: Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as mod...
By Yanshu Li, Jiaqian Li, Canran Xiao, Xi Xiao, Tianyang Wang, Yongtai Liu
LatentHarness unifies memory access and latent reasoning by treating them as sequential latent actions—THINK, RECALL, and EXIT—within a language model. It is trained via counterfactual policy distillation, which evaluates the impact of each action on the emitted token and learns when to recall evidence versus continue reasoning. On six long‑context reasoning benchmarks, a 1.4B‑parameter LatentHarness model outperforms the strongest baselines by 2.8% and 10.0% relative, while running 5.9× faster than the leading long‑context baseline.
By Xiaoqiang Wang, Suyuchen Wang, Bang Liu
BELIEFRAG is a closed‑loop controller that maintains an explicit evidence state—tracking sufficiency, reliability, conflict, uncertainty, evidence gaps, and acquisition cost—to guide adaptive retrieval‑augmented generation (RAG). By updating this state, it selects among retrieval, query rewriting, verification, answering, stopping, and abstention, achieving higher token‑F1 scores and lower token usage on six QA benchmarks compared to fixed iterative retrieval. The approach shows that corrective re‑retrieval and calibrated answerability are key to its performance gains.
By Hongji Pu
OP-CAD introduces a curriculum-based, on-policy clean-audio distillation framework that enhances audio-visual reasoning under environmental noise and competing speech. The method trains a student model from mild to severe noise, using a frozen teacher that provides token-level supervision based on clean audio and verified answers, while selectively weighting positions sensitive to acoustic interference. Experiments show OP‑CAD outperforms existing methods across all noise conditions, preserving clean‑correct answers without sacrificing overall accuracy.
By Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang
Switching Linear Attention (SwiLA) is a new sequence layer that improves upon standard softmax attention by maintaining a fixed-size recurrent state while enhancing representational capacity. It derives its recurrence from a test-time regression framework, using online expectation-maximization in a mixture of linear regressions model. In various benchmarks—including associative recall, in-context language learning, and language modeling—SwiLA achieves strong performance, narrowing the gap to softmax attention and even surpassing it in some settings.
By Hyun Dong Lee, Xavier Gonzalez, Nicolas Zucchet, E. Kelly Buchanan, Emily B. Fox, Scott W. Linderman
The paper introduces SCOUT, a co‑training framework that adapts an off‑policy teacher to better continue from student‑generated prefixes in on‑policy distillation (OPD). By periodically optimizing the teacher’s conditional continuation ability using reinforcement learning with verifiable rewards, SCOUT improves the teacher’s performance on student prefixes. Experiments across various teacher‑student setups, model scales, and reasoning domains show that SCOUT consistently enhances the effectiveness of OPD.
By Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Bl\"obaum, Purak Jain
DiFF is a generative framework that uses Doppler velocity cues from 4D millimeter-wave radar to improve human motion flow estimation. It combines Doppler-informed motion priors with a Kolmogorov‑Arnold Network (KAN) based conditional flow matching model, featuring a KAN‑attention mechanism for expressive feature extraction. Experiments demonstrate that DiFF achieves state‑of‑the‑art performance, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
By Kai Wang, Mingle Zhao
UniEvo‑VL is a self‑evolving framework that lets multimodal models improve themselves by using their own critiques as privileged information. The method trains a single model to act as both teacher and student, minimizing divergence between their diffusion distributions over sampling trajectories. Experiments on Qwen‑image‑2512 show significant gains in image generation metrics, and stronger external critics further raise the improvement ceiling, though results vary across tasks.
By Fang Wu, Da Xing, Yanjie Huang, Junxi Wang, Ji Wang, Hejia Geng, Guancheng Wan, Bowen Zuo, Xiaomin Li, Shixiang Tang, Xinyu Xiang, Zehong Wang, Shiyi Du, Peng Xia, Shuangjia Zheng, Yining Hong, Li Erran Li, Jure Leskovec, Yejin Choi
The paper treats multi‑path large language model reasoning as a diversity‑combining problem, analogous to noisy channel observations in wireless communications. It shows that the optimal symmetric linear combiner of latent embeddings is uniform, justifying majority vote in standard self‑consistency while allowing for weighting or pruning when prompt‑template branches are heterogeneous. By reducing path correlation through prompt‑template diversity, the authors propose an Adaptive‑K rule that selects an optimal number of reasoning paths, preserving most of the accuracy achieved with a fixed large number of paths across multiple models and benchmarks.
By Guangsheng Yu, Litianyi Zhang, Qin Wang, Xu Wang, Mingyuan Li, Shaoxiong Ji, Ren Ping Liu, Massimo Piccardi
The paper investigates self‑distillation techniques for language models by systematically varying three key design choices: the source of rollout tokens (student vs. teacher), the teacher coupling strategy (frozen or exponential moving average), and the KL divergence direction (reverse or forward). Experiments on Qwen2.5‑7B and Ministral‑3‑3B across 1,200 adaptation runs reveal that rollout source mainly affects acquisition on contradictory tasks, teacher coupling most strongly influences acquisition across all tasks, and KL direction impacts retention differently depending on the model. A controlled theoretical model reproduces these empirical trends, offering a unified framework for understanding acquisition‑retention trade‑offs in self‑distillation.
By Luis Zuin, Alexis Huet, Dario Rossi, Zied Ben Houidi
arXiv:2609. 40285v1 Announce Type: new Abstract: On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories.
By Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh