arXiv:2609.38881v1 Announce Type: new
Abstract: Real-time strategy (RTS) games require agents to coordinate economic development, production and construction, base defense, unit organization, and att...
By Xinhe Tian, Xiaoyue Zhang, Ziyou Zhang, Jiacheng Li, Xiaoqiang Jin, Qianchuan Zhao, Gaochen Cui
arXiv:2609.39306v1 Announce Type: cross
Abstract: Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our...
By Shengjie Jin, Hengbo Xu, Zelong Sun, YuJie Guo, Zhiwu Lu
arXiv:2609.39704v1 Announce Type: cross
Abstract: This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without lab...
By Muhammad Zawish, Steven Davy
arXiv:2608.06912v2 Announce Type: replace
Abstract: Selecting the top-$k$ elements is a fundamental operation for inducing sparsity in large-scale models and optimization problems, enabling robust ex...
By Jakub Antczak, Joanna Wojciechowicz, Kamil Ksi\k{a}\.zek, Marcin Mazur, {\L}ukasz Struski, Jacek Tabor
arXiv:2509.23235v3 Announce Type: replace-cross
Abstract: Model inversion is a widely adopted technique in data-free learning that reconstructs synthetic inputs from a pretrained model through iterat...
By Seongsoo Heo, Dong-Wan Choi
arXiv:2603.06350v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constrai...
By Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang
arXiv:2605.28868v2 Announce Type: replace-cross
Abstract: Metagenomic taxonomic annotation is essential for interpreting complex microbial communities, yet reliable annotation remains challenging und...
By Rongye Ye, Lun Li, Zheng Luo, Yiran Zhan, Zhang Zhang, Shuhui Song
arXiv:2609.38812v1 Announce Type: new
Abstract: Terminal agents rely on self-verification to assess and correct their solutions as they solve tasks through interaction with command-line environments....
By Yingfeng Luo, Shaowei Wei, Daixin Wang, Dingyang Lin, Kaiyan Chang, Weiqiao Shan, Tong Zheng, Zhiqiang Zhang, Jingbo Zhu, Tong Xiao
arXiv:2609.38995v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains a student on its own generated responses using feedback from the same model conditioned on privileged informa...
By Di Huang, Hao Li, Yixin Chen, Fuhai Li
arXiv:2609.38325v1 Announce Type: new
Abstract: We introduce modal kinetic typography, which animates a vector glyph to express a semantic concept while keeping it legible. Our key idea is to build m...
By Maham Tanveer, Jiyeon Han, Nanxuan Zhao, Hao Zhang
arXiv:2609.38777v1 Announce Type: new
Abstract: A central goal of vision-language model (VLM) distillation is to transfer both the teacher's language capabilities and its visual understanding. Howeve...
By Yuanhao Sun, Huawei Ji, Jiaxin Ding, Luoyi Fu, Xinbing Wang
arXiv:2609.39096v1 Announce Type: new
Abstract: Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compr...
By Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding, Xingang Pan
arXiv:2609.40037v1 Announce Type: new
Abstract: Few-step autoregressive video generation enables efficient streaming synthesis, but errors introduced in early temporal blocks are reused as context an...
By Fangyu Lin, Xingtong Ge, Lunjie Zhu, Yi Zhang, Zhening Liu, Tianhang Wang, Mengfei Li, Yumeng Zhang, Guanglu Song, Yu Liu, Jun Zhang
arXiv:2609.40362v1 Announce Type: new
Abstract: We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and q...
By Hongyuan Tao, Xinggang Wang, Lianghui Zhu, Yongkang Li, Yunchao Wei, Bin Feng, Shaoyu Chen, Qian Zhang, Chang Huang, Kai Yu
arXiv:2609.38488v1 Announce Type: cross
Abstract: Conventional 3D Gaussian Splatting (3DGS) requires depth sorting and ordered alpha blending to correctly render overlapping Gaussian primitives. Stoc...
By Zijian Huang, Suiliang Mai, Chuankun Zheng, Yuan Meng, Yuchi Huo
arXiv:2609.39757v1 Announce Type: cross
Abstract: Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs int...
By Xiao Cui, Mo Zhu, Yulei Qin, Yuze Wu, Wengang Zhou, Houqiang Li
arXiv:2605.08371v2 Announce Type: replace
Abstract: Multi-view geometry transformers are feed-forward 3D foundation models that jointly predict depth maps, point maps, and camera poses for N images i...
By Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Zi Wang, Qing Guo, Sen He, Huanrui Yang
arXiv:2605.17630v3 Announce Type: replace
Abstract: Frozen segmentation foundation models often fail when the target class appears in a form that is weakly represented during pretraining. To address...
By Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
arXiv:2609.34330v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have demonstrated impressive performance in multimodal understanding, but processing large numbers of visu...
By Tinghao Wang, Yichen Guo, Qizhe Zhang, Yuan Zhang, Weimin Ouyang, Rui Huang, Jiajun Cao, Sixiang Chen, Hao Jiang, Jixian Wu, Zheng Lu, Bofan Zhu, Renyuan Li, Shanghang Zhang
arXiv:2512.22802v2 Announce Type: replace-cross
Abstract: Step distillation accelerates diffusion sampling by training a few-step student to imitate a many-step teacher, but distillation itself remai...
By Amirhossein Tighkhorshid, Zahra Dehghanian, Hamid R. Rabiee