LEGO-OPD introduces a factorized teacher composition for multimodal on‑policy distillation, combining a Language Expert and a Grounding Expert into a single teacher distribution. By treating the language expert as a prior over tokens and the grounding expert as a visual likelihood that updates this prior, the method decouples language reasoning from visual grounding. Adaptive calibration further adjusts the influence of visual evidence at each decoding prefix, preventing over‑ or under‑supervision. Experiments with Qwen3 models demonstrate that LEGO‑OPD outperforms both single‑ and multi‑teacher baselines on multimodal and text‑only reasoning tasks, improving visual perception while preserving language reasoning.
By Jaeyun Shin, Hangeol Chang, Jong Chul Ye
The paper introduces ZeNOVA, a gradient‑free method for aligning initial noise in generative models. It uses annealed soft‑value guidance, manifold‑constrained hyperspherical Langevin dynamics, and Metropolis‑Hastings jumps to address instability in black‑box reward settings. Experiments on image and video models show ZeNOVA outperforms existing zeroth‑order baselines by more stably optimizing noise toward higher rewards.
By Jinho Chang, Jong Chul Ye
This study introduces the first controlled benchmark of generative models for weather data assimilation using real station observations from 11,849 NOAA MADIS stations across the U.S. It evaluates key design choices—diffusion vs. flow matching, pixel vs. latent-space formulations, and inference-time conditioning strategies—against a classical 3D-Var baseline. The benchmark finds that learned generative priors and full-gradient guidance improve RMSE over ERA5, while other design variations offer minimal benefit, especially under sparse observation conditions.
By Ruizhe Huang, Qidong Yang, Jonathan Giezendanner, Sherrie Wang
The paper introduces Kinematic MeanFlow (K-MF), a one‑step action generation policy for Robotic Foundation Models that addresses instability in the MeanFlow framework. By decoupling the time derivative into two sub‑interval terms, K-MF captures early and late denoising dynamics separately, reducing error amplification. Experiments show K-MF achieves faster inference—reducing action‑head latency by 67.5%–74.4% and overall end‑to‑end latency by 30.3%–54.9%—while outperforming multi‑step flow matching on various tasks.
By Jiawei Fan, Sifeng Wang, Yuqing Hou, Anbang Yao
MWOP (Modality-aware Width-wise Operation Pruning) is a method that independently prunes visual‑to‑visual, text‑to‑visual, and text‑to‑text attention paths within each layer of multimodal large language models, and separately selects feed‑forward network channels for visual and textual inputs. It uses a first‑order Taylor criterion to guide pruning, re‑evaluates FFN importance after attention pruning, and applies LoRA‑based recovery training. The approach is paired with path‑sparse Triton attention kernels and compact visual‑side FFN execution to achieve practical acceleration, preserving token sequences while reducing computation.
"whyItMatters":"MWOP achieves a 1.6× prefill speedup on LLaVA‑OneVision‑7B while retaining 99.7% performance, and further boosts token‑compression methods to 2.9× and 2.7× speedups, demonstrating its effectiveness across architectures."
By Xudong Wang, Hao Wu, Haozhe Hu, Peiran Yin, Xinghao Chen, Yunpu Ma, Wei Zhang, Xiaoyu Shen
arXiv:2610.01892v1 Announce Type: cross
Abstract: Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning...
By Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
arXiv:2610.02002v1 Announce Type: cross
Abstract: Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a r...
By Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed
arXiv:2610.02207v1 Announce Type: cross
Abstract: 3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bo...
By Ramazan Fazylov, Stamatis Lefkimmiatis, Ivan Laptev
arXiv:2602.03006v3 Announce Type: replace
Abstract: Deploying Large Language Models (LLMs) for discriminative workloads is often limited by inference latency, compute, and API costs at scale. Active...
By Ziyang Yu, Liang Zhao
arXiv:2607.22629v3 Announce Type: replace
Abstract: Large Reasoning Models produce long, explicit chains of intermediate steps before generating a final answer at inference time. These intermediate t...
By Durgesh Kalwar, Vardhan Palod, Jaya Adithya Pavuluri, Subbarao Kambhampati
arXiv:2606.27824v3 Announce Type: replace-cross
Abstract: Therapeutic peptides are a promising drug modality, but their generation must satisfy multiple therapeutic constraints. We introduce BindSafe...
By Takashi Fujiwara, Hikaru Shindo, Kaushalya Madhawa, Jun Jin Choong, Shuan Chen, Yuna Oikawa, Yiming Zhang, Gyubok Lee, Keisuke Ozawa
arXiv:2609.38342v1 Announce Type: new
Abstract: On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning....
By Zhexi Lu, Subhajit Chaudhury, Tejaswini Pedapati, Keerthiram Murugesan, Lei Yu
arXiv:2609.38599v1 Announce Type: new
Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that...
By Xinyu Wang, Sicheng Lyu, Xiao-Wen Chang
arXiv:2609.38830v1 Announce Type: new
Abstract: Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior...
By Fahao Chen, Linkang Du, Jinhao Zhou, Peng Li, Zhou Su
arXiv:2609.38853v1 Announce Type: new
Abstract: Diffusion distillation is widely adopted to accelerate sampling, and the resulting few-step models are broadly believed to match or even surpass their...
By Yifei Wang, Xiaoyu Wu, Tsu-Jui Fu, Chen Chen, Liang-Chieh Chen, Zhe Gan, Chen Wei
arXiv:2609.38987v1 Announce Type: new
Abstract: Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated f...
By Rui Cai, Wenhui Zhu, Xiwen Chen, Jincheng Cao, Han Yu, Shayan Mohajer Hamidi, Zelin He, Qiyao Ma, Daiwei Chen, Xuanzhao Dong, Yuanda Xu, Jelena Markovic-Voronov, Kayhan Behdin, Zhengze Zhou, Ran He, Alborz Geramifard, Rohit Jain, Zhe Zhao
arXiv:2609.39137v1 Announce Type: new
Abstract: Scaling Large Language Models (LLMs) via Mixture-of-Experts (MoE) enables massive parameter growth with nearly constant per-token computation. However,...
By Peng Jin, Zihan Qiu, Zekun Wang, Bo Zheng, Yang Xu, Tian Xie, Xiao Li, Huaqing Zhang, Haoran Lian, Rui Men, Dayiheng Liu
arXiv:2609.39185v1 Announce Type: new
Abstract: Mamba-style and hybrid language models compress their past into a fixed-size recurrent state that is rewritten at every generated token. Storing this s...
By Snigdha Chandan Khilar
arXiv:2609.40075v1 Announce Type: new
Abstract: Partial Optimal Transport (POT) extends the classical optimal transport problem by relaxing the strict mass conservation constraint, enabling its use i...
By Khoa Nguyen, Dung T. Nguyen, Thong Huynh, Hoang-Hiep Nguyen-Mau, Anh Nguyen, Minh Ngoc Dinh, Juho Kannala
arXiv:2609.40170v1 Announce Type: new
Abstract: MANETs enable flexible infrastructure-less wireless connectivity in dynamic and resource-constrained environments. As modern MANETs exploit multiple fr...
By Tomer Alter, Nir Shlezinger, Michael Segal