The paper introduces dFlowGRPO, a reinforcement learning framework tailored for discrete flow models (DFMs). It generalizes previous work on diffusion large language models by supporting various probability paths and non-masked source distributions, and formulates denoising as a Markov decision process that leverages transition rates and posterior models. Experiments on the multimodal DFM FUDOKI show that dFlowGRPO outperforms existing GRPO methods on text‑to‑image generation and matches continuous flow models on multimodal understanding tasks.
By Zhengyan Wan, Yidong Ouyang, Panwen Hu, Qiang Sun
The paper introduces FedCORE, a federated adaptation framework for multimodal graph foundation models that jointly optimizes perception (Encoder) and reasoning (GNN) modules via a shared low‑dimensional latent state. Unlike prior methods that freeze the Encoder, FedCORE allows both components to adapt together, addressing the dependency between multimodal evidence extraction and graph‑based relational reasoning. Experiments show that FedCORE significantly narrows the Encoder–GNN pairing gap, achieving an 80.7% reduction compared to independent joint adaptation.
By Zekai Chen, Xun Wu, Hailin Zhang, Xunkai Li, Yu Liu, Kairui Yang, Muyan Huang, Xuaner Chen, Rong-Hua Li, Guoren Wang
OmniMed-Jev is a new medical multimodal model that represents each decision as a Choice, Noul, or Score over a runtime-supplied candidate set, returning a full probability distribution for each decision. By unifying diverse imaging modalities and prediction tasks into a single candidate-conditioned probability model, it makes heterogeneous outputs comparable probabilities rather than task-specific strings. In controlled comparisons against a generative baseline, OmniMed-Jev’s reported probabilities align more closely with observed correctness, reducing calibration error by up to an order of magnitude and reliability error by up to two, while maintaining comparable point-prediction performance.
By Luyao Tang, Cheng Chen
The paper introduces CRAFT, a method for improving compositional generalization in vision‑language‑action (VLA) models. It addresses the issue where models rely on visual shortcuts during fine‑tuning, leading them to execute demonstrated skill combinations that match observations rather than the instructed ones. By training with counterfactual instruction–observation pairs and transferring supervision through reusable skill representations, CRAFT enhances success on unseen skill combinations while preserving performance on demonstrated ones across multiple VLA models and benchmarks.
By Taegeun Yang, Youngju Na, Yoonki Cho, Sung-Eui Yoon
The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.
By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
The paper introduces NEEDLE, a training‑free technique for removing backdoors from large language models. After a trigger is identified, NEEDLE estimates a backdoor direction and a refusal subspace using activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal‑related representations. The method requires no clean reference model or original poisoned data and achieves the lowest attack success rate and minimal impact on model performance across multiple model families and attack types.
By Minoo Kim, Vasileios Lampos, George Drayson
The paper presents a training-free diffusion-based motion planner that replaces learned global trajectory scores with analytical local scores derived from obstacle, smoothness, velocity, and inter-agent feasibility terms. By reconstructing trajectory scores through local interactions between neighboring waypoints and nearby constraints, the method decomposes the denoising process while preserving the optimization structure of classical trajectory methods. Experiments demonstrate that this approach generates smooth, feasible trajectories for large multi-agent tasks in complex environments quickly, outperforming learning-based and optimization baselines without requiring training data.
By Michael Y. Fatemi, Jinhao Liang, Ferdinando Fioretto
GLoC-EHR is a multimodal language model that processes electronic health records by combining a fixed-size global memory of the entire patient trajectory with a local memory of selected events. It generates hospital-course summaries and masked concept descriptions, then is fine‑tuned to cite evidence before answering clinical questions, using group relative policy optimization to reward correct, evidence‑supported responses. On MIMIC‑IV outcome tasks, GLoC‑EHR achieves the highest macro AUROC among compared models when answering directly, and maintains strong performance with evidence‑cited reasoning while adding distinct supported findings from the local memory.
By Chaiho Shin, Kwangsoo Kim
ScaffoldM3C is a lightweight, multimodal, auto‑regressive framework that generates stable 3D block constructions by treating the task as a probabilistic next‑block generation problem. It incorporates text, image, and sketch conditioning, introduces a scaffold block token to aid intermediate stability, and uses Sequential Monte Carlo to explore multiple assembly sequences simultaneously. The model is four times smaller than existing baselines, achieving 5‑ to 20‑fold inference speedups while matching or surpassing state‑of‑the‑art construction quality and stability in both simulations and real‑world robot demonstrations.
By Gadiel Sznaier Camps, Chengyang He, Guillaume Sartoretti, Eduardo Montijano, Mac Schwager
The paper demonstrates that reinforcement learning with verifiable rewards (RLVR) can improve vision‑language benchmark performance even when models are trained without visual input. When real images are introduced at test time, models trained blind recover about half of the performance gain at 3B parameters and nearly four‑fifths at 7B, but extended real‑image training can erode grounding while benchmark gains persist. The authors propose a visual resolvability rule and show that requiring visual evidence for correct answers leads to significant improvements in target discovery and generalization to unseen question types, while controls confirm that the gains stem from actual visual grounding rather than artifacts.
By Haocun Ye, Xinlong Jiang, Qile Chen, Bingyu Wang, Teng Zhang, Shubai Chen, Tingyu Wu, Zhenkun Zheng, Yiqiang Chen
The paper introduces a compact framework that transforms continuous multimodal workplace video into a structured Procedural State Memory called a Work Environment Model (WEM). Using event segmentation theory, it detects segment boundaries based on changes in visual context, location, motion, narration, gaze/object interaction, and optional exocentric workspace evidence, then abstracts each segment into an evidence‑linked event card. These event cards incrementally update the WEM, enabling efficient, auditable documentation and retrieval while respecting on‑premise privacy constraints, and the authors evaluate the system on segmentation quality, memory compression, retrieval fidelity, and long‑horizon QA.
By Vivek Chavan, J\"org Kr\"uger
CODesign is a co-design framework that jointly generates protein sequences and structures to improve consistency between them. It introduces a large consistency‑distilled dataset of about 105,000 dimers and employs a multimodal joint flow model with a consistency‑aware resampling strategy to iteratively refine sequences and side chains. The approach achieves state‑of‑the‑art in silico success rates for protein‑ and ligand‑target binder design, with ablation studies showing a 70.9% performance boost from the distilled dataset and further gains from the resampling mechanism.
By Yuanle Mo, Bo Qiang, Haitao Lin, Qinghan Wang, Gang Du, Odin Zhang, Pheng Ann Heng
GeoLatent introduces a geometry‑guided latent structuring approach for 3D reasoning from 2D images, separating position, direction, and global geometry under geometric supervision. It combines Common–Residual Geometry Alignment (CR‑GEO) with routed optimization to prevent latent collapse and to direct visual answer learning through the latents while maintaining full attention. The method achieves state‑of‑the‑art performance on SPAR‑Bench and SPBench, improving geometry effective rank and overall accuracy.
By Yakun Zhu, Yi Bin, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Duo Peng, Jingkuan Song, Heng Tao Shen
ProtoDCS introduces a robust open‑set test‑time adaptation framework for vision‑language models, addressing the challenge of simultaneously handling covariate‑shifted in‑distribution (csID) and out‑of‑distribution (csOOD) data. It replaces brittle thresholding with a double‑check separation using a probabilistic Gaussian Mixture Model and employs an evidence‑driven adaptation strategy that updates prototypes efficiently, reducing overconfidence and computational cost. Experiments on CIFAR‑10/100‑C and Tiny‑ImageNet‑C show state‑of‑the‑art performance, improving both known‑class accuracy and OOD detection metrics.
By Wei Luo, Yangfan Ou, Jin Deng, Zeshuai Deng, Xiquan Yan, Zhiquan Wen, Mingkui Tan
JET (Justification Evaluation in Transformer) leverages pretrained language and vision‑language models to choose among a limited set of answers without extra training. It directly evaluates candidate likelihoods, reuses computation across candidates, and runs experiments on desktop CPUs and consumer GPUs to measure decision accuracy and execution cost. Results show high accuracy on the MMLU test set, significant speedups from prefix reuse and cache management, and a 30.8% reduction in process time through input preparation optimizations, all while maintaining unchanged outputs.
By Shenghao Ding
Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.
By Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan
The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.
By Kaizhen Tan
The paper introduces Concept-Driven Domain Adaptation (CDDA), a three-stage framework that adapts vision‑language models for concept‑to‑example video retrieval in educational settings. CDDA first structures textual embeddings using textbook and teacher‑handbook concept pairs, then transfers this geometry to documentary visuals with a frozen visual encoder, and finally jointly fine‑tunes both encoders with sparse visual concept supervision. On a middle‑school physics benchmark, CDDA outperforms several multimodal baselines in retrieving concept‑driven moments while preserving concrete image‑text alignment.
By Haiming Zhao, Tai Wang, Kun Zhang, Xicheng Peng, Zhiyang Li
The paper investigates whether the sparsity of Mixture-of-Experts (MoE) models leads to intrinsic semantic organization across modalities and domains. It shows that experts naturally specialize semantically even without explicit modular training. The authors propose ExpertLens, a data‑free method that decodes router weights to identify domain‑specialized experts, enabling selective fine‑tuning that matches or exceeds full fine‑tuning while updating only 21.7–47.0% of parameters and achieving a 4.0× speedup, outperforming LoRA in both performance and efficiency.
By Damiano Marsili, Raphi Kang, Aditya Mehta, Pietro Perona, Georgia Gkioxari
The paper introduces OMAF, an Online MARL framework that uses a one-step Transformer-based flow policy to generate coordinated actions efficiently. By replacing costly iterative sampling with a single-step action generation and a joint optimization scheme that couples softmax Q-value estimation with flow policy objectives, OMAF maintains expressive multimodal behavior while improving training speed. Experiments on 10 tasks from MPE and MAMuJoCo demonstrate up to 3.4× higher returns and 10.5× better sample efficiency compared to baseline methods.
By Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang