The paper introduces CrossUAV, a benchmark for joint object detection and instance segmentation in UAV imagery, and proposes Cross-Granularity Socialized Collaboration (CGSC), a framework that regulates hierarchical interactions between tasks. CGSC progressively activates cross-task exchanges and adaptively adjusts interaction strength based on task contribution, aiming to reduce interference and exploit complementary coarse- and fine-grained knowledge. Experiments show consistent improvements on both detection and segmentation tasks, supporting the effectiveness of hierarchical dynamic interaction for cross-granularity collaboration.
By Xinjie Yao, Ruipu Zhao, Yunqi Zhu, Zhihe Fan, Zhoupeng Guo, Weihao Li, Zhen Wang, Qilong Wang, Pengfei Zhu
The paper introduces D$^3$-MOPD, a dynamic domain scheduling method for multi-teacher on‑policy distillation. It adapts the domain mixture during training by monitoring each domain’s reverse‑KL trajectory, thereby allocating more compute to slower‑converging domains and less to those that plateau early. Experiments on a Qwen3.6‑35B‑A3B student show that D$^3$-MOPD closes 97% of the student‑to‑teacher performance gap, matches peak performance with roughly three times fewer rollout steps, and outperforms specialist teachers on most benchmarks.
By Zechen Sun, Zhiwei Zhang, Fei Zhao, Juntao Li, Mu Chuan, Huayu Deng, Guojian Zhan, Wenliang Chen, Yao Hu, Min Zhang
Agile-WAM is a tactile World Action Model that jointly predicts future visual and tactile states and robot actions for contact‑rich manipulation. It encodes visual and tactile observations into a shared latent space and uses a vision‑tactile‑to‑action flow‑matching process to generate action chunks and future latents. The model introduces multi‑horizon multimodal prediction, leveraging the different timescales of vision and touch, and achieves a 29.4 % improvement in real‑world success rates with 11.9 ms inference latency across nine simulated and five real‑world tasks.
By Hanchu Zhou, Brendan Lynch, Raman Goyal, Dechen Gao, Begum Kasap, Boqi Zhao, Junshan Zhang
The paper introduces FRESH‑GEORANGE, a semantic‑spatial range retrieval system that separates source‑watermark freshness from optional record age. It uses geographic cells and semantic microblocks for pruning, and offers an exact mode that guarantees 100% recall and a certified mode that can stop early while providing a deterministic recall lower bound. A CPU pilot on 2,500 OpenFlights airport records demonstrates the system’s correctness, achieving high recall with a modest latency overhead compared to a spatial‑first exact baseline.
By Taimoor Ahmad
The paper investigates how the choice of the π parameter in λπ norm-constrained adversarial attacks influences the sparsity and smoothness of the perturbations. By applying two established sparsity metrics and introducing three new smoothness measures—including one based on first-order Taylor approximations—the authors perform extensive experiments on real-world image datasets and various neural network architectures. Their results indicate that λρ norms with π values between 1.3 and 1.5 consistently provide the best balance between sparsity and smoothness, challenging the common use of λ1 or λ2 norms.
By Christof Duhme, Florian Eilers, Xiaoyi Jiang
The paper introduces Block Parallelism (BP) and Context‑Sharded Block Parallelism (CSBP) to improve training efficiency for Block Diffusion Language Models (BDLMs) with long contexts. By assigning each corrupted‑block computation to a separate rank and sharding the shared clean sequence, CSBP reduces communication overhead and memory usage while preserving training semantics. Experiments on 16 H200 GPUs and 8 H100 GPUs show throughput gains of up to 1.61× and 7.59×, respectively, and higher benchmark pass rates in practical fine‑tuning scenarios.
By Tarun Suresh, Pranshu Chaturvedi, Hangoo Kang, Parth Shroff, Ishan S. Khare, Hermann Kumbong, Azalia Mirhoseini
The paper introduces a compression framework for the Whisper automatic speech recognition model that jointly optimizes six deployment dimensions—model size, temporal resolution, encoder token stride, low‑rank adaptation capacity, weight precision, and sparsity pattern—using NSGA‑III. The optimization targets three objectives: word error rate, inference FLOPs, and memory footprint. Evaluating 1,680 configurations, the study identifies compression combinations that outperform single‑axis scaling and notes that 1:4 structured sparsity cannot maintain acceptable accuracy within the tested budgets.
By Vyom Agarwal, Mokshda Gangrade, Siddharth Pal, Jerry Wu
The paper presents a method for training neural networks on synthetic data to approximate the optimal Bayes estimator for dense emitter localization. By doing so, it demonstrates that neural networks can effectively handle complex localization tasks in high-density scenarios. The study supports future efforts to develop high-throughput, large-field-of-view super‑spatiotemporal resolution single‑molecule localization microscopy (SMLM) systems.
By Yi Sun, Mona Sharifi, Muzna Yumman
The paper introduces AttnPrint, a white‑box fingerprinting method that extracts low‑frequency components of cross‑modal attention distributions to identify multimodal large language models (MLLMs). It also presents DistillTrace, a black‑box auditing tool that uses hypothesis testing of MLLM outputs to detect potential model infringement. Experiments on 154 model instances across 19 architectures show that AttnPrint effectively detects derivative models and remains robust to downstream modifications, while DistillTrace reveals distillation relationships under various techniques.
By Chao Huang, Meng Tong, Kejiang Chen
FCx is a new algorithm that generates counterfactual explanations while explicitly enforcing feasibility constraints. It uses a modified Variational Autoencoder with a multi‑factor loss to produce realistic, low‑cost counterfactuals that satisfy both hard constraints supplied by users and soft constraints inferred via causal inference. Experiments on four public datasets demonstrate that FCx matches state‑of‑the‑art performance across multiple metrics while guaranteeing feasibility.
By Kleopatra Markou, Vana Kalogeraki, Dimitrios Gunopulos
TinyCNN is a lightweight convolutional neural network with only 193,190 trainable parameters designed for on‑device plant disease classification. It achieves 98.88% test accuracy on the 38‑class PlantVillage benchmark, outperforming larger models while consuming far less energy, memory, and cost. The study also explores knowledge distillation to further compress the model and evaluates cross‑dataset robustness, finding a significant performance drop when moving from PlantVillage to PlantDoc due to background‑driven shortcut learning.
By Ngoc-Bao Ho-Lam, Thai-Anh Nguyen
The paper presents a deep‑learning perception framework for selective robotic cotton harvesting, evaluated on 1,008 field images captured under diverse lighting and weather conditions. Detection models from YOLOv8 to YOLOv13 were benchmarked, with GELAN‑s achieving the best trade‑off between accuracy and speed. For segmentation, YOLOv12‑m‑seg outperformed other models, and a detection‑prompted segmentation approach using GELAN‑s bounding boxes further improved localization for SAM variants. Field trials with a UR5e robot and ZED2i camera confirmed YOLOv12‑m‑seg’s real‑time performance for cotton boll detection, segmentation, and selective picking.
By Thevathayarajh Thayananthan, Xin Zhang, Isuru Laddusinghe Badu, Jonathan Harjono, Glen C. Rains, Beiwen Li, Leonardo M. Bastos, Nuwan K. Wijewardane, Vitor S. Martins
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
The paper critiques the common practice of evaluating depth usage in depth‑recurrent language models by truncating depth during inference and measuring performance decline. It argues that this method conflates three distinct effects—fewer block applications, reduced computation, and an out‑of‑distribution readout—yet is usually interpreted as measuring only the second. To address this, the authors introduce the Depth Control Protocol (DCP), a suite of positive and negative controls that isolate each factor, along with a training intervention to confirm causality, specifically tailored for depth‑wise weight‑sharing architectures.
By Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung
The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.
By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu
The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.
By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua
The paper introduces On‑Demand Attention (ODA), a decoding strategy that lets pretrained language models decide when to use global attention based on a lightweight recall head. ODA keeps the original model weights unchanged, only training the recall head, and can be implemented with GPU‑side conditional execution to reduce global reads. Experiments on Qwen, Gemma, and hybrid‑attention models show that ODA largely recovers performance lost by local attention while cutting the number of global attention operations.
By Haibo Feng, Ruiqi Liang, Hanyang Peng, Shiqi Yu
GeLaCo is an evolutionary method for compressing large language models by collapsing layers through parametrized weight merging. It uses population-based search with a fitness function that balances similarity of residual updates and language modeling KL divergence, enabling both single and multi-objective compression. The approach yields Pareto-optimal trade-offs between compression and quality, outperforming existing methods in perplexity and generative evaluations.
By David Ponce, Thierry Etchegoyhen, Javier Del Ser
The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.
By Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
Video DeltaNet (VDN) introduces a hybrid attention mechanism for livestream video generation, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). VDA updates memory once per frame, integrating spatial tokens, while separate output projections and learnable gates balance the two branches. Applied to MiniMax H3, VDN achieves a 14.5× speedup over the dense baseline, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs.
By Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng