The paper demonstrates that language models can acquire new capabilities from post‑training data even when the training text is unrelated to the target task. Using a method called Active Taskless Distillation (ATD), the authors show that a single word from a teacher model can transfer knowledge to a student model without any target‑task examples or teacher logits. Experiments on Qwen2.5-1.5B reveal significant performance gains on HumanEval+ and improvements in scientific knowledge, commonsense reasoning, and reading comprehension across various model families.
By Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong
The paper presents a retrieval‑augmented generation pipeline for answering regulatory compliance questions in finance. It builds a three‑stage retriever on LegalBERT and a compact 2B–12B generator served with 4‑bit quantization, achieving a Recall@10 of 0.774 on the ObliQA benchmark and improving answer quality via RAFT‑LoRA fine‑tuning. However, the adapted models fail to transfer to Australian case‑law questions, and a closed‑book model performs almost as well while lacking verifiable grounding.
By Tobias Deu{\ss}er, Abhishek Pillai, Aurelio F. Bariviera, Dhananjay Bhardwaj, Lorenz Sparrenberg, David Berghaus, Christian Bauckhage, Rafet Sifa
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh
BEHAVE models an interacting human group as a complex dynamical system called a HumanSystem, whose state is partly encoded in the interaction structure rather than individual tracks. By incorporating interaction evidence, the method improves group discrimination and captures differences in neighbor-level organization during bottleneck scenarios. The framework derives routing, local dynamics, and stability metrics, enabling real-time querying of group state and critical modes for Physical AI applications.
By Helene Malyutina
The paper introduces ρ_p-LoRA, a rank-allocation strategy for low-rank adaptation (LoRA) in large language models that uses ρ_p regularization (0 < p < 1) to induce sparsity in rank-one components. By regularizing the energy of each component, redundant parts are encouraged to vanish while important ones are retained, leading to an implicit thresholding criterion derived from a two-dimensional proximal subproblem. Experiments on natural language understanding and question-answering tasks show that ρ_p-LoRA achieves performance comparable to existing LoRA baselines.
By Zebang Xie, Chuanyang Zheng, Yik-Chung Wu, Yihang Gao
The paper introduces GBFRVFL, a fuzzy granular-ball random vector functional link network designed to improve robustness in noisy, imbalanced, or uncertain data settings. It employs granular-ball computing to group raw samples into adaptive balls and proposes two membership assignment schemes: F-GBRVFL, which uses fuzzy membership to gauge ball reliability, and SDAP-GBRVFL, which introduces a statistical density‑adaptive Pythagorean membership that adjusts based on class variance, local sparsity, and ball compactness. Experiments on 37 UCI and KEEL datasets show that these models outperform baseline methods in both clean and noisy conditions, achieving higher accuracy and stability.
By A. Quadir, A. Rahaman, P. N. Suganthan, M. Tanveer
arXiv:2609. 29696v1 Announce Type: new Abstract: We construct, for every function class $\mathcal{F}\subseteq[0,1]^{\mathcal{X}}$ and every accuracy $0<\alpha\le 1$, an agnostic sample compression scheme for the empirical squared loss: for every finite sample $S\in(\mathcal{X}\times[0,1])^m$ with arbitrary (noisy) labels, the scheme stores at most $O(\mathrm{fat}(\mathcal{F},c'\alpha)\cdot\log^3(2/\alpha))$ original labeled examples and auxiliary bits, independent of the sample size $m$, and reconstructs a function $\hat f$ with $L_2(\hat f,S)\le\inf_{f\in\mathcal{F}}L_2(f,S)+\alpha$.
By Guangjian Zhang
FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.
By Wanqi Yang, Shiwei Liu
The paper introduces a physics‑guided, data‑driven framework for reconstructing dense ultrasound RF data from sparse acquisitions. It trains an end‑to‑end interpolation network with a hybrid RF‑ and beamforming‑domain loss, stabilized by exponential moving average, and employs random‑skip masking to generalize across varying sparsity patterns and channel configurations. On a held‑out test set, the method achieves a mean SSIM of about 0.95 across decimation factors from ×2 to ×13, consistently improving RF reconstruction and post‑beamforming image quality.
By Luoyuan Zhang, Yiyang You, Ananya Tandri, Yinan Feng, Hyunwoo Song, Jeeun Kang, Youzuo Lin
The paper evaluates post‑training quantization (PTQ) for text‑to‑speech (TTS) models across multiple architectures using a unified protocol. It shows that reducing weights to 4‑bit per‑channel can significantly lower predicted mean opinion scores (UTMOS) and that even 8‑bit per‑tensor scaling can cause severe degradation, with the impact varying by model. A staged ablation identifies the sensitive components, and per‑layer GPTQ can recover performance to within 0.1 UTMOS, while real int8 and int4 kernels confirm the simulated results on hardware, demonstrating that each configuration must be validated on the target runtime.
By Se Un Park, Yutae Kim, Junyoung Park
The paper introduces Routide, a Swift/MLX runtime that runs a quantized Qwen3.6-35B-A3B model on iPhone by keeping expert weights on device storage and a byte‑budgeted subset in memory. It evaluates cache‑policy effects, showing that a 512 MiB LRU cache yields 0.00% demand hits while a 576 MiB LRU reaches 38.58% hits across five 128‑token workloads, indicating that capacity limits depend on policy and workload. The study also reports memory footprints, thermal events, and power estimates, demonstrating that flash‑backed MoE inference is feasible within bounded resources but has measurable limitations.
By Musa Shams
The paper investigates why vision‑language models that tokenize images with vector‑quantized (VQ) codebooks frequently hallucinate objects on grounded yes/no tasks. By applying activation patching across 25 models from eight large‑language‑model families, the authors uncover an early‑layer attention routing circuit shared by VQ‑tokenized VLMs. They develop a three‑gate diagnostic that isolates ten models carrying this circuit, show that swapping a single architectural component (VQ+Linear) introduces the circuit, and demonstrate that ablating the early‑layer ($L_0$) component reduces hallucinations in open‑ended generation by 31 % while other decoding‑time fixes do not.
By Shamanthak Hegde, Xiangrui Liu, Maitreya Patel, Yezhou Yang
The paper presents a Teacher-Student distillation framework for continuous online fault detection in mobile robots. An offline foundation model (TSPulse) generates pseudo‑labels from augmented time‑series data, while a lightweight MiniRocket Student, enhanced with a Recursive Least Squares estimator, performs real‑time inference with a 4.30 ms CPU latency. The Student adapts online to domain shifts, improving VUS‑PR scores from 0.26 to 0.75 and uses an uncertainty‑guided active learning strategy to request minimal operator interventions.
By Jordan Levy, Nicolas Verstaevel, Vincent Talon, Benoit Gaudou
Lossless Anti-Distillation Sampling (LADS) is a defense that keeps the generation process unchanged while reducing the effectiveness of model distillation. It achieves this by coupling latent randomness across accounts, so that a single user experiences the same output as without defense, but a multi‑account distiller receives dependent data that hurts its generalization. Experiments on image, math, and code generation show that LADS degrades distilled model performance while preserving statistical fidelity for individual users.
By Zibo Diao, Jingchu Gai, Xinyue Ai, Zhang Zhang, Zhenyu He, Di He
The paper investigates how post‑training compression techniques—such as pruning, quantization, and distillation—affect demographic fairness in Whisper speech‑recognition models. It finds that pruning and INT4 quantization significantly widen word‑error‑rate gaps between demographic groups, especially for Black/AA and Asian speakers, while distillation tends to reduce these gaps. The study introduces a temporal‑taxation metric to quantify the increased correction effort required for marginalized speakers after compression.
By Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers, Srinivasan Parthasarathy
BanglaTurn is a new corpus of 35,374 Bangla podcast speech samples, each 3 to 15 seconds long, labeled for end‑of‑turn detection through speaker diarization, an LLM pass, and human verification. A Whisper‑based model with task‑specific classification heads achieves 84.33 % accuracy on a balanced test set, outperforming the Smart‑Turn v3 baseline (69.28 %) and reducing the false‑negative rate from 51.57 % to 7.55 %, though with a higher false‑positive rate. The study also details the contributions of encoder‑layer fine‑tuning, multi‑scale pooling, INT8 quantization, and reports inference latency of 165–191 ms on CPU.
By Mizbaul Haque Maruf
DeltaWAM introduces a new approach to world-action models (WAMs) for bimanual manipulation by jointly predicting visual deltas and actions instead of dense future frames, thereby reducing redundant modeling of unchanged content and mitigating nuisance appearance variations. The method employs three architectures with varying representation and computation sharing, and incorporates Streaming Delta Memory (SDM) to update cached anchor context using compact observed deltas, which cuts heavy video-expert processing. Experiments on RoboTwin show that DeltaWAM with SDM raises average success rates from 81.3% to 85.4% in clean settings and from 75.8% to 83.9% under visual randomization, while also reducing training FLOPs by up to 23.77% and inference latency by 36.57%.
whyItMatters":"DeltaWAM improves both performance and computational efficiency for bimanual manipulation tasks by focusing on visual deltas and efficient memory updates, as demonstrated by higher success rates and lower FLOPs on RoboTwin."
By Han Yan, Zishang Xiang, Haokai Jiang, Zeyu Zhang, Qilin Wang, Weiyu Guo, Yandong Guo, Boxin Shi, Hao Tang
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
By Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang
Spectral Amplitude Purification in Distribution Matching for Diffusion Distillation (SAP‑DMD) is a plug‑and‑play method that adaptively modulates the amplitude spectrum of the Distribution Matching Distillation (DMD) directional field. By suppressing the dominant low‑frequency tail, SAP‑DMD reduces low‑frequency dominance and promotes more effective recovery of fine structures and textures. Experiments on PixArt‑α, SD3, and SD3.5 show that SAP‑DMD accelerates training convergence and improves generation quality under both 2‑step and 4‑step sampling.
By Zhenyu Zhou, Can Wang, Chun Chen, Zeyu Zheng, Defang Chen
The paper introduces ImCorr, a method for sub‑pixel semantic correspondence that uses an implicit feature field to eliminate grid‑based quantization errors inherent in patch‑based vision transformers. By training a FiLM‑conditioned decoder to embed precise positional information, ImCorr queries continuous coordinates on the source side and decodes onto a denser grid on the target side, thereby reducing representation‑level errors. Experiments on SPair‑71k and AP‑10K show significant gains at fine‑grained thresholds, with a 6.2‑point improvement at PCK@0.01 over the previous state of the art.
By Yusung Choi