The paper investigates whether sink-aware attention head selection remains valid after 4‑bit NF4 weight‑only post‑training quantization. Using Sink Topology Consistency metrics, it finds that global rank preservation stays high across Qwen2.5 and Llama‑3.2 models, yet top‑k head overlap drops to 61–79% and layer‑specific sink‑mass shifts can be substantial. The study also shows that cross‑domain calibration degrades more than within‑domain precision and that recalibration with a small number of samples can recover most of the stability, though full‑map stability may require updating more layers.
By Kuanlin Chen, Chen-Wei Kuo, Cheng-En Ou
The study investigates whether learned scoring functions can outperform hand‑tuned heuristics in Kubernetes node‑selection. Two models—a Random Forest on engineered features and a graph neural network on job dependency graphs—were trained on a large production trace; both achieved modest regression gains (R²≈0.042) but lagged behind a simple free‑CPU heuristic in Top‑1 ranking accuracy (65‑66% vs. 74‑84%). The authors attribute this gap to objective mismatch, noting that pointwise regression rather than a ranking‑specific loss likely limits performance.
By Wang Xuying, Zhibek Sarypbekova
The paper presents a smartphone‑compatible, retrieval‑augmented language model tailored to Bangladeshi statutory law. By distilling a 9‑billion‑parameter Gemma‑2 teacher into a 2‑billion‑parameter student using supervised fine‑tuning and QLoRA, the authors achieve significant gains in ROUGE‑L and BERTScore on an English benchmark while keeping the model lightweight (1.6 GB) and operable offline on a Pixel 6. The system retrieves from 36,029 statutory passages using a hybrid dense/BM25 approach, and cross‑lingual evaluation shows effective Bangla query handling against an English‑only corpus, with a practicing lawyer rating the responses highly in a pilot study.
By MD. Nafis Kamal, Mahadi Hasan Fahim, Talha Ridwan, Nadifa Zaman, Fariha Roushon Florin, Farig Yousuf Sadeque, Saadat Rafid Ahmed
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
AnalogDepth adapts the Depth Anything 3 (DA3) visual geometry model to the spatially structured noise of analog video transmission (VTX) used in FPV drones. The method employs a parameter‑efficient training pipeline that uses student‑teacher knowledge distillation with Low‑Rank Adaptation (LoRA) on the DINOv2 backbone, and trains on a noise bank built from real FPV recordings rather than synthetic Gaussian noise. Experiments on six real FPV flight sequences across three indoor scenes show that this real‑noise training consistently lowers per‑frame depth RMSE and 3D reconstruction Chamfer distance compared to the pretrained DA3 baseline and Gaussian noise baselines.
By Andr\'e Amorim, Pedro F. Proen\c{c}a
SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.
By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
WorldCrafter is a video world model that introduces a camera‑queryable implicit 3D‑aware memory to improve long‑horizon consistency and viewpoint control. The model compresses multi‑view evidence into a limited token budget shaped by the requested viewpoint, integrating historical observations via a memory encoder and pose‑conditioned readout before denoising. Experiments on static and dynamic scenes demonstrate significant gains in consistency and camera‑control accuracy while maintaining visual quality during minute‑scale exploration.
By Wangbo Yu, Kunhao Liu, Wenbo Hu, Shenghai Yuan, Chaoran Feng, Haiyang Zhou, Yukun Huang, Yiran Wang, Wang Zhao, Yingmin Luo, Ying Shan
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
The paper addresses the problem of quantization artifacts in spectral data from the Medtronic Percept PC deep brain stimulation device, which stores local field potential amplitudes as 16‑bit integers. By treating dequantization as an interval‑censored subspace estimation problem, the authors evaluate five correction methods and find that quantized probabilistic PCA most effectively reduces spurious spectral peaks while preserving true peaks and maintaining a low noise floor. The study demonstrates that over 20% of detected peaks in clinical spectra are artifacts, highlighting the need for accurate dequantization in biomarker pipelines.
By Shreesh Karjagi, Elif Ceren Fitoz, Maryam Khalid, Tanya Nauvel, Parisa Sarikhani, Helen S. Mayberg, Christopher J. Rozell, Sankaraleengam Alagapan
The paper investigates selective on‑policy distillation, where a student model is trained only on token positions chosen by a selector. It demonstrates that the commonly used shared learning rate is not neutral: performance varies significantly with the learning rate for different selectors, leading to inconsistent comparisons. The authors attribute this selector‑rate entanglement to the selection process itself and recommend reporting the full arm‑by‑rate matrix for fair evaluation.
By Chencheng Zhu
The paper introduces a Hybrid Quadratic Unconstrained Binary Optimization (QUBO) framework for structured neural network pruning that integrates task‑aware sensitivity metrics (first‑order Taylor and Weight‑Fisher) into the objective’s linear term and optionally uses activation similarity for quadratic interactions. It controls pruning cardinality via a binary search over a capacity incentive rather than an explicit penalty and further refines the pruning mask with a two‑stage QUBO–Tensor‑Train strategy that employs gradient‑free black‑box optimization. Experiments on SIDD image denoising with a Half‑UNet model demonstrate that this Hybrid QUBO outperforms Taylor and L1‑based QUBO baselines in PSNR and SSIM, while also revealing computational and deployment challenges of mask‑based pruning.
By Osama Orabi, Artur Zagitov, Hadi Salloum, Viktor A. Lobachev, Yaroslav Kholodov
The paper introduces Calibrated Clipping, a dynamic method to align FP8 quantization bounds with high‑precision BF16 distributions, thereby mitigating training instability in full‑pipeline FP8 reinforcement learning for large language models. It identifies that compounded FP8 noise distorts importance ratios, causing entropy surges and garbled outputs. Experiments across GRPO and DAPO algorithms on 8B‑32B models show the technique restores performance to BF16 levels.
By Fanchao Chen, Ziheng Jiang, Ziyun Wei, Zheng Zhong, Du Li, Chi Zhang, Haibin Lin, Shivaram Venkataraman
The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.
By Leonard Twagirayezu, Prasenjit Mitra
The paper introduces Strategy Accumulation and Guided Execution (SAGE), a two-stage framework that makes automated fine-tuning of large language models cumulative. In the first stage, a multi-agent pipeline uses Monte Carlo Tree Search to explore training strategies while a Distillation Agent records task-specific insights and cross-task confidence scores into a structured repository. In the second stage, SAGE retrieves relevant experience from this repository to guide training on new tasks, achieving a 12.4‑percentage‑point improvement over a baseline pipeline without accumulated experience on nine unseen tasks.
By Haoran Zhao, Wei Du, Dingwen Yang, Jixuan Huang, Junlin Shang, Lingyong Fang, Ya Guo, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2609.23084v1 Announce Type: new
Abstract: Over-the-air federated learning lets edge devices transmit their local updates simultaneously, reducing the communication overhead. The resulting wavef...
By Jonggyu Jang, Hyeonsu Lyu, Hyun Jong Yang
arXiv:2609.24042v1 Announce Type: new
Abstract: Edge deployment motivates forecasting models with compact parameter storage and low-bit representations. Deep equilibrium models (DEQs) obtain implicit...
By Ruotong Yang, Hongdong Zhu, Qi Gao, Yin Ma, Hai Wei, Kai Wen
arXiv:2609.24141v1 Announce Type: new
Abstract: On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods i...
By Keye Zheng, Hanyu Li, Zhan Cheng, Yuan Gao
arXiv:2609.24197v1 Announce Type: new
Abstract: Speculative decoding losslessly accelerates large language model inference by having a lightweight draft model predict future tokens for verification b...
By Weifan Jiang, Krishna Teja Chitty-Venkata, Megan Flynn, Reed Meyerson, Zhenting Qi, Tianyu Wu, Eldar Kurtic, Minlan Yu, Alexandre Marques
arXiv:2609.24338v1 Announce Type: new
Abstract: An Intraoperative Hypotension (IOH) event is a frequent complication during administration of general anaesthesia with serious downstream consequences,...
By Rithin Nagaraj, Sudiksha Chindula, Bhaskarjyoti Das
arXiv:2609.24370v1 Announce Type: new
Abstract: Self-attention is central to modern Transformer architectures, but its dense dot-product formulation makes it difficult to identify which internal dire...
By Vasileios Arampatzakis, Vasileios Sevetlidis, George Pavlidis