Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. E...
arXiv:2609.25149v1 Announce Type: new
Abstract: Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph sparsificati...
By Tianfeng Chen, Xianyue Li
Semantic Self-Distillation (SSD) is a method that distills the semantic dispersion of sampled answers from large language models into lightweight student models. These students estimate a prompt-conditioned density before answer generation, providing a prompt-level uncertainty signal via entropy and an answer-level reliability measure through probability density. Experiments on TriviaQA and MMLU show that SSD matches the teacher’s uncertainty estimates while enabling additional tasks such as hallucination prediction, out-of-domain detection, and multiple-choice answer selection.
By Edward Phillips, Sean Wu, Fredrik K. Gustafsson, Boyan Gao, David A. Clifton
CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.
By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
The paper introduces Geometry‑Aware Hyperbolic Residual Quantization, a method that adapts residual vector quantization to hyperbolic space while preserving its telescoping structure. It achieves this by using Hyperbolic Residual Aggregation in the forward pass and a discounted Hyperbolic Straight‑Through Estimator in the backward pass, thereby avoiding geometric inconsistencies and unstable gradients. Experiments on hierarchical prediction, recommendation, image tokenization, and neural audio coding demonstrate improved stability and structural organization of hyperbolic residual codes, with a noted trade‑off between compression efficiency and hierarchical organization.
By Alessio Colombo, Melika Ayoughi
The paper introduces block‑triangular joint drifting, a method that applies a projected drift field to the joint distribution of consecutive states, enabling one‑step generative surrogate models for stochastic transition dynamics. This architecture preserves the current‑state marginal while directly sampling the conditional distribution of next states, allowing stochastic trajectories to be generated with a single model evaluation per time step. Experiments show that the approach achieves accurate marginal and trajectory‑dependent statistics with favorable accuracy‑cost tradeoffs compared to deterministic, diffusion, flow, and distillation‑based generative surrogates.
By Nicholas Geissler, Shreya Jha, Ricardo Baptista, Benjamin Peherstorfer
The paper investigates how different 4‑bit quantization techniques affect privacy when fine‑tuned small language models are deployed. It finds that methods using a calibration corpus, such as Activation‑aware Weight Quantization (AWQ) and Gradient‑based Post‑Training Quantization (GPTQ), prevent the reproduction of planted private records, whereas a calibration‑free format (GGUF Q4_K_M) leaks 5.3% of them. Across models ranging from 0.5 to 7 billion parameters, AWQ consistently leaks the least while maintaining minimal accuracy loss, indicating that the choice of 4‑bit method is a privacy decision as well as a performance one.
By Cristhian Kapelinski, Diego Kreutz
The paper investigates how entropy over chain‑of‑thought tokens influences policy decisions such as gradient application, pruning, and collapse detection. By separating scaffold tokens from substantive content, the authors analyze entropy, Kullback–Leibler divergence, and entropy velocity for each channel, proving differences between raw and content conventions and bounding answer diversity. Empirical results across 23 configurations show that scaffold tokens can account for up to 41% of high‑entropy positions, with entropy share growing through distillation, while content conventions outperform raw surprisal on compression tasks and reveal significant answer leakage in re‑fed chains.
By Marios Papamichalis, Regina Ruane
The paper introduces MIND, a method that distills specialist geospatial model embeddings into a single generalist coordinate embedding with adjustable spatial granularity, using nested supervision across multiple embedding dimensions. MIND’s design allows downstream predictors to use only leading chunks or apply a Chunked Penalty to downweight finer details without retraining the INR. The authors evaluate MIND on CoordBench, a large INR benchmark of 52 datasets and 78 targets, and report that MIND and its Chunked Penalty variant achieve the highest regression and classification scores, especially under regional holdout, establishing a new state‑of‑the‑art for geographic implicit neural representations.
By Isaac Corley, Arjun Rao, Esther Rolf, Konstantin Klemmer, Evan Shelhamer, Nils Lehmann, Marc Ru{\ss}wurm, Gengchen Mai, Nathan Jacobs, Hannah Kerner
RootQuantV2 adapts a frozen DINOv3 Vision Transformer to predict root length and surface area directly from minirhizotron images, eliminating the need for manually traced masks. By training only 11.9 M parameters (3.78 % of the model), it achieves R² values of 0.950 for length and 0.930 for area, improving RMSE by 24.3 % and 20.7 % over the previous RootQuant CNN approach. The method repurposes existing numeric archives of root traits for high‑throughput, automated phenotyping in field‑grown crops.
By Kinjalk Parth, Sebastian Varela, Andrew D. B. Leakey
Ladders-of-Thought (LoT) is a framework that enhances reasoning in small- to mid-scale large language models by automatically generating easier variants of reasoning problems and organizing them into difficulty buckets. It uses a self‑evolving bandit scheduler to adaptively allocate training, improving performance across math and multi‑hop reasoning tasks on 1–8 B models. LoT achieves significant gains (e.g., +32 pp on AddSub, +16 pp on QASC) and converges faster than staged curricula.
By Minghui Liu, Thomas Magelinski, Dehao Yuan, Qi Yu, Furong Huang
Magnitude Profile Pruning introduces a training‑free, calibration‑free method for removing attention heads in Transformer models by statistically detecting outliers in weight row norms. Heads whose projection weights fall within the bulk of the distribution are pruned, while outlier heads are retained. Across several models, the MP‑G variant achieves superior perplexity at various sparsity levels and yields significant parameter and FLOP reductions without requiring forward passes, calibration data, or gradient computations.
By Kasun Dewage, Marianna Pensky, Heranga K. Rathnasekara, Suranadi De Silva
The paper investigates whether domain‑matched calibration data is necessary when compressing large language models for financial tasks. It finds that if compression causes little task damage, the choice of calibration corpus has minimal impact, whereas significant damage—especially from pruning—can be mitigated by using a finance‑specific calibration set (FinMix). The study tests this across multiple models, compression settings, and financial tasks, consistently supporting the link between damage and recovery.
By Junyi Ye, Mengjia Yu, Debapriya Hazra, Guiling Wang
PP‑Net is a hybrid physical‑prior neural network designed to remove scattered light from biomedical images on embedded devices. It combines a denoising network (DFN‑Net), a scattering‑map estimator (ASAP), and a refinement network (GF‑Net) to fuse a physics‑based prior with denoised observations. The method uses progressive synthetic training and cross‑domain transfer to reduce reliance on paired ground truth, achieving significant PSNR, SSIM, and NIQE improvements over baselines while maintaining an inference latency of about 200 ms per 512×512 image on edge hardware.
By Yongfei Guo, Tingjin Chu, Mengzhuo Liu, Hongwei Lou, Yuanhao Gong
The paper introduces Gated Token Recurrence (GTR), a softmax‑free recurrent vision backbone that replaces global softmax attention with gated linear attention, alternating scan directions, and enhanced SwiGLU blocks. GTR is distilled from a DINOv3 teacher using only final‑layer patch‑token alignment, and achieves strong performance on COCO object detection (58.9 box AP) with very low latency (1.908 ms on an RTX 4090). The backbone also transfers to multiple dense prediction tasks and runs efficiently on edge hardware via a specialized CUDA operator and TensorRT deployment.
By Zhe Feng, Longfei Liu, Wei Liu, Kai Chen, Jiangjiang Kong, Wei Zhou, Yifeng Qian, Dexiong Chen, Xuanlong Yu, Xi Shen
This paper introduces a hierarchical Bayesian multitask learning model that assumes a shared sparsity structure across different binary classification tasks. The authors develop a variational inference algorithm for efficient posterior approximation and evaluate the method on synthetic data and pooled microbiome studies. Results show superior support recovery in synthetic experiments and robust, well‑calibrated predictions with informative taxa selection in microbiome classification.
By Haonan Zhu, Andre R. Goncalves, Camilo Valdes, Hiranmayi Ranganathan, Boya Zhang, Jose Manuel Mart\'i, Car Reen Kok, Monica K. Borucki, Nisha J. Mulakken, James B. Thissen, Crystal Jaing, Alfred Hero, Nicholas A. Be
arXiv:2607. 25967v2 Announce Type: replace-cross Abstract: Singular Value Decomposition (SVD) underlies matrix factorisation tasks across many fields, with imaging applications demanding real-time processing.
By Christopher Hahne
The study investigates how the number of prompts and the strategy of refreshing rollout responses affect on‑policy distillation (OPD). Using a 3×3 experiment with 14,080 trajectories and 110 optimizer updates, the authors find that with ten policy snapshots, eight prompts achieve 24.09% accuracy—nearly matching the 24.51% obtained with 14,080 distinct prompts. However, when responses are frozen at the initial policy, increasing prompt breadth actually reduces accuracy, whereas per‑update refresh raises it, producing a 4.07‑point interaction effect. Comparisons with two teacher models show that periodic models excel in short‑budget accuracy and answer completion, but frozen‑response models surpass them in overall accuracy at a 32K output limit, using 1.7–1.8× more response tokens.
whyItMatters":"The findings demonstrate that prompt efficiency in OPD is contingent on both the refresh strategy and the inference budget, informing how to design more effective distillation pipelines."
By Lingxiang Hu, Tianle Xia, Ming Xu, Yiding Sun, Linfang Shang
The paper introduces LatentPort, a method that allows a language model to transfer its live memory to another model without requiring the receiver to reread the context. Experiments on a Qwen3.5 4B-to-9B sibling pair show that adding a Gated DeltaNet (GDN) persistent-state package reduces negative log‑likelihood by 0.747 nats/token and improves performance across 64 PG19 documents. The study also demonstrates that direct recurrent and convolution reuse outperforms learned GDN maps, and a 434,176‑parameter correction further narrows the performance gap to the native 9B model.
By Simon P. Villani
The paper introduces a four‑stage framework—SFT, PG‑CoT, Dynamic, and K‑RL—to improve large language models for Traditional Chinese Medicine prescription generation. It addresses three key gaps: lack of auditable reasoning (SR Gap), failure to adjust prescriptions over time (LA Gap), and non‑enforcement of absolute contraindication rules (SC Gap). Experiments on 12 fine‑tuned models and 6 zero‑shot baselines show that the framework, particularly a 7B Mistral model, outperforms zero‑shot GPT‑5 on all three TCM evaluation metrics.
By Zheng Chen, ZhiCheng Du, Haoxuan Li, Peiwu Qin