arXiv Machine Learning

FRAME: Learning the Adaptation Domain with a Mixture of Fractional-Fourier Experts

arXiv:2607. 00162v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) reparameterizes weight updates in a fixed basis: low-rank adapters operate in the spatial domain, while a recent line of spectral methods operates in a fixed Fourier domain.

arXiv Machine Learning
1d ago

Score the Update, Not the Token: Descent-Aligned Routing for Combinatorial LoRA Experts

The paper introduces VANE, a new routing strategy for Mixture-of-LoRA-experts that scores experts based on the expected loss reduction of their updates rather than token similarity. By decomposing each expert into a reader and writer, VANE evaluates all reader–writer pairs using a low‑rank compass that predicts the descent direction, enabling efficient top‑k selection with additive gates. Experiments on Llama‑3 models show VANE outperforms twelve PEFT and MoE‑LoRA baselines while using fewer trainable parameters and providing router scores that better reflect expert usefulness.

By Priya Nair, Lukas Brenner, Maya Lindqvist, Daniel Whitmore, Wen-Hsuan Liu, Tom Saliencro, Amara Okonkwo, Rohan Desai
arXiv AI
Jun 12

The Hidden Power of Scaling Factor in LoRA Optimization

arXiv:2606. 12883v1 Announce Type: new Abstract: In Low-Rank Adaptation (LoRA), the scaling factor $\alpha$ is often treated as a mere complement to the learning rate, yet its role in optimization remains poorly understood.

By Zicheng Zhang, Haoran Li, Jiaxing Wang, Guoqiang Gong, Anqi Li, Yudong Hu, Ting Xiong, Yurong Gao, Junxing Hu, Zhida Jiang, Yifeng Zhang, Pengzhang Liu, Qixia Jiang
arXiv AI
Sep 24

Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation

The paper introduces FHEAT, a fractional diffusion operator derived from the discrete cosine transform, to learn how much spectral mixing each stage of a 3D medical segmentation network should perform. By reparameterizing the operator with a semigroup time, the authors enable the optimizer to decide whether global mixing is needed, resulting in a lightweight U‑shaped architecture (Light‑UNETR) paired with a Kolmogorov‑Arnold mixer (KAN3D). In semi‑supervised experiments, FHEAT‑Seg achieves state‑of‑the‑art Dice scores while dramatically reducing FLOPs through learned spectral sparsification.

By Yi-Hui Shen, Tie-Qiang Li