arXiv:2607. 18026v1 Announce Type: new Abstract: Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment?
By Jiahe Fan, Yinghao Hou, Si Chen, Aiyuan Zhang, Hong Xie, Defu Lian
arXiv:2607. 16252v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning (PEFT) method for large language models.
By Yupeng Chang, Yuan Wu, Yi Chang
arXiv:2607. 20301v1 Announce Type: new Abstract: Fine-tuning has been widely used to adapt large language models (LLMs) for domain-specific tasks.
By Abigail Woodring, Adrian Chan, Rana Muhammad Shahroz Khan, Sukwon Yun, Chau-Wai Wong, Tianlong Chen
arXiv:2608. 05164v1 Announce Type: cross Abstract: Independently trained large language models may develop shared internal representations of semantic concepts despite architectural differences -- but whether this geometric similarity has functional consequences for cross-model behavioural control remains untested.
By Ayushi Agarwal
arXiv:2606. 01717v1 Announce Type: new Abstract: Instruction tuning aligns large language models, including multimodal ones, with diverse user intents, but scaling to heterogeneous mixtures is hindered by gradient interference and bandwidth-heavy synchronization.
By Minsik Choi, Geewook Kim
arXiv:2602. 23638v3 Announce Type: replace-cross Abstract: Federated LoRA provides a communication-efficient mechanism for fine-tuning large language models on decentralized data.
By Haoran Zhang, Dongjun Kim, Seohyeon Cha, Haris Vikalo
The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.
By Baban Gain, Trilok Nath Singh, Asif Ekbal
The paper introduces Align‑LoRA, a unified LoRA framework for multi‑task learning that replaces complex, isolated adapter designs with a single‑adapter model enhanced by a higher rank and an explicit alignment loss. It demonstrates that a router‑free, multi‑head model with high inter‑head redundancy can outperform more elaborate baselines, and that a unified LoRA can achieve competitive performance while enabling weight merging and zero inference latency. Extensive experiments and theoretical analysis confirm that Align‑LoRA surpasses prevailing approaches, offering a simpler, production‑friendly paradigm for parameter‑efficient fine‑tuning of large language models.
By Jinda Liu, Yi Chang, Yuan Wu
arXiv:2609.10305v1 Announce Type: new
Abstract: Language models under one million parameters matter for edge deployment, domain adaptation, and reproducible research, yet a two-layer LSTM or Transfor...
By Fang Li
The paper investigates how large language models (LLMs) share a common Fisher‑Rao geometry in their next‑token probability distributions, revealing that behaviour largely determines this geometry while activation geometry depends on coordinate choices. Across transformer, state‑space, and recurrent architectures, output geometries align more closely than activation geometries, and this shared structure facilitates semantic‑category transfer and improves agreement with human word choices as models scale and train. The study further demonstrates that geometry can guide minimum‑disturbance interventions, enabling reusable control that preserves behaviour better than Euclidean methods and enhances steering, editing, attribution, dictionary learning, and fine‑tuning.
By Dario Picozzi
arXiv:2605.07111v3 Announce Type: replace-cross
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
By Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li, Virginia Smith, Kevin Kuo
arXiv:2606. 05613v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems.
By Long P. Hoang, Yiran Zhao, Wei Lu, Wenxuan Zhang