Hugging Face Blog

LoRA training scripts of the world, unite!

arXiv AI
6d ago

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

The paper shows that pretrained transformers often stop using their depth early, following only a few lines of context. A small rank‑8 LoRA applied to an early layer can extend this chain‑following ability, enabling models like Qwen3‑8B to achieve near‑perfect accuracy on 24‑line chains and significantly longer chains with further training. The LoRA acts as a relay, passing chain identity through middle layers and allowing frozen heads to read further up the chain, with the last useful intervention layer identified in most held‑out models.

By Zehao Jin, Ruixuan Deng, Junran Wang
Hugging Face Trending Papers
Jul 13

LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments

This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components.