arXiv Machine Learning By Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

Read the original on arXiv Machine Learning →

The paper investigates why the train‑validation performance gap widens during fine‑tuning of pretrained models. It proposes a dynamic structural explanation: as training proceeds, updates shift from broadly reusable features to more example‑specific ones, increasing gradient heterogeneity and the gap. Experiments on synthetic ResMLP hierarchies, NLP models (RoBERTa, DeBERTa, Qwen) across six datasets, and vision models (ResNet‑18) confirm that higher reliance on private features correlates with larger accuracy gaps, supporting the proposed account.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 20

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

The study measured the impact of a single training example on a GPT‑2 model by running 24 counterfactual experiments. 32 models were trained from scratch on OpenWebText, and at a specific training step a single batch row was replaced with a 194‑token passage under three conditions (fluent prose, fabricated subject, random characters) or left unchanged. Results showed that the passage was learned from one exposure and decayed, with measurable differences in cross‑entropy up to 50 steps after injection but no lasting effect at the final step.

By Zachary Speck, Asa Shepard
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong