arXiv AI By Zinan Tang, Yukun Zhang, Shaomian Zheng, Zhuoshi Pan, Qizhi Pei, Dingnan Jin, Jun Zhou, Yujun Wang, Biqing Huang

CausalMix: Data Mixture as Causal Inference for Language Model Training

Read the original on arXiv AI →

arXiv:2607. 01104v1 Announce Type: cross Abstract: In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 7

A Unified Framework for In-Context Learning with Causal and Masked Language Models

arXiv:2607. 04081v1 Announce Type: new Abstract: In-context learning (ICL) has emerged as a central capability of pretrained language models, yet its theoretical analysis has focused primarily on causal language models trained by left-to-right autoregressive prediction, such as GPT-style models.

By Chenrui Liu, Chuanlong Xie, Falong Tan, Yicheng Zeng, Lixing Zhu