arXiv AI By Liangkai Hang, Junjie Yao, Zhiyu Li, Feiyu Xiong, Hongkang Yang, Zhi-Qin John Xu

Small Initialization Matters for Large Language Models

Read the original on arXiv AI →

arXiv:2606. 17945v1 Announce Type: new Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 27

Emergent Abilities in Large Language Models: A Survey

Emergent Abilities in Large Language Models: A Survey reviews how scaling LLMs leads to previously unseen capabilities such as advanced reasoning, in-context learning, coding, and problem-solving. The paper critically examines definitions, inconsistencies, and the conditions that foster these abilities, including scaling laws, task complexity, pre‑training loss, quantization, and prompting strategies. It also discusses the extension to Large Reasoning Models and highlights safety concerns like deception, manipulation, and reward hacking, calling for improved evaluation and governance.

By Leonardo Berti, Flavio Giorgi, Gjergji Kasneci
arXiv AI
Aug 28

Syntax vs. Semantics: How Transformers Learn Deep Dependencies

The paper investigates how transformers acquire deep semantic dependencies, proposing a mechanistic framework that frames learning as a competition between surface statistics and deep semantics. It identifies a "Gradient Starvation" effect that suppresses error signals for sparse semantic dependencies early in training, delaying structural reasoning until a sudden phase transition. The study also explains the success of Chain-of-Thought strategies and introduces a topology‑aligned contrastive objective that improves variable binding performance by more than twice the gain of standard fine‑tuning.

By Jiangrui Zhao, Xiaoting Du