Berkeley AI Research

What exactly does word2vec learn?

Read the original on Berkeley AI Research →

What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Berkeley AI Research.

arXiv Computation and Language
Sep 11

Data-Efficient Language Modeling: From Frontier Advancement to Principle-Guided Model Improvement

The paper reports a data‑efficient language modeling study conducted by Qiushi Engine on the BabyLM 2026 Strict‑Small benchmark, using only 10 million corpus words and 100 million cumulative presentations. It describes a three‑stage research program: Stage I built a frontier model via compact restatements and incremental learning; Stage II identified that exact repetition versus aligned restatement affect context use and proposed a principle for organizing experience around contextual dependencies; Stage III applied selective supervision and preservation techniques, achieving a modest overall score increase from 42.02 to 42.25 and the highest public Strict‑Small result as of 8 September 2026. The work also discusses further studies on compression, relational anchors, shared representations, and measurement, and makes models and code publicly available.

By Shuxing Yang, Kaihao Zhu, Junjie Yang, Rui Zhao, Junyao Wu, Yize Wang, Wenhao Li, Fujia Chen, Taowen Deng, Shenzhan Hong, Yaqi Li, Zichen Li, Jincheng Mi, Yuang Pan, Hongsheng Chen, Yihao Yang
arXiv Machine Learning
3d ago

Beyond Linear Concepts: Discovering and Aligning Non-Linear Concept Manifolds in Large Language Models

The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.

By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff
Google AI Blog
Mar 7, 2024

Social learning: Collaborative learning with large language models

Posted by Amirkeivan Mohtashami, Research Intern, and Florian Hartmann, Software Engineer, Google Research Large language models (LLMs) have significantly improved the state of the art for solving tasks specified using natural language, often reaching performance close to that of people. As these models increasingly enable assistive agents, it could be beneficial for them to learn effectively from each other, much like people do in social settings, which would allow LLM-based agents to improve each other’s performance.

By Google AI
arXiv AI
5d ago

NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space

The paper introduces NinaXander, a method for composing frozen language models from different architecture families by inserting a trained shared‑latent adapter between their layers. By running the initial layers of one model, converting the intermediate representation with the adapter, and then continuing with the remaining layers of another model, multiple composed models can be created without retraining. Experiments with RWKV and Pythia show that while some compositions preserve syntactic quality and reduce memory usage, none match the parent model’s accuracy and language‑modeling performance drops on out‑of‑domain data.

By Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri
Google AI Blog
Mar 14, 2024

Cappy: Outperforming and boosting large multi-task language models with a small scorer

Posted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-following framework. This paradigm is exemplified by recent multi-task LLMs, such as T0 , FLAN , and OPT-IML .

By Google AI