arXiv Computation and Language By Christoph Wigbels, Ali Abusaleh, Markus T. Jansen, Alexander Mehler, Markus J. Hofmann

From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora

Read the original on arXiv Computation and Language →

The paper explores how individual text corpora (ITCs) can be integrated into small language models (SLMs) using DoRA fine‑tuning. By training a DoRA adapter for each of 150 participants, the authors show that the adapter can encode a participant’s own ITC into the model’s weights, improving fit to that participant’s held‑out text. However, while the adapter improves log‑loss on a generalized knowledge test, it does not enhance accuracy under a bias‑corrected PMI readout, and adding retrieval does not provide further benefit.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 14

EAR: Entity-Aware Partitioning Approach for Retrieval-Augmented Generation Development

The paper introduces EAR, an Entity‑Aware Partitioning approach that improves retrieval‑augmented generation for multiple‑choice question answering by extracting normalized surface anchors from questions, answers, and the corpus. EAR retrieves local windows around matching anchors and can attach a larger parent passage via an extractive summary, reducing retrieved words by 37.5‑40.2% compared to fixed‑size chunks. Experiments on a cleaned MMLU‑style subset with Mistral, Gemma, and DeepSeek show modest accuracy changes, none statistically significant, highlighting EAR’s methodological contribution of compact, inspectable retrieval units.

By Cenab Batu Bora, Oylum Alatl{\i}, Sebnem Bora, Oguz Dikenelli