arXiv AI By Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim

TELLME: Test-Enhanced Learning for Language Model Enrichment

Read the original on arXiv AI →

arXiv:2608. 11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

DKL: Decoupled Knowledge Learning for Instruction-Tuned Language Models

The paper introduces DKL, a method for adding new knowledge to instruction‑tuned language models without compromising their instruction‑following abilities. DKL performs extended pre‑training on a base LLM to embed knowledge, then merges these weights into the instruction‑tuned model, avoiding costly instruction fine‑tuning. Experiments show DKL raises RAG accuracy from 54.17% to 79.26% on retrieval failure cases while using far less training data than previous approaches.

By Kushagra Bhushan, Meghanadh Pulivarthi, Sai Krishna Reddy Sathi, Gaurav Pandey, Sonam Gupta, Vineet Kumar, Jaydeep Sen, Yatin Nandwani, Sachindra Joshi, Dinesh Raghu
arXiv Machine Learning
Sep 25

A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs

The paper addresses look‑ahead bias in large language models (LLMs) used for financial prediction, which arises because LLMs are trained on long time‑series data. It proposes a low‑cost solution that adjusts the logits of a base model at inference time using two smaller, specialized models—one fine‑tuned to forget certain information and another to retain it. Experiments show that this method removes both verbatim and semantic knowledge, corrects biases, and outperforms previous approaches.

By Humzah Merchant, Bradford Levy
arXiv AI
Jul 13

Self-Guided Test-Time Training for Long-Context LLMs

arXiv:2607. 09415v1 Announce Type: cross Abstract: Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does not guarantee effective utilization of long inputs.

By Xinyu Zhu, Zhe Xu, Xiaohan Wei, Yunchen Pu, Fei Tian, Chonglin Sun, Kaushik Rangadurai, Hua Zhi, Frank Shyu, Sandeep Pandey, Luke Simon, Yu Meng, Xi Liu
arXiv Computation and Language
6d ago

PALM: Point-in-Time Adaptation for Financial Language Models

arXiv:2609.30316v1 Announce Type: cross Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already obse...

By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
arXiv Computer Vision
Sep 14

TestDG: Test-time Domain Generalization for Continual Test-time Adaptation

This paper introduces TestDG, an online test-time domain generalization framework for continual test-time adaptation (CTTA). TestDG learns features invariant to both current and past test domains during testing, using a new model architecture, adaptation strategy, and prototype selection/update mechanisms. It achieves state‑of‑the‑art results on four CTTA benchmarks and demonstrates superior generalization to unseen test domains.

By Sohyun Lee, Nayeong Kim, Juwon Kang, Seong Joon Oh, Suha Kwak