arXiv AI By Jiahao Lu, Mohan Kankanhalli

Agentic-TTT: Training test-time policy for test-time training

Read the original on arXiv AI →

Agentic‑TTT introduces a test‑time policy that decides when and how to apply test‑time training (TTT) to large language models. By treating TTT procedures as tools and learning from the utility gains of its decisions, the policy can trade off performance improvements against computational cost. On a benchmark, Agentic‑TTT nearly doubles the utility of the base model, adapts to unseen domains, and demonstrates autonomous self‑improvement capabilities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

A Survey on Self-Improving Test-Time Intelligence: Feedback-Driven Adapting, Learning, and Scaling at Inference

The paper surveys the emerging field of self‑improving AI systems that refine their behavior during deployment. It introduces a unified framework called feedback‑driven Test‑Time Intelligence (TTI) to connect two previously separate research directions: model state modification via test‑time signals and prediction enhancement through additional inference resources. The survey reviews key methods, applications across vision, language, multimodal learning, generative models, robotics, and healthcare, and outlines open challenges and a research roadmap.

By Shuaicheng Niu, Guohao Chen, Yaofo Chen, Zhiquan Wen, Jinwu Hu, Zeshuai Deng, Deyu Chen, Shuhai Zhang, Renjie Chen, Zihao Lian, Shoukai Xu, Gang Dai, Yunbei Zhang, Wei Luo, Yifan Zhang, Mingkui Tan, Cheng Deng
arXiv Machine Learning
Jul 10

TTHE: Test-Time Harness Evolution

arXiv:2607. 08124v1 Announce Type: cross Abstract: The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures.

By Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, Bo Han
arXiv AI
Jul 7

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

arXiv:2607. 05202v1 Announce Type: new Abstract: Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable procedures for searching, debugging, and verification.

By Xingze Gao, Chuanrui Hu, Hongda Chen, Pengfei Yao, Zhao Wang, Yi Bai, Zhengwei Wu, Yunyun Han, Xiaofeng Cong, Jie Gui, Yafeng Deng, Teng Li