arXiv AI By Erik Nijkamp, Anurag Koul, Egor Pakhomov, Bo Pang

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

Read the original on arXiv AI →

The paper proposes a hierarchical architecture for long-horizon language‑model agents that must operate over days or weeks without forgetting. It introduces three key components: time‑scale levels that store bounded summaries, a clocked tick as the basic action unit, and cascaded intelligence that escalates tasks to more capable models only after review failures. A ten‑day experiment demonstrated that the agent maintained continuity across context resets, adapted its behavior based on early knowledge, and identified where learned components could be integrated.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.

By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren
arXiv AI
Aug 26

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

The paper introduces Recuris, a recursive Experiential‑Working Memory architecture that lets long‑horizon agents track task progress and select skills based on current needs rather than full history. By coupling working memory with experiential memory, execution becomes structured evidence that localizes failures to specific memory components, enabling a bounded recursive memory‑evolution loop. Across four benchmarks and ten models, Recuris improves task success in 35 of 37 model‑benchmark pairs, raising state‑of‑the‑art performance on tau‑bench and SkillFlow and reducing common long‑horizon failures by up to 80%.

By Zhaochen Yu, Yingcheng Wu, Zhenfei Yin, Kaiyuan Chen, Zhe Zhao, Mengdi Wang, Shuicheng Yan, Ling Yang