arXiv AI By Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

Read the original on arXiv AI →

ReLiveGym is a diagnostic environment that evaluates long‑lived language‑model agents over weeks of chronologically replayed real‑world streams such as news, market data, and social media. The tasks vary in time sensitivity, reasoning depth, and recurrence, and the study tests eight base language models to see how model choice and harness design—especially action timing—affect performance. Continuous learning from hindsight feedback is also examined to address failure modes in these long‑term tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence

The paper proposes a hierarchical architecture for long-horizon language‑model agents that must operate over days or weeks without forgetting. It introduces three key components: time‑scale levels that store bounded summaries, a clocked tick as the basic action unit, and cascaded intelligence that escalates tasks to more capable models only after review failures. A ten‑day experiment demonstrated that the agent maintained continuity across context resets, adapted its behavior based on early knowledge, and identified where learned components could be integrated.

By Erik Nijkamp, Anurag Koul, Egor Pakhomov, Bo Pang