arXiv Machine Learning By Ji'an Lei, Jian Huang

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

Read the original on arXiv Machine Learning →

TACIT-Switch is a cost‑aware routing method that decides when to hand off from a cheaper, smaller language‑model agent to a more expensive, larger one. It learns permanent handoff policies using Teacher‑Annotated Censored Intervention Times (TACIT), treating each annotation as an interval‑censored observation on a cumulative‑risk scale. In controlled simulations, TACIT‑Switch improves success rates by 7.4–11.1 percentage points over other routing baselines while keeping cost comparable, and it achieves the highest held‑out success on ALFWorld and DABench datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 30

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.

By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.