arXiv Machine Learning

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

TACIT-Switch is a cost‑aware routing method that decides when to hand off from a cheaper, smaller language‑model agent to a more expensive, larger one. It learns permanent handoff policies using Teacher‑Annotated Censored Intervention Times (TACIT), treating each annotation as an interval‑censored observation on a cumulative‑risk scale. In controlled simulations, TACIT‑Switch improves success rates by 7.4–11.1 percentage points over other routing baselines while keeping cost comparable, and it achieves the highest held‑out success on ALFWorld and DABench datasets.

arXiv AI
Sep 30

Learn Now, Use Next, Trust Later: Prequential Test-Time Learning for LLM Agents

The paper introduces StepLearn, a nonparametric framework for prequential test‑time learning in large language model agents. StepLearn separates immediate use of informative transitions from persistent trust, turning each transition into a hypothesis that guides the next step and only reusing it after prospective validation across episodes. Experiments on WebArena‑Lite and ALFWorld show StepLearn improves success rates by 2.2–12.7 percentage points over the strongest baseline, with benefits evident from the first task attempts.

By Tong Zhao, Reed Li, Yuyang Hu, Yutao Zhu, Haijin Liang, Haibo Shi, Yu Lu, Zhicheng Dou
Hugging Face Trending Papers
Jul 7

Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade

Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable. We show that failure is predictable early from the agent's internal representations: lightweight per-round probes on hidden activations anticipate eventual episode failure as early as the first interaction round, where scorers reading only the agent's observable behavior are barely better than chance.

arXiv AI
1d ago

Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

The paper introduces STEPGATE, an uncertainty‑aware handoff framework that evaluates each step of a small language model (SLM) agent and selectively escalates difficult steps to a stronger model. On a 52‑task single‑step benchmark, the Qwen2.5‑1.5B/7B pair achieved 82.7% task success with only 30.8% escalation, outperforming local‑only and random escalation baselines. In multi‑turn tests, STEPGATE reached 69.0% trajectory success and 84.0% action success while using only 30.0% cloud actions, demonstrating that step‑level escalation can close much of the performance gap to a stronger backend with fewer remote tokens.

By Abolfazl Younesi
arXiv AI
3d ago

Reach or Solve? Deep Diving into Agentic RL Gains with Checkpoint Handoffs

The paper introduces checkpoint handoff, an evaluation protocol that separates an agent’s ability to reach useful states from its ability to complete tasks in reinforcement learning. By using one checkpoint as a reacher up to a handoff point and another as a solver from the same replayed history, the authors can measure Reach (how often states within a fixed number of actions from success are achieved) and Solve (how often the task is completed from those states). Experiments on TravelPlanner and ALFWorld show that switching the solver from supervised fine‑tuning to RL yields larger gains when RL is used as the reacher, indicating that RL more effectively finds solvable states.

By Xuan Liu, Jingbin Qian