arXiv AI By Abolfazl Younesi

Do I Need the Cloud? Uncertainty-Aware Step-Level Handoff for Small Language Model Agents

Read the original on arXiv AI →

The paper introduces STEPGATE, an uncertainty‑aware handoff framework that evaluates each step of a small language model (SLM) agent and selectively escalates difficult steps to a stronger model. On a 52‑task single‑step benchmark, the Qwen2.5‑1.5B/7B pair achieved 82.7% task success with only 30.8% escalation, outperforming local‑only and random escalation baselines. In multi‑turn tests, STEPGATE reached 69.0% trajectory success and 84.0% action success while using only 30.0% cloud actions, demonstrating that step‑level escalation can close much of the performance gap to a stronger backend with fewer remote tokens.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 24

DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers

arXiv:2607. 20531v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate them score the final answer or a fixed "ground-truth" list of tools, both of which are fragile once the underlying data is live and stateful.

By Jerzy Kami\'nski, Ilya Galyukshev, Artem Kuznetsov, Sergey Chuprin, Kirill Redko, Aidar Shumbalov, Anna Kalyuzhnaya
arXiv AI
Sep 30

TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing

arXiv:2605.18859v3 Announce Type: replace-cross Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...

By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi