arXiv AI

To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

arXiv AI
Sep 2

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

The paper introduces an action‑class diagnostic framework for multi‑turn tool‑calling in large language model agents, breaking failures into action‑class miscalibration and action‑execution failure across a four‑class action space (TOOL_CALL, ASK, REFUSE, CONFIRM). It defines a self‑revealing upper bound (Acc GAR) to expose state‑grader masking of miscalibration and shows that miscalibration is a significant, previously hidden failure mode, especially for heavily tool‑trained families. The study demonstrates that calibration can be reshaped by context‑only perturbations, but the effects vary widely across models and perturbation mechanisms, underscoring the need for diagnostics beyond aggregate accuracy.

By Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin
arXiv AI
Aug 25

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

CallScreenBench is a benchmark for evaluating small, on-device language models that act as phone secretaries, focusing on their ability to handle unknown inbound calls without a cooperative task. The benchmark measures owner endorsement through five call-and-note metrics, each paired with counter-metrics and uncertainty estimates, and includes guardedness diagnostics to identify safe, tool‑free proxies. Results across 4‑bit checkpoints of 0.6‑4 B parameter models show varying performance on service, recall, plausibility, and triage discrimination, highlighting trade‑offs between quality and guardedness.

By Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
arXiv Machine Learning
Sep 30

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv AI
Aug 26

The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents

The paper investigates the cost–quality trade-offs involved when coding agents switch between low‑cost, low‑capability (LC) and high‑cost, high‑capability (HC) language models during long‑running tasks. By experimenting with different handoff directions, timings, and interfaces—full‑trajectory transfer, compaction, and trajectory removal—the authors find that full‑trajectory escalation recovers less than half of the LC‑to‑HC quality gap while adding significant cost, a penalty they call the handoff tax. Conversely, downshifting from HC to LC yields a more favorable cost‑quality balance, and the optimal interface depends on the direction of the handoff. whyItMatters":"The study quantifies how model handoffs impact both performance and expense, offering guidance for designing more efficient coding agents that balance cost and quality."

By Roy Ganz, Mor Shpigel Nacson, Adi Kalyanpur, Ron Litman
arXiv AI
Jul 7

Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry

arXiv:2607. 03702v1 Announce Type: new Abstract: Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle to effectively leverage failed trajectories: full retries incur high interaction costs, while experience retrieval tends to dilute critical experience signals.

By Weiyang Guo, Zesheng Shi, Longhui Zhang, Zeen Zhu, Min Zhang, Jing Li