arXiv AI By Xiaoqing Wu, Xingyu Fan, Feifei Li, Wenhui Que

When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories

Read the original on arXiv AI →

arXiv:2608. 06057v1 Announce Type: new Abstract: Tool-calling agents infer task state from accumulated dialogue and tool traces.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 6

Multi-Turn On-Policy Distillation with Prefix Replay

We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories.