← Back to all news
Hugging Face Trending Papers July 15, 2026

Set-shifting Behavioral Test for Harnessed Agents

Read the original on Hugging Face Trending Papers →

What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from cognitive psychology to study how well agents adapt to hidden reliability shifts.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.

  • llms
  • agents
  • benchmarks

Related stories

arXiv AI
Jul 16

Set-shifting Behavioral Test for Harnessed Agents

arXiv:2607. 13396v1 Announce Type: new Abstract: What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session?

By Ziwei Ye
llmsagentsbenchmarks
More like this →
arXiv AI
Aug 6

OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents

arXiv:2608. 05013v1 Announce Type: cross Abstract: LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life.

By Jingsheng Zheng, Xinyuan Fang, Jintian Zhang, Zhengke Gui, Huajun Chen, Ningyu Zhang
llmsagentsmultimodal
More like this →
Hugging Face Trending Papers
Jul 7

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts.

llmsagentsbenchmarkssafety
More like this →
arXiv AI
Jul 8

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents

arXiv:2607. 05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons.

By Wael Albayaydh, Rui Zhao, Ivan Flechais
llmsagentsbenchmarkssafety
More like this →
Hugging Face Trending Papers
Jun 4

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks.

llmsagents
More like this →
arXiv Machine Learning
Jun 5

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

arXiv:2606. 05922v1 Announce Type: cross Abstract: AI agents rely on a harness of skills, tools, and workflows to solve complex problems.

By Wenbo Pan, Shujie Liu, Chin-Yew Lin, Jingying Zeng, Xianfeng Tang, Xiangyang Zhou, Yan Lu, Xiaohua Jia
llmsagents
More like this →