arXiv Machine Learning By Jiaxing Li, Lei Song, Rui Dong, Youyong Kong

Learning from Failures: Heterogeneous Graph Memory for Small Language Model Tool-Using Agents

Read the original on arXiv Machine Learning →

The paper introduces FRESH, a Failure-aware Retrieval framework that uses Experience-Structured Heterogeneous graphs to transform past successes and failures into structured external memory for tool‑using agents. By explicitly modeling dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH enables frozen language models to reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on τ‑Bench and AppWorld with multiple open‑source models demonstrate that FRESH consistently improves task success and tool‑use reliability compared to no‑memory agents and other memory‑based baselines.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
1d ago

AnyAct: Universal Action for Self-Evolving Agents

AnyAct introduces a universal action layer that consolidates diverse tool capabilities into a self‑evolving action space for AI agents operating in open‑world environments. It tackles the scale dilemma, tool non‑stationarity, and heterogeneous feedback by using hierarchical progressive retrieval and test‑time reliability evolution, while a heterogeneous observation grounding module unifies multi‑modal feedback. Evaluations on LiveMCPBench and the newly created OSMCP benchmark show state‑of‑the‑art performance, with significant gains in task success rate and reduced execution steps, especially for models with limited native capabilities.

By Lingrui Xu, Yangqin Jiang, Jiachang Zhang, Xubin Ren, Chao Huang