Trace2Tower is a transition‑aware EigenTrace framework that transforms raw execution traces of large language model agents into a robust skill hierarchy. By abstracting step‑level interactions into canonical events and constructing a unified graph based on semantic compatibility, transition dynamics, and outcome evidence, it isolates stable, success‑aligned behavioral modes through contrastive spectral decomposition. These modes populate a dynamic skill tower of action templates, procedural routines, and overarching task strategies, which are continuously refined via verifier‑guided feedback, achieving superior performance on ALFWorld and WebShop benchmarks.
By Jiazheng Sun, Boyu Yang, Binhao Yuan, Mingxuan Li, Xin Peng
SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure‑mode attribution. It employs a two‑stage training pipeline to align evaluation with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on industrial short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and significantly advances attribution diagnosis, leading to its deployment at Kuaishou with over 400 million daily active users.
By Zhongxin Huang, Songyang Li, Renzhe Zhou, Feiran Zhu, Chenglei Dai, Zhen Xiao, Xuanping Li, Jingwei Zhuo
arXiv:2608. 12847v1 Announce Type: new Abstract: Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed.
By Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu
MOSAIC is a training‑free framework that adapts Graph Retrieval‑Augmented Generation (GraphRAG) to each query by converting query‑specific evidence needs into a bounded policy over seed selection, traversal, stopping, and evidence selection. It keeps the corpus graph, indexes, scoring, grounding, and answer generation shared, while an LLM analyzer tailors the exploration strategy per query. On GraphRAG‑Bench, MOSAIC improves answer correctness by over 5 points on Medical and 4 points on Novel, achieves high evidence recall and context relevancy, and reduces path and evidence evaluations compared to fixed policies.
By EunKyeong Lee, Kyeong-Jin Oh, Jinwon Kim, Hye Woo Lee, Minsang Song, Hyeongjun Jang, Junyoung Youn
SEEK (Skill‑Routed Evaluation with Evolvable Knowledge) is a framework that externalizes search evaluation criteria into a skill bank, dynamically routes relevant skills for each query‑result pair, and uses a task‑adapted listwise evaluator to generate page‑level judgments and failure mode attribution. It employs a two‑stage training pipeline to align the evaluator with human preferences and a replay‑gated skill bank to incorporate new evaluation knowledge without retraining the model. Experiments on Kuaishou’s short‑video search demonstrate that SEEK improves listwise quality evaluation accuracy and enhances attribution diagnosis, leading to better online search evaluation at scale.
The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model.
"whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."
By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
arXiv:2608. 02011v2 Announce Type: replace Abstract: Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candidate snippets but finalize without inspecting them.
By Daeyoung Roh, Donghee Han
The paper introduces Specified-Foil Counterfactuals for temporal graphs, a method that seeks low‑cost past‑event interventions to make a user‑specified alternative outcome the top prediction. It uses trace‑guided search to compare completed executions of the original prediction with reconstructed incomplete executions of the foil, mapping differences to operations such as DELETE, INSERT, REWIRE, RELABEL, and SHIFT, and then verifies the foil through exact replay. Experiments on continuous‑time dynamic graphs and temporal knowledge graphs show that the approach retains most greedy successes while dramatically reducing predictor evaluations and achieving the specified foil in a majority of cases.
By Minwoo Yu, Young-guk Ha
The paper investigates how trajectory fine‑tuning can enhance small language models (SLMs) as next‑action controllers in retrieval‑augmented question answering. By building a seven‑way action‑prediction task from teacher search traces, the authors fine‑tune SLMs and cross‑lingual SLMs (xSLMs) using LoRA and evaluate on 1,646 held‑out examples, achieving a macro‑F1 of 0.6536 with Granite 4.1 3B. In an end‑to‑end controller/generator swap experiment on 149 trajectories, the fine‑tuned model improves Exact Match from 0.7530 to 0.7946 and token F1 from 0.7783 to 0.8295, demonstrating that trajectory supervision boosts action prediction and evidence‑recording behavior.
By Mohammed Al-Maamari, Saber Zerhoudi, Michael Granitzer, Jelena Mitrovi\'c
arXiv:2607. 13884v1 Announce Type: new Abstract: Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of states, actions, and observations.
By Wenjun Wang, Yuchen Fang, Fengrui Liu, Zibo Liang, Kai Zheng
The paper introduces FRESH, a Failure-aware Retrieval framework that uses Experience-Structured Heterogeneous graphs to transform past successes and failures into structured external memory for tool‑using agents. By explicitly modeling dependencies among tasks, actions, errors, repairs, and execution conditions, FRESH enables frozen language models to reuse reliable strategies, avoid recurring failures, and make safer decisions in stateful tool interactions. Experiments on τ‑Bench and AppWorld with multiple open‑source models demonstrate that FRESH consistently improves task success and tool‑use reliability compared to no‑memory agents and other memory‑based baselines.
By Jiaxing Li, Lei Song, Rui Dong, Youyong Kong
arXiv:2608. 09153v1 Announce Type: new Abstract: Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps.
By Yikai Zhao, Pradeep Kumar Misra, Saurabh Pandey