arXiv AI

From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps

arXiv:2607. 23942v1 Announce Type: new Abstract: Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs.

arXiv AI
2d ago

The Cognitive Continuity Test: Verifying Governed State Transitions in Persistent AI Agents

The paper introduces the Cognitive Continuity Test (CCT), a policy-relative contract designed to verify state transitions in persistent AI agents. CCT uses scoped authority, provenance, deterministic application, semantic predicates, and candidate-persistence receipts to classify transitions as verified, violated, or requiring further evidence. Experiments with the IdentityLineageBench dataset show that the CCT can accurately match canonical labels and identify invalid fixtures, while also providing performance metrics for valid-path execution.

By Jun He, Deying Yu
arXiv AI
Aug 24

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

The paper introduces Spec-Driven Agentic Development (SDAD), a framework that leverages large language models to ingest extensive functional requirement documents and repository context in a single workflow, turning specification quality into the engine for autonomous software delivery. SDAD blends disciplined upfront formalisation with rapid implementation, encompassing intent capture, machine‑readable specifications, agentic synthesis, and multi‑agent verification with human sign‑off. It positions AI‑code as a fourth production paradigm, compares it to traditional Waterfall and Agile approaches, and extends the model to team role evolution, quantitative governance metrics, and a staged migration blueprint for practical adoption.

By Vu Hung Nguyen, Thanh Nguyen
arXiv AI
Sep 25

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

The paper introduces Growing Harness, a training method that transforms recurring control logic in large language model agents into reusable executable code, reducing reliance on the model for task-specific decisions. By using strategy-free scaffolds, failure-guided code repair, and success-first gating, the approach learns a shared harness that improves performance across multiple benchmarks and model sizes. Experiments on BrowseComp-Plus and WebArena-Verified show significant gains in success rates and substantial reductions in LLM calls and inference cost compared to traditional tool‑calling agents.

By Laizhen Li, Jiarui Li, Juanjuan Zhao, Kejiang Ye, Ye Li, Cheng-zhong Xu, Xitong Gao
arXiv AI
Sep 7

From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use Agents

The paper introduces an online skill‑evolution framework that transforms interaction traces and evaluator feedback into a persistent, versioned library of reusable procedures for computer‑use agents. By executing each iteration against a frozen library snapshot, the system updates skills without altering the underlying model parameters. Experiments across four OSWorld domains show that the evolving library consistently outperforms an empty‑library baseline, with gains ranging from 5.7 to 18.6 percentage points, while also revealing domain‑specific temporal stability and challenges in skill retrieval and revision.

By Longtao Hu, Xiao Liang, Linchao Zhu
arXiv AI
Aug 24

Terminal Agents: A Survey of AI Agents in Command-Line Environments

The paper surveys AI agents that operate primarily through command-line terminals, defining them as systems whose main action loop involves executing terminal commands, receiving textual feedback, and interacting with a stateful environment. It introduces a seven‑dimensional terminal competence profile to link system architecture, learning, and evaluation, and highlights how behavior is jointly shaped by the model, interface, harness, runtime, and environment. The authors argue for explicit reporting of system and runtime conditions, supported by replayable traces and process‑level evidence, to better understand and benchmark terminal‑mediated agency.

By Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen
Hugging Face Trending Papers
Aug 18

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It first checks functional claims against evidence and evaluates artifacts against nine safety properties, then tests admitted skills in a controlled environment to capture execution traces and identify failures. The approach achieves perfect precision and recall in vulnerability detection, significantly reduces attack success rates, and boosts task effectiveness and security rates in skill generation benchmarks.