arXiv AI By Yi Bin, Xiaoyang Yuan, Haoxi Zeng, Wencheng Ye, Wenqi Shao, Chen Qian, Wei Ye, Yujuan Ding, Zheng Wang, Pengpeng Zeng, Jingkuan Song, Heng Tao Shen

Terminal Agents: A Survey of AI Agents in Command-Line Environments

Read the original on arXiv AI →

The paper surveys AI agents that operate primarily through command-line terminals, defining them as systems whose main action loop involves executing terminal commands, receiving textual feedback, and interacting with a stateful environment. It introduces a seven‑dimensional terminal competence profile to link system architecture, learning, and evaluation, and highlights how behavior is jointly shaped by the model, interface, harness, runtime, and environment. The authors argue for explicit reporting of system and runtime conditions, supported by replayable traces and process‑level evidence, to better understand and benchmark terminal‑mediated agency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

arXiv:2606. 28480v1 Announce Type: cross Abstract: As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding.

By Shoufa Chen, Luyuan Wang, Xuan Yang, Zhiheng Liu, Yuren Cong, Yuanfeng Ji, Feiyan Zhou, Xiaohui Zhang, Fanny Yang, Belinda Zeng
arXiv AI
Aug 20

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

FACET is a framework for synthesizing terminal tasks that preserves source intent and ensures cross‑artifact consistency. It reconstructs agent skills into coherent scenarios, repairs the execution environment, and uses the resulting container state as shared grounding for the instruction, solution, and verifier. By validating and repairing artifacts through execution, FACET produces complex tasks with dense executable checks and data‑efficient supervision, improving performance on Terminal‑Bench 2.1.

By Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, Feng Zhao