arXiv:2606. 06893v1 Announce Type: new Abstract: Large language model agents increasingly rely on Skills to encode procedural knowledge, yet high-quality Skills remain costly to hand-write.
By Yuyang Zhang, Xinyuan Han, Xudong Jiang, Run Wang
arXiv:2606. 03103v1 Announce Type: new Abstract: Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses.
By Wenkai Wang, Tao Xiong, Jingchen Ni, Yunpeng Bao, Xiyun Li, Tianqi Liu, Hongcan Guo, Zilong Huang, Shengyu Zhang
arXiv:2609.22249v1 Announce Type: new
Abstract: This paper treats prompt engineering as a discipline for turning informal human intent into structured AI work specifications. It develops the practice...
By Erfan Loweimi, Hadi Daneshvar, Samira Loveymi, Samir Ouelha, Zhengjun Yue, Hajar Mozaffar, Saturnino Luz
arXiv:2608. 10765v1 Announce Type: new Abstract: Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy.
By Farnaz Soleimani (LISSI), Abdelghani Chibani (LISSI), Yacine Amirat (LISSI), Ghazaleh Khodabandelou (LISSI)
arXiv:2606.08091v2 Announce Type: replace
Abstract: Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents ei...
By Jianhui Wei, Yan Zhang, Jie Tan, Hengchuan Zhu, Xiaotian Zhang, Ziyi Chen, Daoan Zhang, Wei Xu, Yeying Jin, Zuozhu Liu
The paper proposes a new way to describe benchmarks for AI systems that perform knowledge work, outlining four explicit fields: represented activity, tested setting, required work product, and evaluated result. It builds an inventory of 18 work activities from O*NET to enable activity-level reporting across occupations, and evaluates these activities for semantic coherence, algorithm sensitivity, ontology legibility, and human interpretability. The authors apply their framework to three existing benchmarks—GDPval, OfficeQA Pro, and APEX-SWE—to show how different aspects of work can be captured within the same representation.
By Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian
arXiv:2606. 29437v1 Announce Type: cross Abstract: The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a key question: how can we evaluate not only AI-assisted outputs, but also the interaction process that produced them?
By Mohammed Bousmah
RefactorPlatform is an open‑source harness that standardizes the evaluation of repository‑scale refactoring agents by fixing the environment and systematically varying design choices such as model backbone, execution regime, and prompt specificity. Each run operates in an isolated workspace, logs detailed telemetry, and verifies changes with AST‑based checks. Experiments on 100 RefactorBench tasks show that AST‑aware chunking improves performance by 25‑30%, a lean retrieval‑augmented single agent outperforms a sub‑agent configuration, and retrieval’s accuracy gains offset its token overhead, keeping cost per successful refactoring unchanged.
By Aziz Ben Amor, Drish Mali, Mann Acharya, Vijayasri Iyer, S\'ebastien Brati\`eres
The paper introduces a Symbolic-RAG-Generative architecture called GRACE for goal‑oriented conversational systems. GRACE transforms business intent into a fixed objective set and uses a constrained policy to update the conversation state based on visitor‑authored evidence while ensuring visitor utility. Evaluation on real‑estate and professional‑cleaning dialogues shows high accuracy in state transitions, evidence precision/recall, and monotonicity.
By Ramon Gonzalez (Mentomy AI), Antonio Diaz (Mentomy AI)
The paper introduces CoSLR, a Human‑AI collaborative system for systematic literature reviews that incorporates mandatory human checkpoints within a three‑phase pipeline using large language models and Retrieval‑Augmented Generation. In a survey of 63 participants, 42.9 % rated the system’s usability highly, yet 34.9 % indicated they would trust AI‑generated summaries without further human verification after brief interaction. The study highlights that effective human oversight in AI‑assisted literature reviews depends on users’ willingness to engage with the checkpoints, underscoring a calibration issue that interface design must directly address.
By MD Aidul Islam, Malik Abdul Sami, Muhammad Waseem, Zeeshan Rasheed, Kai-kristian Kemell, Zheying Zhang, Pekka Abrahamsson
arXiv:2606. 11835v1 Announce Type: cross Abstract: Collecting participants' lived experiences is central to design research.
By Zhiqing Wang, Steven Dow
The paper introduces RIPPLE, a method for adapting workflow-synthesizing agents through prompt-policy editing without retraining the underlying model. RIPPLE diagnoses failed execution trajectories, maps failures to specific policy segments, and restricts edits to those segments. It then evaluates candidate edits in isolation and replays only those that remain safe after composition, achieving up to a 23.1% improvement in validation success on a synthetic benchmark and positive gains on additional language‑model backbones.
By Manqing Mao, Hong Wang, Samson Koelle, Jie Yuan, Zhuoer Wang, James Feng, Yanjun Lin, Daniel Edmiston, Nikki Lijing Kuang, Zhecheng Sheng, Wei Niu