arXiv:2608.27831v1 Announce Type: cross
Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, a...
By Gyuhyeong Kim, Hyojung Gwon, Jeonghyeon Kim, Kyuhong Shim, Sunjae Lee
arXiv:2609.22249v1 Announce Type: new
Abstract: This paper treats prompt engineering as a discipline for turning informal human intent into structured AI work specifications. It develops the practice...
By Erfan Loweimi, Hadi Daneshvar, Samira Loveymi, Samir Ouelha, Zhengjun Yue, Hajar Mozaffar, Saturnino Luz
arXiv:2606. 17164v1 Announce Type: cross Abstract: Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals, constraints, and expected outputs are often buried in prose or left implicit.
By Enkhzol Dovdon
The study investigates how users interact with ChatGPT for code generation beyond simple function-level tasks, focusing on project-level benchmarks that involve multi-class dependencies. A user study with 36 participants examined prompting patterns, screen recordings, and chat logs to identify Human‑LLM Interaction (HLI) features that influence productivity. The results highlight three consistently supportive HLI features, five guidelines to boost productivity, and a taxonomy of 29 runtime and logic errors with mitigation strategies.
By Sangwon Hyun, Hyunjun Kim, Jinhyuk Jang, Hyojin Choi, M. Ali Babar
arXiv:2607. 20773v1 Announce Type: cross Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges.
By Zeshu Zhu, Natalie Friedman, Kevin Weatherwax, Emily Eiben
arXiv:2609.14726v1 Announce Type: cross
Abstract: Large language models are increasingly used to scale codebook-based annotation in scientific research, but existing workflows provide limited support...
By Boqin Yuan, Xiaoyi Gu, Fiona Li, Chang Wan, Angel Hsing-Chi Hwang, Jieyu Zhao
arXiv:2608. 01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts.
By B. Sankar, Pawni Yadav, Srinidhi Ranjini Girish, Amogh A. S
Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving. However, when users provide incomplete descriptions or plausible but unverified explanations, LLMs may prematurely align with these assumptions and propose solutions before collecting sufficient evidence.
arXiv:2606. 13220v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as interactive assistants for technical problem solving.
By Fabrizio Marozzo, Pietro Li\`o
arXiv:2607. 14105v1 Announce Type: cross Abstract: For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints.
By Cedric Richter, Salah Ghamizi, Mike Papadakis
arXiv:2607. 20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.
By Junzhi Chen, Harsh Trivedi, Jane Pan, Michael JQ Zhang, Tejas Srinivasan, Niranjan Balasubramanian, Ashish Sabharwal
IDRBench is a benchmark designed to evaluate the interactive capabilities of deep research agents that use large language models. It introduces controlled opportunities for clarification within a common workflow, comparing autonomous and interactive trajectories by measuring task‑specific report alignment and interaction cost. Experiments on 100 tasks with seven LLMs show that interaction consistently improves alignment, though its effectiveness varies depending on the agents’ questions and feedback integration.
By Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang, Jun Yu, Wei Chen, Anthony K. H. Tung