arXiv AI

HARP: The Human--AI Research Platform

arXiv:2607. 20773v1 Announce Type: cross Abstract: Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exchanges.

arXiv AI
Aug 11

How to Ask the AI: A User Perspective Survey for Large Language Model Prompting

arXiv:2608. 07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by typing requests, such as ``plan a three-day Vienna trip'', ``solve the attached mathematical problem'', ``draft an email to inquire review progress'', etc.

By Yiqun Zhang, Yunfan Zhang, Mingjie Zhao, Sen Feng, Yiu-ming Cheung
arXiv AI
Sep 11

Pairit: A Platform for Live Experiments on Human-AI Collaboration

Pairit is an online platform that enables researchers to design, test, and deploy live experiments on human-AI collaboration. Using a single YAML configuration file, users can specify an executable experiment graph—including pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components—and combine any number of humans and AI agents in real-time sessions. The platform has been validated through multiple live deployments, including peer-reviewed studies, and captures high-resolution process traces of communication, negotiation, and collaborative work in human-AI dyads.

By Harang Ju, Sinan Aral
arXiv AI
Jun 30

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

arXiv:2606. 30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time.

By Rahul Khedar, Mayank Malhotra, Avinash Karn, Mouli V, Prakhar Mehrotra
arXiv Computation and Language
Sep 21

Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes

This scoping review examined 48 studies on generative AI chatbots designed to deliver motivational interviewing (MI). It found that most systems were text‑based and disembodied, with about half incorporating dynamic adaptation, and that safety reporting was inconsistent. While user perceptions were generally positive and many studies reported MI‑consistent interactions, evidence for sustained behavioral or functional change remains limited.

By Runze Hu, Jingqi Kong, Yang Yang, Yihang Yang, Jingyao Liu, Haizhou Tang, Shanghang Zhang, Zheng Liu
arXiv AI
Sep 15

Personalizing Personal Health Interfaces: Co-Design with Generative AI

The paper explores how generative AI can lower the barrier to personalizing health dashboards by enabling users to co-design interfaces in Figma Make. In a study with 14 participants, redesigns of Google and Apple Health focused on personal context, future planning, and interactive experiences, though conversational AI designs tended toward chat-window conventions. AI facilitated the materialization of loosely articulated ideas, yet model defaults and generation latency influenced iteration, and the process highlighted interpretability and accountability over privacy, trust, and emotional safety.

By Karthik S. Bhat, Vidhi Shah, Vedika Agnihotri, Dong Whi Yoo, Koustuv Saha
arXiv AI
Sep 4

Efficient Test-Time Adaptation through Human-AI Interaction

The paper introduces Test-Time Adaptation through Human‑Agent Interaction (TAHI), a method that uses iterative human feedback to adapt AI agents to individual users’ criteria. By integrating cross‑session interaction data into agent context and weights, and building an evolving rubric module, the authors demonstrate that agents can improve task success by 4.5–20.9% after only a few interactions. The evolving rubric also serves as a scalable annotation tool, detecting 16.0–22.3% more failures than language models or humans alone, and personalized agents can even generalize improvements up to 8.8% across users.

By Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao, Aspen Chen, Jonas Mueller, Zhiqi Liang, Jett Chen, Michael Ryan, Qianou Ma, Luxi He, Zhoujun Cheng, Andre He, Seungone Kim, Jiayi Geng, Mingqian Zheng, Weiwei Sun, Zheyuan Zhang, Xinran Zhao, Yike Wang, Abe Hou, Liwei Jiang, Pang Wei Koh, Diyi Yang, Graham Neubig, Daniel Fried