arXiv AI By Dan C. Hsu, Luke Lu

Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita

Read the original on arXiv AI →

arXiv:2607. 02975v1 Announce Type: new Abstract: Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by acquired information.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

Salesforce Koa: An Enterprise Language Model for Agentic Tool Use

arXiv:2609.15066v1 Announce Type: cross Abstract: We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcem...

By Zixiang Chen, Sufeng Niu, Yingchi Liu, Wenting Zhao, Akshara Prabhakar, Shubham Mehrotra, Bin Bi, Zhujun Lan, Katherine Tan, Mohammad Ramezanali, Tulika Manoj Awalgaonkar, Monojit Banerjee, Jielin Qiu, Shiva Kumar Pentyala, Zhepeng Cen, Anupam Tripathi, Ali Ziaei, Regunathan Radhakrishnan, Darvish Lee Shadravan, Shelby Heinecke, Sitaram Asur, Silvio Savarese, James Zhu, Phil Mui, Huan Wang
arXiv AI
Jun 30

Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization

arXiv:2602. 11351v2 Announce Type: replace Abstract: Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completion beyond passive instruction following and making them essential for real-world, user-centric applications.

By Yihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu, Zuxin Liu, Jiacheng Zhu, Zhang-Wei Hong, Laixi Shi, Ding Zhao
arXiv AI
Sep 12

Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge

The paper argues that deploying generative AI agents requires more than isolated task success; they must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows. The authors introduce two complementary evaluation aspects—operational resilience and considerate participation—to assess how agents recover from blocked work, communicate limits, and adapt to affected people and role boundaries. Using 120 simulated healthcare trajectories across two AI models and twelve stakeholder-derived tasks under varying challenge levels, the study finds that agents shift toward greater human dependence and increased workload as challenge accumulates, while also broadening from task-focused adaptation to task reframing and wider coordination.

By Yuanchen Bai, Zijian Ding, Angelique Taylor