arXiv AI By Hainiu Xu, V\'{i}tor N. Louren\c{c}o, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta, Akash Chandrayan, Luca D'Angelo

ReFract: Benchmarking Perspective Awareness in Language Model Agents with Text World Models

Read the original on arXiv AI →

ReFract is a new benchmark that tests whether language model agents can act appropriately based on a user’s role, a capability called Perspective Awareness. It contains 150 expert‑validated entries derived from anonymized domain support queries, and uses Text World Models to simulate the agents’ operating environments and generate perspective‑aware action trajectories. Current state‑of‑the‑art LLMs solve at most 69% of the tasks, with over half of their trajectories attempting perspective‑violating actions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran