arXiv Computation and Language By Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

Read the original on arXiv Computation and Language →

The paper introduces INTENT-AS-A-TOOL, a method that equips large language models with intent-targeted tools to provide a fine-grained, judge‑free signal of their commitment to specific behaviors during reasoning. By monitoring the probability of calling these intent tools, the authors can track how intent evolves throughout generation, complementing chain‑of‑thought monitoring and expanding post‑hoc labels into dense trajectories. The approach identifies critical steps for online intervention, demonstrating that action preferences are useful for detecting agentic misalignment in autonomous agents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran
arXiv AI
Jun 6

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

arXiv:2606. 05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow/deny decisions, risk categories, and/or explanatory rationales about potential policy violations.

By Yuhao Sun, Jiacheng Zhang, Shaanan Cohney, Zhexin Zhang, Feng Liu, Xingliang Yuan