arXiv AI

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

arXiv:2607. 02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

arXiv Computation and Language
4d ago

INTENT-AS-A-TOOL Makes it Easy to Track Agentic Misalignment

The paper introduces INTENT-AS-A-TOOL, a method that equips large language models with intent-targeted tools to provide a fine-grained, judge‑free signal of their commitment to specific behaviors during reasoning. By monitoring the probability of calling these intent tools, the authors can track how intent evolves throughout generation, complementing chain‑of‑thought monitoring and expanding post‑hoc labels into dense trajectories. The approach identifies critical steps for online intervention, demonstrating that action preferences are useful for detecting agentic misalignment in autonomous agents.

By Yutong Zhang, Jianshuo Dong, Peng Xu, Long Wang, Jie Zhang, Tianwei Zhang, Xiaoping Zhang, Han Qiu