arXiv AI
Sep 11

LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents

LexAgentHallu is a new benchmark that profiles hallucinations in legal agents across multi-step interactions. It contains 3,414 instances spanning 17 legal categories and 6 task types, each annotated with a dual-layer taxonomy of 7 high-level and 27 fine-grained hallucination categories. The benchmark introduces fine-grained metrics to quantify and localize failures along an agent’s execution path, revealing patterns such as the Right-Answer-Wrong-Reason effect and clustered hallucination subclasses.

By Yujin Zhou, Mingxuan Zheng, Chuxue Cao, Huang Yidan, Jiale Chen, Yike Guo, Sirui Han
arXiv AI
Sep 25

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

The paper examines how hallucinations arise in multi-stage video‑understanding agents by aligning existing benchmarks with the stages of temporal grounding, visual observation, and reasoning. It introduces a causal stage‑intervention protocol that isolates each stage while keeping the downstream task constant, revealing that grounding errors dominate downstream hallucinations and that correct region location matters more than precise temporal overlap. The study also shows that current benchmark scores poorly predict causal sensitivity and can fail under distribution shift, advocating for stage‑aware evaluation methods.

By Shuzhi Gong, Fengze Sun, Yuansan Liu