Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
LexAgentHallu is a new benchmark that profiles hallucinations in legal agents across multi-step interactions. It contains 3,414 instances spanning 17 legal categories and 6 task types, each annotated with a dual-layer taxonomy of 7 high-level and 27 fine-grained hallucination categories. The benchmark introduces fine-grained metrics to quantify and localize failures along an agent’s execution path, revealing patterns such as the Right-Answer-Wrong-Reason effect and clustered hallucination subclasses.
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricated holdings and misci...
arXiv:2608. 19206v1 Announce Type: cross Abstract: Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity.
arXiv:2608. 10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty.
arXiv:2512. 21577v3 Announce Type: replace-cross Abstract: Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs.
The paper examines how hallucinations arise in multi-stage video‑understanding agents by aligning existing benchmarks with the stages of temporal grounding, visual observation, and reasoning. It introduces a causal stage‑intervention protocol that isolates each stage while keeping the downstream task constant, revealing that grounding errors dominate downstream hallucinations and that correct region location matters more than precise temporal overlap. The study also shows that current benchmark scores poorly predict causal sensitivity and can fail under distribution shift, advocating for stage‑aware evaluation methods.