arXiv Computer Vision

PROVIA: Procedure State Tracking for Online Mistake Detection in Egocentric Videos

PROVIA is a system for online mistake detection in egocentric videos that tracks procedure state by maintaining a factual state and a learned summary of performed steps, including mistakes. It uses Bayesian state merging to create an automaton from correct demonstrations and applies a sequential test to convert per‑frame mistake probabilities into alarms. Evaluated on datasets such as CaptainCook4D, IndustReal, HoloAssist, and IMPACT‑ego, PROVIA outperforms baseline methods while operating at 58–70 fps.

arXiv Machine Learning
Aug 4

Real-Time Detection and Repair of LLM Agent Failures

arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.

By Sunny Dubey
arXiv Computation and Language
6d ago

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.

By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv AI
Sep 2

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

The paper "trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories" examines the limitations of outcome-only evaluation for large language model agents. Using a deterministic tool‑using support‑desk environment with a scripted oracle policy and a fault injector, the authors compare five different judging approaches—programmatic rules, outcome‑only, step‑rubric at two model sizes, and a self‑consistency ensemble—on metrics such as detection, step localisation, fault typing, calibration, and cost across 400 trajectories. The study finds that outcome‑only judges miss many silent faults and generate false positives, while step‑rubric judges achieve higher recall with no false alarms but at greater cost, and that none of the judges read the final reply, allowing fabricated promises to evade detection. "whyItMatters":"The findings highlight that current production‑default outcome‑only evaluations can overlook critical failures in agent behavior, underscoring the need for more nuanced, step‑level judging methods to ensure reliable LLM agent performance."

By Hadi Mohammadi
arXiv AI
Jul 8

StepShield: When, Not Whether to Intervene on Rogue Agents

arXiv:2601. 22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when.

By Gloria Felicia (University of Virginia), Zitha Sasindran (Indian Institute of Science Bangalore), Jinfeng He (Cornell University), Michael Eniolade (University of the Cumberlands), Hemant Kumar (University of Arizona), Milan Hussain Angati (California State University Northridge)
arXiv Machine Learning
Sep 25

Don't Read the Log: Execution Traces Contaminate Verifiers in Video-Generation Agents

The paper investigates how providing execution traces to multimodal judges in agentic video‑generation systems can bias their verdicts. On a benchmark of 109 two‑event clips, traces that falsely report successful tool calls cause large‑language‑model judges to incorrectly accept 78–90 % of failures, while contradictory traces lead to 100 % rejection of correct clips. The effect persists even when judges are instructed to consider only the video frames, indicating that the vulnerability stems from the judges’ learned trust in tool logs rather than the visual content itself.

By Jian Xu
arXiv Machine Learning
Aug 31

CURA: Certified Runtime Alarms for Computer-Use Agents

arXiv:2608.27808v1 Announce Type: cross Abstract: Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. O...

By Divake Kumar, Sina Tayebati, Devashri Naik, Amanda Sofie Rios, Nilesh Ahuja, Omesh Tickoo, Ranganath Krishnan, Amit Ranjan Trivedi
Hugging Face Trending Papers
Sep 2

VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement

VeriPhy is an auditable physical‑verification system that evaluates generated video by compiling prompts into typed physical obligations and a statically validated execution plan before any frames are observed. During execution, it gates calls to frozen low‑level experts (e.g., segmentation, tracking, counting, depth, OCR, audio‑event detection) and returns provenance‑carrying evidence records, which are mapped to a three‑valued state (supported, contradicted, unknown) with full traceability. On a 1,500‑clip corpus of human‑annotated flaw records, VeriPhy accounts for 228 failures out of 304, outperforming a published question‑decomposition evaluator that accounts for 164, while also providing auditable evidence for each verdict.

arXiv AI
Sep 4

Instruction Duplication as an Inference-Time Control Primitive

The paper introduces instruction duplication, a simple inference‑time control that repeats the procedural instruction without retraining or decoding changes. Across seven instruction‑tuned models and 16,800 scheduled generations, duplicating the instruction improves deterministic All‑8 diagnostic‑response success from 90.22% to 93.17% and reduces failures by 30.2%. In downstream Answer Engineering scenarios, duplication further boosts success rates, demonstrating its practical impact on systems that rely on the generated trajectory.

By Victor Lavrenko (PeaceTech VC, Israel)
Hugging Face Trending Papers
Sep 3

Instruction Duplication as an Inference-Time Control Primitive

Instruction duplication is a simple, inference‑time control that repeats the procedural instruction in a language‑model output without retraining or decoding changes. In experiments across seven instruction‑tuned models on 300 medical multiple‑choice questions, duplicating the instruction increased the proportion of deterministic All‑8 diagnostic‑responses from 90.22 % to 93.17 % and reduced failures by 30.2 %. The technique also improved pre‑provisional TF‑IDF recall and, in downstream Answer Engineering scenarios, significantly raised success rates for specific endpoints.