Hugging Face Trending Papers

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

Read the original on Hugging Face Trending Papers →

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer to evaluate large language models (LLMs) for fault diagnosis and recovery. It demonstrates that a frontier LLM outperforms locally deployable models in identifying a mass‑shift fault, and shows that successful diagnosis depends on following a complete diagnostic procedure rather than premature conclusions. The study provides an architecture and ensemble evaluation methodology for LLM‑assisted mission management on low‑power autonomous underwater vehicles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 18

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

The paper introduces SPAR, a closed‑loop simulation platform that couples real‑time AUV control software with a higher‑level orchestration layer for fault injection, prompting, and evaluation of large language models (LLMs) in diagnosing and recovering from anomalies. SPAR enables ensemble testing of LLMs, comparing a frontier model with three locally deployable LLMs on a mass‑shift fault scenario across 480 trials, revealing that model choice significantly affects diagnostic accuracy. The study demonstrates that while the frontier model consistently ranks the correct fault mechanism among its top hypotheses, local models succeed mainly when they follow the full diagnostic procedure, and overall diagnosis and operational decisions appear decoupled in this dataset.

By Khalid Halba, Kylie Cooper, James G. Bellingham
arXiv AI
Aug 18

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

arXiv:2608. 14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails.

By Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran