Hugging Face Trending Papers

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

Read the original on Hugging Face Trending Papers →

Consumer device repair is an important but underexplored testbed for large language models (LLMs). Repair tasks require reasoning over incomplete problem descriptions, hardware-specific diagnostics, actionable troubleshooting, and safety-critical decisions, where incorrect advice can cause device damage, battery hazards, or permanent data loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 23

Faithful Autoformalization via Roundtrip Verification and Repair

The paper introduces a roundtrip verification method for ensuring that large language models (LLMs) produce faithful formalizations of natural language statements. By formalizing a statement, translating it back to natural language, re-formalizing, and checking logical equivalence with a formal tool, the approach detects inconsistencies without needing ground-truth annotations. When inconsistencies are found, a diagnosis localizes the error to a specific translation step, and a scoped repair operator attempts to correct it. The framework is evaluated on the Texas Transportation Code and Texas Parks and Wildlife Code using Claude Opus and GPT-5, showing that diagnosis-guided scoped repair is most effective and that rules failing the equivalence check exhibit significantly more natural language inference drift.

By Daneshvar Amrollahi, Jerry Lopez, Clark Barrett
arXiv AI
Aug 19

LLM-Only PDDL Domain Repair with Open-Weight Models

The paper evaluates how well open-weight large language models can repair Planning Domain Definition Language (PDDL) models using only LLMs. Experiments show that while the best LLM achieves an F1 score of 0.87—an improvement of 0.38 over a symbolic baseline—it still fails to reliably satisfy test constraints, with a mean test pass rate of only 0.82 and as low as 0.06 on the Thoughtful domain. The study concludes that current open-weight models cannot guarantee the necessary test constraint satisfaction for dependable automated model repair.

By Nader Karimi Bavandpour, Pascal Bercher