arXiv AI By Atm Mizanur Rahman (University of Illinois Urbana-Champaign), Md Arid Hasan (University of Toronto), Syed Ishtiaque Ahmed (University of Toronto), Sharifa Sultana (University of Illinois Urbana-Champaign)

Evaluating LLMs' Effectiveness on Real-World Consumer Device Repair Questions

Read the original on arXiv AI →

arXiv:2606. 03331v1 Announce Type: cross Abstract: Consumer device repair is an important but underexplored testbed for large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Faithful Autoformalization via Roundtrip Verification and Repair

The paper introduces a roundtrip verification method for ensuring that large language models (LLMs) produce faithful formalizations of natural language statements. By formalizing a statement, translating it back to natural language, re-formalizing, and checking logical equivalence with a formal tool, the approach detects inconsistencies without needing ground-truth annotations. When inconsistencies are found, a diagnosis localizes the error to a specific translation step, and a scoped repair operator attempts to correct it. The framework is evaluated on the Texas Transportation Code and Texas Parks and Wildlife Code using Claude Opus and GPT-5, showing that diagnosis-guided scoped repair is most effective and that rules failing the equivalence check exhibit significantly more natural language inference drift.

By Daneshvar Amrollahi, Jerry Lopez, Clark Barrett
arXiv AI
Aug 19

LLM-Only PDDL Domain Repair with Open-Weight Models

The paper evaluates how well open-weight large language models can repair Planning Domain Definition Language (PDDL) models using only LLMs. Experiments show that while the best LLM achieves an F1 score of 0.87—an improvement of 0.38 over a symbolic baseline—it still fails to reliably satisfy test constraints, with a mean test pass rate of only 0.82 and as low as 0.06 on the Thoughtful domain. The study concludes that current open-weight models cannot guarantee the necessary test constraint satisfaction for dependable automated model repair.

By Nader Karimi Bavandpour, Pascal Bercher