Hugging Face Trending Papers

Automated Semantic Fault Localization in SysML v2: A Human-in-the-Loop Framework Using Knowledge-Graph Augmented LLMs

SysML v2's textual syntax enables compiler-based validation of model structure and language conformance. However, semantic mistakes that preserve syntactic validity but violate domain rules cannot be detected through compilers.

arXiv Computation and Language
Sep 1

SkillForge: Compositional Skill Synthesis with Verification-in-the-Loop for Generating Formally Verified Dafny Programs

SkillForge is a framework that breaks down formal code synthesis into reusable atomic skills, each handling a specific subtask such as specification inference, body synthesis, invariant generation, error diagnosis, or repair. A verification-driven harness coordinates these skills by submitting candidates to the Dafny verifier, diagnosing failures, and routing them deterministically to the appropriate repair skill until correctness is achieved or a budget is reached. On a curated benchmark, SkillForge outperforms state‑of‑the‑art agentic and iterative baselines, requiring fewer tokens and lower latency, with ablation studies showing each skill’s measurable contribution and rapid convergence.

By Yanming Liu, Xinyue Peng, Jiannan Cao, Xinyi Wang, Jinbo Su
arXiv AI
Sep 15

Natural-Language to SysMLv2 Translation via Conformance-Driven Iterative Refinement

The paper introduces a framework that translates natural‑language descriptions into SysMLv2 models using a generate‑check‑repair loop driven by a SysMLv2 conformance checker. By embedding the checker as an oracle, the system iteratively repairs generated models until they achieve zero conformance errors, ensuring they are deployable in industrial modeling environments. Evaluation on 151 prompts across four large language models shows the approach raises production‑conformance acceptance from 51.16% to 100%.

By Chance LaVoie, Eladio Andujar Lugo, Taylan G. Topcu, Levent Burak Kara
arXiv AI
Jun 17

EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning

arXiv:2511. 01650v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoning capabilities imperative.

By Ayesha Gull, Muhammad Usman Safder, Rania Elbadry, Fan Zhang, Veselin Stoyanov, Preslav Nakov, Zhuohan Xie
arXiv AI
Sep 3

Dictionary-Guided Mutation Operators for Automated HDL Repair

The paper introduces a dictionary-guided HDL repair system that uses ANTLR-derived mutation vocabularies and a simulation-divergence fault localization module to generate syntactically valid Verilog mutations. The mutation operator performs token substitutions, insertions, and deletions via regex matching, while the fault localization scores source lines based on proximity to diverging output wires, guiding the search. Evaluated on the CirFix benchmark, the approach achieves correct repairs on 14 bug variants, including a multi-bug case, and outperforms CirFix with an 18x speedup on a two-edit benchmark.

By Maisha Mastora, Dean Sullivan
arXiv Machine Learning
Jun 16

Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification

arXiv:2601. 22642v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic systems avoid.

By Chuxue Cao, Jinluan Yang, Haoran Li, Kunhao Pan, Zijian Zhao, Zhengyu Chen, Yuchen Tian, Lijun Wu, Conghui He, Sirui Han, Yike Guo
arXiv AI
Sep 11

Grounded Evaluation and Repair for NL-to-PDDL Problem Generation

The paper presents an end‑to‑end pipeline for translating natural language planning descriptions into PDDL problem instances using large language models. It incorporates multiple checks—syntactic parsing, planner success, domain conformance, an LLM critic, and iterative repair—to ensure faithfulness to the original task. Experiments on Planetarium, AutoPlanBench, and curated PDDL~2.1 problems reveal that operational success can diverge from benchmark‑reference reconstruction, and that structured repair improves outcomes while PDDL~2.1 remains challenging for reference reconstruction.

By Joana Rosa, Pedro Santos, Valdemar Oliveira, Rom\~ao Silva, L. Miguel Silveira, Bruno Martins