arXiv AI By Prajwal S. Venkateshmurthy

TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair

Read the original on arXiv AI →

arXiv:2608. 07617v1 Announce Type: new Abstract: Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimiters, mismatched environments, broken imports, or package conflicts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 24

Escaping Python Dependency Hell: A Hybrid Replay-and-Repair Pipeline for Python Dependency Resolution

The paper introduces PLLM+, a hybrid pipeline for resolving Python dependency conflicts that combines deterministic steps—such as static AST inference, replaying known successful configurations from a solutions database, and live PyPI validation—with an LLM-based repair fallback. Evaluated on the HG2.9K benchmark of 2,891 failing snippets, PLLM+ successfully fixes 1,500 cases, outperforming the baseline PLLM and reducing average runtime from 368.7 to 71.8 seconds per snippet. The majority of fixes (1,495) come from replaying existing configurations, while the LLM fallback contributes only five additional solutions.

By Veronica Poweska, Ariana Oyanguren, Jessica Pourleyli, Sourena Khanzadeh, Manar Alalfi
arXiv AI
Sep 7

When LLM Decompilers Recompile More and Preserve Less

The paper examines how large‑language‑model (LLM) decompilers, which produce clean, idiomatic C code, are currently evaluated mainly on recompilability and passing shipped tests. It shows that these metrics can mask significant behavioral differences: a decompiled function may recompile and pass all tests yet diverge on other inputs or lose disclosed vulnerabilities. To address this, the authors propose Decompile‑Diverge, a behavioral oracle that synthesizes drivers, fuzzes inputs, and compares the decompiled code’s behavior to the original, revealing divergences in up to 13% of cases and exposing gaps in current evaluation suites.

By Chang Liu, Edward Raff, Kristopher Micinski