arXiv AI

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.

arXiv AI
Aug 18

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

arXiv:2602. 15983v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations---a feasibility--correctness gap reaching 90 percentage points on compositional problems.

By Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang, Hanzhang Qin, Chung-Piaw Teo
arXiv AI
Aug 26

Constraint-Guided Enterprise Data Mapping with Large Language Models

The paper introduces Constraint‑Guided Enterprise Data Mapping (CGM), a neuro‑symbolic approach that uses schema‑grounded admissibility constraints to steer large language models (LLMs) in aligning enterprise data. CGM operates in three stages: defining constraints with metadata, generating candidates under relaxed constraints to ensure feasibility, and ranking them with a bounded LLM. Experiments show that hard constraints dramatically reduce candidate space and improve F1 scores, enabling small models to match or surpass large LLMs at a fraction of the cost while reducing expert effort.

By Sebastian Monka, Pramod Anantharam, Thien Vo Minh, Lavdim Halilaj
arXiv AI
Aug 5

IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation

arXiv:2608. 02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost.

By Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu
arXiv AI
Sep 12

SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics

SemVerBench is a benchmark that evaluates how well large language models (LLMs) understand and apply version-constraint resolution semantics, such as determining whether a version satisfies constraints like ^1.2.3 or >=2.0. The study finds that many models struggle with certain corner cases, with GPT‑5.1 performing poorly while Claude and Opus perform much better. The authors suggest that the failures stem from an activation/application gap rather than a lack of knowledge, and recommend that coding agents delegate version resolution to a dedicated resolver tool.

By Qibai Chen, Zeming Liu
arXiv AI
2d ago

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

The paper introduces a benchmark for evaluating whether off‑the‑shelf small language models (SLMs) can reliably perform microtasks that support a large language model (LLM) planner, such as auto‑approving shell commands, writing memory, selecting tools, and ranking past turns. Using fixed prompts and confidence‑interval‑aware eligibility thresholds, the authors test several Qwen3 models (0.6/1.7/4/8 B) in FP16 with no tuning and find that none of the 16 configurations meet the eligibility criteria. Quantization to 4‑bit precision further degrades performance, with the eligibility gap tracking model size rather than precision, and the issue persists across different models (e.g., Llama‑3.x) and prompt variations.

By Jundong Hu, Shekar Ramachandran