arXiv AI

Improving Constraint Models with LLM Agents

arXiv:2608. 08127v1 Announce Type: new Abstract: The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints, global constraints, constraint reformulation, and variable representation.

arXiv Computation and Language
Sep 22

CCTU: A Benchmark for Tool Use under Complex Constraints

CCTU is a new benchmark designed to evaluate large language models (LLMs) on their ability to use tools under complex constraints. It includes 200 test cases that average seven constraint types and 4,700‑token prompts, covering resource, behavior, toolset, and response dimensions. An executable validation module performs step‑level checks, and nine state‑of‑the‑art LLMs were tested, revealing that none exceed a 20% task completion rate when strict constraints are enforced, with frequent violations and limited self‑refinement.

By Junjie Ye, Guoqiang Zhang, Wenjie Fu, Zelin Li, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Aug 5

IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation

arXiv:2608. 02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittle: schema, indexing, and semantic errors can cause compilation failures, infeasible models, or incorrect objectives, while iterative repair, search, and multi-agent workflows increase inference cost.

By Penglin Zhu, Linhai Zhang, Jungang Xu, Xinchi Wei, Xiuqi Wu
arXiv AI
3d ago

TACIT: Optimization Models that Learn from Their Mistakes

arXiv:2609.38434v1 Announce Type: cross Abstract: Real-world optimization problems are difficult to model accurately because many objectives and constraints reside in domain experts' tacit knowledge,...

By Maxime Bouscary, Marco Molinaro, Sirui Li, Saurabh Amin, Ishai Menache, Konstantina Mellou
arXiv AI
Sep 4

RECAST: Expanding the Boundaries of LLMs' Complex Instruction Following with Multi-Constraint Data

RECAST is a new framework that generates datasets with far more constraints per example than existing benchmarks, aiming to push large language models (LLMs) to better follow complex instructions. The authors built RECAST-30K, a 30,000‑instance dataset covering 19 constraint types extracted from real prompt‑response pairs, and showed that fine‑tuning on it improves LLMs’ ability to handle complex tasks without harming general performance. RECAST also provides rule‑based and LLM‑based validators for automatic constraint verification, enabling reward‑based reinforcement learning to further enhance model performance on challenging tasks.

By Zhengkang Guo, Wenhao Liu, Mingchen Xie, Jingwen Xu, Zisu Huang, Muzhao Tian, Jianhan Xu, Yuanzhe Shen, Qi Qian, Muling Wu, Xiaohua Wang, Changze Lv, He-Da Wang, Hu Yao, Xiaoqing Zheng, Xuanjing Huang
arXiv AI
Jul 7

OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement

arXiv:2607. 05346v1 Announce Type: new Abstract: We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-ready mathematical formulation as well as executable code.

By Adriana Laurindo Monteiro, Nayse Fagundes, Gabriel Mattos Langeloh, Gustavo de Oliveira Kanno, Priscila Louise Aguirre, Thiago Costa Rizuti da Rocha, Victor Leme Beltran
arXiv AI
3d ago

Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling

The paper introduces OptTips, a knowledge base of 50 expert optimization modeling techniques, and OptDachshund, a multi‑agent framework that generates mathematical models and solver code from natural‑language problem descriptions. Using these tools, the authors create the EfficientOpt benchmark, comprising 561 expert‑reviewed tasks with paired reference implementations, to evaluate large language models (LLMs) on both correctness and computational efficiency. Their experiments with 11 LLMs show a consistent efficiency gap: even when LLMs produce correct solutions, the resulting programs often take longer to solve than expert‑crafted counterparts, highlighting the need to assess both accuracy and runtime performance in LLM‑based optimization modeling.

By Zhong Li, Xin Huang, Jinhui Wan, Xiangyi Wang, Shenkai Zhang, Ruiqi Chen, Wenyu Liu, Zaiwen Wen, Ziyan Luo
arXiv AI
Sep 7

Dynamic Adaptation of the LLM Context for Generating Routines with Coupled Semantics

The paper addresses the failure of large language models (LLMs) in code generation when routine correctness relies on execution-dependent coupling—situations where the meaning of one routine depends on another’s runtime behavior. It introduces a dynamic context adaptation framework that iteratively validates generated code, extracts diagnostic information from execution traces, and guides generation using a knowledge graph and simulated annealing. Experiments show the method surpasses zero‑shot, Reflexion, and OpenEvolve on most benchmark problems, especially where runtime coupling is critical.

By Gnaneswar Villuri, Hashmath Shaik, Alex Doboli
arXiv Computation and Language
Aug 25

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

arXiv:2602.03318v4 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling--a slow and fragile process ill-suited to novel scenarios. While large language models (L...

By Yifan Shi, Jiayi Wang, Minyi Wu, Ye Fan, Jialong Shi, Jianyong Sun