arXiv AI

OR for AI That Does OR: Routing LLMs up the Escalator inside the OSCAR Framework

The paper introduces OSCAR, an LLM‑based framework that translates business descriptions into accurate optimization models while verifying and improving them through a simulator, coder, and reviewer. OSCAR uses a cost‑ordered escalation strategy to select among LLMs of varying price and capability, achieving 95–100% accuracy on benchmark problems with local, open‑weight models. The framework also provides competitive guarantees and token‑cost advantages over existing LLMs like Codex and Claude Code.

arXiv AI
Jun 2

FrontierOR: Benchmarking LLMs' Capacity for Efficient Algorithm Design in Large-Scale Optimization

arXiv:2605. 25246v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines.

By Minwei Kong, Chonghe Jiang, Ao Qu, Wenbin Ouyang, Zhaoming Zeng, Xiaotong Guo, Zhekai Li, Junyi Li, Yi Fan, Xinshou Zheng, Xi Jing, Yikai Zhang, Zhiwei Liang, Seonghoo Kim, Runqing Yang, Zijian Zhou, Sirui Li, Han Zheng, Wangyang Ying, Ou Zheng, Chonghuan Wang, Jinglong Zhao, Hanzhang Qin, Cathy Wu, Paul Pu Liang, Jinhua Zhao, Hai Wang
arXiv AI
Aug 28

LLMs Can Design Near-Optimal OR Algorithms

The paper investigates whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems, focusing on inventory control, queueing network control, and assortment optimization. Two usage levels are examined: (1) the model receives a single problem instance and outputs a solution, and (2) the model receives only a problem class description and returns a general algorithm mapping instance parameters to solutions. Using a single untuned prompt and a Python sandbox, the strongest tested model, gpt-5.6-sol, matches or surpasses existing methods on nearly all evaluated instances, even when the algorithm is fixed before seeing evaluation cases, and performance improves markedly across models released within eight months.

By Jackie Baek
arXiv Machine Learning
Aug 11

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

arXiv:2608. 08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries, escalations, and developer wait time are included.

By Srinivasan Manoharan, Junhua Zhao, Fangbo Tu, Haifeng Wu, Jian Wan, Maliah Rajan M, Ashwin Hegde, Mithun Sasidharan, Kalyan Chakravarthi Podamekala
arXiv AI
Sep 10

PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving

PlannerForge is a unified LLM‑agent framework that covers the entire scenario‑based testing pipeline for autonomous driving systems, from scenario generation to ADS assessment, and adds ADS enhancement and benchmarking stages. It was evaluated with ten off‑the‑shelf LLMs across all tasks and five prompt conditions, achieving best‑per‑task scores between 0.88 and 1.00 and matching commercial APIs with open‑source models such as Qwen3.6:35B. The end‑to‑end chaining retains 83% of seed queries for commercial backends and 78% for open‑source, outperforming existing tools like Scenario Factory 2.0 and BM25 in natural‑language generation, attribute realization, and physically valid edits. whyItMatters":"PlannerForge demonstrates that a single LLM‑based system can streamline and improve the fragmented scenario‑based testing workflow for autonomous driving, achieving high performance without domain‑specific fine‑tuning."

By Yuan Gao, Sebastian M\"uller, Mattia Piccinini, Marc Kaufeld, Yuchen Zhang, Finn Rasmus Sch\"afer, Qunying Song, Johannes Betz