arXiv Machine Learning

Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language Models

arXiv:2608. 00019v1 Announce Type: new Abstract: Deploying large language models (LLMs) for operations research (OR) tasks remains challenging because correctness depends on a coherent modeling process, not merely a correct final answer.

Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.

arXiv Computation and Language
Aug 25

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

arXiv:2602.03318v4 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling--a slow and fragile process ill-suited to novel scenarios. While large language models (L...

By Yifan Shi, Jiayi Wang, Minyi Wu, Ye Fan, Jialong Shi, Jianyong Sun
arXiv AI
Aug 25

Chain of Operators: An Inference-Time Harness for In-Context Operator Learning

Chain of Operators (CHOP) is a framework that enables frozen scientific foundation models to tackle out‑of‑distribution (OOD) tasks without any parameter updates. By leveraging In‑Context Operator Networks (ICON), CHOP breaks down unfamiliar problems into sequences of explicit mathematical operations and model calls, effectively translating OOD queries into the model’s learned regime. Experiments on PDE benchmarks and air‑quality forecasting show that CHOP consistently reduces inference errors while maintaining full interpretability and cross‑equation generalization.

By Minghui Yang, Chenghan Wu, Ling Guo, Liu Yang
arXiv Machine Learning
Aug 19

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models proposes a new method for fine‑tuning LLMs after training. The approach models each group gradient as a random variable, estimates its probability distribution, and uses Dirichlet‑based gradient uncertainty to weight each group’s contribution during policy updates. Experiments on multiple benchmarks show that this uncertainty‑aware aggregation improves the effectiveness of post‑training policy optimization.

By Peizheng Guo, Jianqi Zhang, Xingyu Zhang, Yun Fan, Jiahuan Zhou, Changwen Zheng, Wenwen Qiang
arXiv Machine Learning
Jul 17

Models Can Model, But Can't Bind: Structured Grounding in Text-to-Optimization

arXiv:2605. 21751v2 Announce Type: replace Abstract: Text-to-optimization requires two separable capabilities: modeling -- choosing the right optimization structure -- and binding -- grounding every coefficient, index, and parameter in the concrete problem data.

By Zhiqi Gao, Albert Ge, Alexander Berenbeim, Nathaniel D. Bastian, Frederic Sala