The paper evaluates coreset selection methods by incorporating both selection and training time into a unified wall‑clock budget, using a standardized benchmark across four datasets and multiple selectors. Across numerous budget anchors, simple random or full‑data training consistently outperforms sophisticated selectors, and selection costs are dominated by a full‑dataset scan that cannot be amortized. The study also identifies when subset reuse can justify selection and reports several correctness fixes in a popular codebase.
By Yangze Liu, Zhongyi Han
The paper investigates how to design portfolios of agentic AI workflows that vary in reasoning strategy, verification structure, and compute cost. It proposes a portfolio-and-selector framework where multiple workflow executions are run and the best output is chosen, balancing additional compute with potential gains in accuracy. The authors develop exact and approximate optimization methods, evaluate them on three datasets, and show modest improvements over the best single workflow.
By Mojtaba Abdolmaleki, Stefanus Jasin, Boyu Wang
Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment policy is to use the workfl...
arXiv:2605. 25246v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used for optimization modeling and solver-code generation, yet practical operations research and optimization problems often require a harder capability: designing scalable algorithms that exploit problem structure and outperform direct formulation-and-solve baselines.
By Minwei Kong, Chonghe Jiang, Ao Qu, Wenbin Ouyang, Zhaoming Zeng, Xiaotong Guo, Zhekai Li, Junyi Li, Yi Fan, Xinshou Zheng, Xi Jing, Yikai Zhang, Zhiwei Liang, Seonghoo Kim, Runqing Yang, Zijian Zhou, Sirui Li, Han Zheng, Wangyang Ying, Ou Zheng, Chonghuan Wang, Jinglong Zhao, Hanzhang Qin, Cathy Wu, Paul Pu Liang, Jinhua Zhao, Hai Wang
arXiv:2607. 24647v1 Announce Type: new Abstract: AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks.
By Haiqian Yang, Yuan Cao
The paper introduces OSCAR, an LLM‑based framework that translates business descriptions into accurate optimization models while verifying and improving them through a simulator, coder, and reviewer. OSCAR uses a cost‑ordered escalation strategy to select among LLMs of varying price and capability, achieving 95–100% accuracy on benchmark problems with local, open‑weight models. The framework also provides competitive guarantees and token‑cost advantages over existing LLMs like Codex and Claude Code.
By Jinzhi Bu, Haixin Tang, Huanan Zhang
arXiv:2608. 06808v1 Announce Type: new Abstract: The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenarios when solving complex combinatorial optimization problems.
By Shaofeng Zhang, Shengcai Liu, Zhiyuan Wang, Ke Tang
arXiv:2603. 21180v4 Announce Type: replace Abstract: Sequential experimental design under expensive, gradient-free objectives is a central challenge in computational statistics: evaluation budgets are tightly constrained and information must be extracted efficiently from each observation.
By Foo Hui-Mean, Yuan-chin I Chang
arXiv:2406. 06629v2 Announce Type: replace Abstract: This survey examines key advancements in designing features to represent optimization problem instances, algorithm instances, and their interactions within the context of single-objective continuous black-box optimization.
By Gjorgjina Cenikj, Ana Nikolikj, Ga\v{s}per Petelin, Niki van Stein, Carola Doerr, Tome Eftimov
arXiv:2606.22826v2 Announce Type: replace
Abstract: Evaluating LLMs across many model variants---quantized, fine-tuned, or deployment-specific---requires running large benchmarks repeatedly, a proces...
By Devleena Das, Rajeev Patwari, Vikram Kumar Bukka, Nithin Kumar Guggilla, Elliott Delaye, Ashish Sirasao
arXiv:2607. 20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.
By Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.
By Jingbo Wen, Liang He, Ziqi He