arXiv Machine Learning By Frank Hu, Shriram Chennakesavalu, David Graff

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

Read the original on arXiv Machine Learning →

Frontier large language models (LLMs) are examined as batch optimizers in both continuous and discrete settings. The study finds that while LLMs perform competitively in zero‑shot optimization of numerical test functions, their performance is less robust than classical non‑LLM methods. However, LLMs excel in semantically rich, discrete spaces that resemble their pretraining data, demonstrating strong batch optimization behavior in such contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 7

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

The paper surveys efficient reasoning in large language models, contrasting fast intuitive (System 1) and slow deep (System 2) reasoning. It analyzes why System 2 is computationally costly yet more accurate, and why System 1 is efficient but less effective. The survey covers causes of inefficiency, patterns of reasoning behavior, and potential solutions to balance performance and computational budgets, offering actionable insights and an open‑source repository for ongoing research.

By Rui Wang, Hongru Wang, Boyang Xue, Jianhui Pang, Shudong Liu, Yi Chen, Jiahao Qiu, Derek Fai Wong, Heng Ji, Kam-Fai Wong
arXiv Computation and Language
Sep 23

ClusterFewshot: Improving Few-shot Optimization for LLMs workflow

ClusterFewshot is a new strategy for selecting few‑shot demonstrations in large language model workflows. It combines semantic structuring with utility‑aware scoring to build representative demonstration sets, improving accuracy over prior bootstrap‑based methods. In DSPy‑based pipelines, it substantially reduces optimization cost across multiple benchmarks while consistently outperforming earlier approaches in both standalone prompt tuning and hybrid prompt‑weight optimization.

By Omri Bar Haim, Shahar Katz, Lior Wolf