arXiv Computation and Language By Omri Bar Haim, Shahar Katz, Lior Wolf

ClusterFewshot: Improving Few-shot Optimization for LLMs workflow

Read the original on arXiv Computation and Language →

ClusterFewshot is a new strategy for selecting few‑shot demonstrations in large language model workflows. It combines semantic structuring with utility‑aware scoring to build representative demonstration sets, improving accuracy over prior bootstrap‑based methods. In DSPy‑based pipelines, it substantially reduces optimization cost across multiple benchmarks while consistently outperforming earlier approaches in both standalone prompt tuning and hybrid prompt‑weight optimization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Sep 4

Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings

Frontier large language models (LLMs) are examined as batch optimizers in both continuous and discrete settings. The study finds that while LLMs perform competitively in zero‑shot optimization of numerical test functions, their performance is less robust than classical non‑LLM methods. However, LLMs excel in semantically rich, discrete spaces that resemble their pretraining data, demonstrating strong batch optimization behavior in such contexts.

By Frank Hu, Shriram Chennakesavalu, David Graff
arXiv Computation and Language
Aug 27

InternBootcamp: Boosting LLM Reasoning with Verifiable Task Scaling

InternBootcamp is an open‑source framework that offers over 1,000 domain‑diverse task environments for large language model (LLM) reasoning research. It introduces Bootcamp‑Eval, an automatically generated benchmark for comprehensive performance assessment. Experiments show that training on InternBootcamp significantly improves reasoning performance, with a 32B model achieving state‑of‑the‑art results on Bootcamp‑Eval and other established benchmarks, demonstrating that scaling the number of training tasks yields consistent gains.

By Peiji Li, Jiasheng Ye, Yongkang Chen, Linyang Li, Yichuan Ma, Zijie Yu, Ganqu Cui, Haozhan Li, Jiacheng Chen, Chengqi Lyu, Wenwei Zhang, Qipeng Guo, Dahua Lin, Bowen Zhou, Kai Chen
arXiv AI
Jun 30

Scaling Textual Gradients via Sampling-Based Momentum

arXiv:2506. 00400v4 Announce Type: replace-cross Abstract: LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective method for automatic prompt engineering.

By Zixin Ding, Junyuan Hong, Zhan Shi, Jiachen T. Wang, Zinan Lin, Li Yin, Meng Liu, Zhangyang Wang, Yuxin Chen
arXiv AI
Sep 24

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.

By Hamed Taherkhani, Mohammad Abdollahi, Melika Sepidband, Hridya Dhulipala, Tien N. Nguyen, Hadi Hemmati