Optimization story: Bloom inference
Related stories
Making thousands of open LLMs bloom in the Vertex AI Model Garden
Fast Inference on Large Language Models: BLOOMZ on Habana Gaudi2 Accelerator
Beyond the Node: Clade-level Selection for Efficient MCTS in Automatic Heuristic Design
arXiv:2602. 00549v2 Announce Type: replace Abstract: While Monte Carlo Tree Search (MCTS) shows promise in Large Language Model (LLM) based Automatic Heuristic Design (AHD), it suffers from a critical over-exploitation tendency under the limited computational budgets required for heuristic evaluation.
Scientific discovery as meta-optimization: a combinatorial optimization case study
Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluation criterion based on quality, novelty, and validity. Large language models (LLMs) have enabled automated exploration of this space, but we argue that simultaneous modification of the evaluation criteria is equally important.
Scientific discovery as meta-optimization: a combinatorial optimization case study
arXiv:2606. 26728v1 Announce Type: new Abstract: Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluation criterion based on quality, novelty, and validity.
DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel. Diffusion-based drafters further reduce proposal latency by predicting an entire token block in parallel, but their position-wise distributions are marginal rather than conditioned on tokens selected along each draft path.
Principle-Evolvable Scientific Discovery via Uncertainty Minimization
arXiv:2602. 06448v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based scientific agents have accelerated scientific discovery, yet they often suffer from significant inefficiencies due to adherence to fixed initial priors.
DARTree: Speculative Diffusion Decoding with Autoregressive Draft Trees
arXiv:2608. 13524v1 Announce Type: new Abstract: Speculative decoding losslessly accelerates autoregressive language models by verifying multiple draft tokens in parallel.
An In-depth Study of LLM Contributions to the Bin Packing Problem
arXiv:2510. 27353v2 Announce Type: replace Abstract: Recent studies have suggested that Large Language Models (LLMs) could provide interesting ideas contributing to mathematical discovery.
Importance-Aware Scheduling for High-Dimensional Hyperparameter Optimization
arXiv:2606. 10068v1 Announce Type: cross Abstract: Hyperparameter Optimization (HPO) is essential for building high-performing ML/DL models, yet conventional optimizers often struggle in high-dimensional spaces where evaluations are costly and progress is diluted across many low-impact variables.
The Program Is Still There: A Conservation Law for Program Discovery
arXiv:2606. 13799v1 Announce Type: cross Abstract: Finding the shortest program that generates a sequence is uncomputable, and for six decades that fact has been mistaken for a wall around finding any generating program.