arXiv AI By Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran

HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Aug 27

Beam Search, Self-Consistency, and the Limits of Inference-Time Scaling for Grammar-Constrained Text-to-SQL in Small Language Models

The paper investigates how increasing inference-time computation—via wider beam search or sample‑plus‑vote—affects performance on grammar‑constrained text‑to‑SQL tasks for small language models. Using the Qwen2.5‑Instruct family (0.5B–7B parameters) on the Spider benchmark, the authors find that larger models consistently outperform higher inference compute on the same model size, and that beam search yields better accuracy than sample‑plus‑vote under matched budgets. These results suggest that, unlike unconstrained settings, scaling inference compute does not compensate for smaller model size when strict grammar constraints are applied.

By Ty Chermsirivatana, John MacCormick
arXiv AI
2d ago

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

The paper introduces Galahad, a memory layer that stores a transformer language model’s key‑value state for blocks of text, allowing subsequent requests to reuse previously computed attention rather than recomputing it. On seven real‑world datasets, 98.7% of prompt tokens were already read, and with Galahad the model could attend to an entire 97,000‑token corpus, achieving 98–100% recall on a 100‑fact test while reducing inference time and energy consumption dramatically. The approach was validated across 30 models and all runtimes, demonstrating that stateful inference can replace stateless serving without loss of accuracy.

By Sietse Schelpe