arXiv Computation and Language

Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.

arXiv Machine Learning
5d ago

Benchmarking System One decision models against trained classifiers and language models for automated decision gates

The paper evaluates System One decision models—typed models that output probabilities for branching decisions—against supervised classifiers and generative language models on automated decision gate tasks. Eight checkpoints from six families, including the hosted model Jev, were benchmarked on workflow, intent, and social‑science items, showing that small trained classifiers match or slightly outperform decision models on intent and workflow when labels are available, while decision models outperform zero‑shot classifiers when labels are absent. The study also explores calibration, risk thresholds, and cost‑efficiency trade‑offs, providing condition‑dependent design guidelines for automated decision gates.

By Amir Rafe, Subasish Das
arXiv AI
3d ago

Keeping JEPA World Models Plannable When Little of the Frame Moves

The paper introduces SLIM, a pushing benchmark that tests language‑based goal specification for latent world models. It shows that a vanilla LeWM fails almost entirely due to an action‑insensitive encoder, but adding an inverse‑dynamics auxiliary loss restores encoder sensitivity and dramatically improves success rates. With the repaired latent, a lightweight language‑goal head can plan from natural language sentences, achieving high performance on navigation and pushing tasks without retraining the world model.

By Florian Strohm, Patrick Wagner, Jannik Schwab, Marco Huber
arXiv AI
Aug 28

Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

The paper evaluates a manager‑worker scaffold that uses a shared filesystem workspace to orchestrate multi‑agent large language model (LLM) coding tasks without training or tuning. Across nine models—including five open‑weight and four closed‑weight systems—the scaffold yields statistically significant accuracy gains for some models (e.g., Qwen3.8‑27B, GPT‑5.6‑Luna, GPT‑5.6‑Terra, Kimi‑K3, Minimax‑M3) while producing null or negative effects for others (e.g., Qwen3.6‑35B). The study shows that the manager can triple token usage but still achieves higher accuracy at a fraction of the cost compared to larger single‑pass models, with key mechanisms identified as context management and problem decomposition.

By Victor Gao (Sang Won), Vida Khosrowshahi (Sang Won), Ali Khosrowshahi (Sang Won), Xihao Sun (Sang Won), Juhyun Lee (Sang Won), Simon (Sang Won), Lee