StreamDecisionBench: Evaluating Decisions in Force on Evolving Language Streams
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
The paper investigates how test‑time computation can enhance language models and at what cost, introducing the SELF‑POT benchmark to evaluate this across competition mathematics, competitive programming, and agentic workflows. SELF‑POT separates candidate coverage from final accuracy, tracks correctness transitions under revision, and measures protocol completion alongside task success. Using a unified budget rule, the study compares direct inference, parallel sampling, and self‑revision across five low‑cost reasoning models, revealing that selection rules and failure handling significantly influence gains and cost savings.
arXiv:2603. 28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs.
FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.
The paper introduces a coroutine-bridge harness that lets a language model emit a Python program to manage tool calls in the CAR-bench evaluation. By decoupling model invocations from tool round-trips, the approach reduces model calls to a median of two per task while maintaining seven agent turns, achieving a median latency of 1.8 s on a Cerebras gpt‑oss‑120b. The harness achieved 60.0 % Pass³ on the official hidden evaluation, outperforming the baseline by 4.5× and matching frontier-model agents on GPT‑5.5, all while keeping the prompt largely cached and minimizing input compute.
arXiv:2609.37647v1 Announce Type: cross Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...