arXiv AI

Can a Cacheable Decision Model Follow Rules?

The paper evaluates Certo, a small non‑generative decision model that scores candidate actions based on text. It compares a joint scorer that processes state, rules, and candidates together with a cacheable encoder that pre‑encodes candidates to reduce cost. Experiments show the cacheable approach loses rule sensitivity, while targeted counterfactual supervision can recover performance on synthetic tasks; however, on real rules the joint scorer still outperforms the cacheable version, and cross‑domain mixtures do not improve accuracy.

arXiv AI
3d ago

PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

The paper introduces PACT, a tuning method for single-token typed-decision models that leverages contrastive pair data to add four training terms—difference-in-differences margin, permutation-consistency, evidence-necessity, and ordinal transport cost—without requiring new annotations. PACT achieves comparable accuracy to existing recipes while reducing position bias and ordinal error, and it improves robustness and stability across seeds. The authors provide code, data splits, and trained adapters for reproducibility.

By Yida Lin
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv AI
Sep 18

The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents

The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.

By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie
arXiv Machine Learning
Sep 25

Certified Predictive Value-of-Advice Gating for Cost-Aware Language-Model Guidance in Reinforcement Learning

The paper proposes a method for selectively querying language‑model advice in reinforcement learning by predicting the value of potential responses and only querying when the expected benefit outweighs the cost. It introduces a certified, response‑contingent metareasoning framework that guarantees near‑optimal advice usage under certain assumptions, and demonstrates that a calibrated controller with Qwen2.5 advisors can improve task performance while drastically reducing the number of advice calls on the BabyAI benchmark.

By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan