arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
By Chenyu Zhou
The paper introduces PACT, a tuning method for single-token typed-decision models that leverages contrastive pair data to add four training terms—difference-in-differences margin, permutation-consistency, evidence-necessity, and ordinal transport cost—without requiring new annotations. PACT achieves comparable accuracy to existing recipes while reducing position bias and ordinal error, and it improves robustness and stability across seeds. The authors provide code, data splits, and trained adapters for reproducibility.
By Yida Lin
arXiv:2609.37647v1 Announce Type: cross
Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
By Tobias Deu{\ss}er, Lorenz Sparrenberg, Rafet Sifa
arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.
By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv:2608. 15445v1 Announce Type: new Abstract: When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure known as goal misgeneralization.
By Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
arXiv:2607. 04281v1 Announce Type: cross Abstract: Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar queries, but most existing methods do not model the time-varying freshness of open-web evidence.
By Muhammad Mansoor, Tahir Ahmad, Yeo-Chan Yoon
The paper introduces State‑Conditioned Minimal Sufficient Evidence Recovery (SER), a method that, given a coding agent’s current state, reconstructs a compact set of evidence passages that collectively provide all facts needed for the agent’s next decision. Using the SERBench dataset of 500 held‑out states from 45 repositories, the authors show that their MSS‑Complement approach recovers a complete evidence set for 73.0 % of states with five items and 80.6 % with eight, outperforming baseline ranking methods. The study also demonstrates that this set‑level policy improves downstream performance on AMA‑Bench and highlights the importance of retrieving missing facts rather than merely re‑ranking similar passages.
By Zhexi Feng, Ruiyi Zhang, Yongbo Yang, Pengtao Xie
arXiv:2607. 02104v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used as cheap, scalable judges that compare candidate outputs pairwise.
By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv:2608. 07813v1 Announce Type: new Abstract: An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships.
By Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen
The paper proposes a method for selectively querying language‑model advice in reinforcement learning by predicting the value of potential responses and only querying when the expected benefit outweighs the cost. It introduces a certified, response‑contingent metareasoning framework that guarantees near‑optimal advice usage under certain assumptions, and demonstrates that a calibrated controller with Qwen2.5 advisors can improve task performance while drastically reducing the number of advice calls on the BabyAI benchmark.
By Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan
arXiv:2607. 19453v1 Announce Type: cross Abstract: We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positive Binance Spot paper policies after assumed costs.
By Ayoub Jadouli
Agentic systems have widened the gap between producing candidate outputs and reviewing them. This paper asks a practical architectural question: should domain specialization be built into an evaluator's weights, or into the rule that decides when its judgment can be trusted?