arXiv Machine Learning

Mirror Horizon: Viable Path Entropy as a Measure of Bounded Reflection

arXiv:2607. 11937v1 Announce Type: new Abstract: Mirror Theory proposes that an intelligent system should be studied not only by what it represents, but by what coherent continuations it can sustain under repeated reflection.

arXiv AI
Aug 20

Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models

The paper introduces EvoResearcher, a training‑free, inference‑time protocol that enables a frozen large language model to perform cost‑bounded self‑reflection and early stopping. By iterating through generate → self‑critique → revise steps until a maximum depth or a CONFIRMED sentinel is reached, the model can self‑verify its answers within a strict compute budget. The protocol incorporates four self‑reflective meta‑reward components—correctness, efficiency, reflection depth, and tool‑call diversity—implemented as prompt‑level mechanisms, and is validated on Big‑Bench Hard, GSM8K, and MATH benchmarks, achieving comparable accuracy while terminating 82‑88% of items early with only about 2.1 generations per question.

By Wei Yu, Suxing Liu, Minjie Yu, Jiahao Wang, Zhijian Zheng, Haocheng Deng, Bing Li
arXiv AI
Sep 11

Grounded Continuation: A Linear-Time Runtime Verifier for LLM Conversations

Grounded Continuation introduces a runtime verifier that classifies each utterance in an LLM conversation into one of eight epistemic operations and uses a symbolic engine to maintain a dependency map of claims and their supports. The verifier checks whether a new continuation is grounded by walking this map, a linear-time process that requires no additional LLM calls. On benchmarks such as ReviseQA and MemoryAgentBench, the verifier improves single-hop accuracy for several QA models, even enabling a 7B model to outperform GPT‑4o when guided by the verifier.

By Qisong He, Jinwei Hu, Xinmiao Huang, Changshun Wu, Yi Dong, Xiaowei Huang
arXiv AI
Sep 18

Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs

The paper introduces a new evaluation protocol called checkpoint handoff to disentangle the contributions of reaching a target state and solving the task in reinforcement learning agents. By cloning states reached by one checkpoint and handing them to another without retraining, the authors separate the REACH metric (how often a policy arrives at a state confirmed to be a fixed number of actions from success) from the SOLVE metric (how often it finishes from that identical state). Across two benchmarks and pipelines, the analysis shows that RL history benefits RL solvers more than SFT solvers, and that independent REACH and SOLVE gaps predict overall performance.

By Xuan Liu, Jingbin Qian