arXiv AI

Terminal Symmetry as a Carrier of Asymmetric Process Knowledge: Statewise Refinement for Anytime Verified Construction

arXiv AI
Aug 20

Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair

The study demonstrates that governance records—structured logs linking task contracts, model attempts, verifier decisions, and outputs—can serve as effective supervision for bounded AI models. Using a verifier-selected self‑training approach, the authors show that a Qwen3‑14B model trained on plans accepted by an independent VAL verifier achieved significant gains in plan acceptance across numerous PlanBench replanning cases, outperforming other selection strategies. The results highlight the feasibility of one‑shot execution and cumulative learning without relying on oracle targets or stronger teachers.

By Jesus Salas
arXiv Machine Learning
Aug 19

Elimination Geometry

The monograph introduces Elimination Geometry (EG), a typed, native‑loss, audit‑oriented framework that investigates when locally optimal objects can be realized by a shared deployment rule. EG examines how elimination and compression can erase distinctions needed for prediction, inference, control, or representation, and it separates local solvability, global realizability, and finite‑sample certifiability. The work synthesizes tools from geometry, optimization, information theory, statistics, and machine learning to address regular, coordination, singular, compositional, and resource‑limited mechanisms, and demonstrates applications in sparse model selection, distribution‑free prediction, observational treatment policies, routed expert and retrieval systems, and learned score fields.

By Mian Huang, Xueqin Wang
arXiv AI
Aug 14

TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems

arXiv:2608. 13221v1 Announce Type: new Abstract: The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and allocate reasoning resources--that is, how they organize search.

By Shunwen Bai, Ziping Ma, Chaoyang Zhang, Yarong Wang, Jiale Liu, Zhen Qin, Qingpei Guo
arXiv AI
Aug 7

Recursive Synthesis for Long-Horizon Terminal Tasks

arXiv:2608. 05466v1 Announce Type: new Abstract: High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutually consistent.

By Zhongzhi Li, Yucheng Shi, Zongxia Li, Ruhan Wang, Anhao Li, Zixun Huang, Junyao Yang, Lei Ke, Ninghao Liu, Haitao Mi, Leowei Liang