arXiv:2608.11318v2 Announce Type: replace-cross
Abstract: Many sequential construction tasks have exact terminal symmetries even though execution is directed and depends on history. Process evidence...
By Yi Liu
arXiv:2608. 04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software.
By Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.
By Pedro Santos
arXiv:2608. 14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update.
By Jiaxing Guo, Lei Ye
The paper introduces RefineCut, an open‑weight planner that edits a typed video timeline by applying structured patches for clip selection, trimming, ordering, transitions, music alignment, and duration. A deterministic verifier checks each patch against an explicit constraint ledger, and the planner is trained via verifier‑replayed distillation and a second evolutionary stage (RefineCut‑Evo) that uses the verifier and a task rubric to generate high‑margin preference pairs. On the RefineCut‑Bench dataset, the 8‑billion‑parameter planner improves from a Video‑Editing Score of 0.620 to 0.924, matching or exceeding its frontier teachers in a closed verifier loop, and the gains transfer to other large models such as Llama‑3.1‑8B and GLM‑4‑9B.
By Haoyu Wang, Cheng Feng, Liuyang Bian, Ruiyang Huang, Lei Wei, Yafei Wen, Xiaoxin Chen, Xiaoying Tang
The paper introduces Verifier-Certified Rule Transport (VCRT), a method that uses native verifiers to replay adjacent operation pairs and identify commutation certificates or anti-diamonds, thereby distinguishing true logical dependencies from mere serialization choices in reinforcement learning with verifiable rewards. VCRT assigns policy credit based on the total probability mass of each certified orbit and imposes constraints on post-swap consistency, source retention, and policy drift. In leave-one-environment-out transfer experiments across ProofWriter, CLRS, and Lean, VCRT achieves a 77.60% macro pass rate, outperforming the strongest baseline by 13.06 points, with the largest gains observed in Lean.
By Bang Xie, Hao Liu, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Senjian Zhang, Wei Chen