arXiv:2608.11318v2 Announce Type: replace-cross
Abstract: Many sequential construction tasks have exact terminal symmetries even though execution is directed and depends on history. Process evidence...
By Yi Liu
arXiv:2608. 04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable software.
By Ruitong Li, Binjie Guo, Aisheng Mo, Guowei Su, Han Wang, Jie Li, Ru Zhang
The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.
By Pedro Santos
arXiv:2608. 14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update.
By Jiaxing Guo, Lei Ye
The paper introduces RefineCut, an open‑weight planner that edits a typed video timeline by applying structured patches for clip selection, trimming, ordering, transitions, music alignment, and duration. A deterministic verifier checks each patch against an explicit constraint ledger, and the planner is trained via verifier‑replayed distillation and a second evolutionary stage (RefineCut‑Evo) that uses the verifier and a task rubric to generate high‑margin preference pairs. On the RefineCut‑Bench dataset, the 8‑billion‑parameter planner improves from a Video‑Editing Score of 0.620 to 0.924, matching or exceeding its frontier teachers in a closed verifier loop, and the gains transfer to other large models such as Llama‑3.1‑8B and GLM‑4‑9B.
By Haoyu Wang, Cheng Feng, Liuyang Bian, Ruiyang Huang, Lei Wei, Yafei Wen, Xiaoxin Chen, Xiaoying Tang
The paper introduces Verifier-Certified Rule Transport (VCRT), a method that uses native verifiers to replay adjacent operation pairs and identify commutation certificates or anti-diamonds, thereby distinguishing true logical dependencies from mere serialization choices in reinforcement learning with verifiable rewards. VCRT assigns policy credit based on the total probability mass of each certified orbit and imposes constraints on post-swap consistency, source retention, and policy drift. In leave-one-environment-out transfer experiments across ProofWriter, CLRS, and Lean, VCRT achieves a 77.60% macro pass rate, outperforming the strongest baseline by 13.06 points, with the largest gains observed in Lean.
By Bang Xie, Hao Liu, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Senjian Zhang, Wei Chen
UnitBoost proposes a non‑generative merge operator to manage compound LLM systems, replacing the opaque higher‑level LLM that traditionally coordinates worker outputs. The operator maps worker outputs to slot‑value proposals, uses a constrained argmax to assemble the final answer, and explicitly tracks unfilled slots as residuals for subsequent rounds, thereby achieving order‑free processing and unit provenance. Across three held‑out benchmarks, UnitBoost outperforms both gold‑label‑selected candidates and input‑matched generative managers, improving compound‑system performance by up to 0.182 points and raising FanOutQA cell F1 from 0.4778 to 0.5524.
By Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang, Peiyang He
The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.
By Tanay Chowdhury, Saeideh Shahrokh Esfahani
arXiv:2607. 18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes?
By Jan Kirin
The paper presents a block‑wise differentiable Sinkhorn attention mechanism designed for long‑context balanced entropic optimal transport on TPU hardware. By stopping a $T$‑step Sinkhorn solve and unrolling a short refinement tail, the authors derive an exact surrogate gradient that achieves efficient block‑wise cost and memory usage. Experimental results on synthetic masked problems and a Pfam protein‑family screen demonstrate high numerical accuracy and sustained throughput on TPU v6e‑8, with notable improvements in reconstruction and sparse cross‑entropy metrics.
By Dylan Forde
arXiv:2608. 02680v1 Announce Type: cross Abstract: Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure with retries, exploration, accidental ordering, and repeated lookups.
By Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies)
arXiv:2607. 17240v1 Announce Type: new Abstract: When does a committed intermediate stage in an LLM reasoning pipeline earn its cost?
By Honglin Li (ShanghaiTech University)