arXiv AI

Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction

arXiv:2608. 11318v1 Announce Type: cross Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent.

arXiv AI
Aug 25

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.

By Pedro Santos
arXiv Computation and Language
Aug 27

Plans You Can Check: Verifier-Grounded Learning of an Open-Weight Planner for Executable Video-Editing

The paper introduces RefineCut, an open‑weight planner that edits a typed video timeline by applying structured patches for clip selection, trimming, ordering, transitions, music alignment, and duration. A deterministic verifier checks each patch against an explicit constraint ledger, and the planner is trained via verifier‑replayed distillation and a second evolutionary stage (RefineCut‑Evo) that uses the verifier and a task rubric to generate high‑margin preference pairs. On the RefineCut‑Bench dataset, the 8‑billion‑parameter planner improves from a Video‑Editing Score of 0.620 to 0.924, matching or exceeding its frontier teachers in a closed verifier loop, and the gains transfer to other large models such as Llama‑3.1‑8B and GLM‑4‑9B.

By Haoyu Wang, Cheng Feng, Liuyang Bian, Ruiyang Huang, Lei Wei, Yafei Wen, Xiaoxin Chen, Xiaoying Tang
arXiv Machine Learning
Sep 24

From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization

The paper introduces Verifier-Certified Rule Transport (VCRT), a method that uses native verifiers to replay adjacent operation pairs and identify commutation certificates or anti-diamonds, thereby distinguishing true logical dependencies from mere serialization choices in reinforcement learning with verifiable rewards. VCRT assigns policy credit based on the total probability mass of each certified orbit and imposes constraints on post-swap consistency, source retention, and policy drift. In leave-one-environment-out transfer experiments across ProofWriter, CLRS, and Lean, VCRT achieves a 77.60% macro pass rate, outperforming the strongest baseline by 13.06 points, with the largest gains observed in Lean.

By Bang Xie, Hao Liu, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Senjian Zhang, Wei Chen
arXiv Computation and Language
Sep 10

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

UnitBoost proposes a non‑generative merge operator to manage compound LLM systems, replacing the opaque higher‑level LLM that traditionally coordinates worker outputs. The operator maps worker outputs to slot‑value proposals, uses a constrained argmax to assemble the final answer, and explicitly tracks unfilled slots as residuals for subsequent rounds, thereby achieving order‑free processing and unit provenance. Across three held‑out benchmarks, UnitBoost outperforms both gold‑label‑selected candidates and input‑matched generative managers, improving compound‑system performance by up to 0.182 points and raising FanOutQA cell F1 from 0.4778 to 0.5524.

By Xing Zhang, Guanghui Wang, Yanwei Cui, Mengdie Flora Wang, Peiyang He
arXiv AI
Aug 20

Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson

The paper presents a cost‑effective approach for industrial explainable‑recommendation systems by decoupling explanation generation from selection. Candidate explanations are pre‑generated using six prompt styles and two commodity LLMs, then a lightweight CPU‑resident selector (e.g., LambdaRank) chooses the best one at request time, achieving sub‑100 ms latency without GPUs. Experiments on a 2,958‑pair Google Local subset and a 300‑pair MovieLens‑1M split show that pairwise ranking methods outperform single‑action RL baselines, while KG‑path selectors achieve near‑perfect user satisfaction scores.

By Tanay Chowdhury, Saeideh Shahrokh Esfahani
arXiv Machine Learning
Sep 10

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge

The paper presents a block‑wise differentiable Sinkhorn attention mechanism designed for long‑context balanced entropic optimal transport on TPU hardware. By stopping a $T$‑step Sinkhorn solve and unrolling a short refinement tail, the authors derive an exact surrogate gradient that achieves efficient block‑wise cost and memory usage. Experimental results on synthetic masked problems and a Pfam protein‑family screen demonstrate high numerical accuracy and sustained throughput on TPU v6e‑8, with notable improvements in reconstruction and sparse cross‑entropy metrics.

By Dylan Forde