arXiv Computation and Language

UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model

UnitBoost proposes a non‑generative merge operator to manage compound LLM systems, replacing the opaque higher‑level LLM that traditionally coordinates worker outputs. The operator maps worker outputs to slot‑value proposals, uses a constrained argmax to assemble the final answer, and explicitly tracks unfilled slots as residuals for subsequent rounds, thereby achieving order‑free processing and unit provenance. Across three held‑out benchmarks, UnitBoost outperforms both gold‑label‑selected candidates and input‑matched generative managers, improving compound‑system performance by up to 0.182 points and raising FanOutQA cell F1 from 0.4778 to 0.5524.

arXiv AI
Aug 25

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

The study evaluates how to best split tasks among large‑language‑model agents for cross‑border VAT determination, comparing one broad agent to configurations ranging from one to five narrow agents. Across 4,400 runs—including token‑matched and failure‑injection scenarios—the intermediate configurations achieved the highest accuracy but did not surpass the fine‑endpoint benchmark, leaving the optimal decomposition hypothesis unconfirmed. The pilot provides a preregistered heuristic for right‑sizing decomposition, along with an oracle, dataset, and analysis pipeline.

By Pedro Santos
arXiv AI
Sep 7

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

SiLR introduces a structure‑preserving admission and process reward mechanism for large language model (LLM) tool agents. Unlike traditional scalar‑score gates that can trap agents in plateau trajectories, SiLR shadow‑executes each proposal and admits it based on a product order over branch‑level violation states, ensuring safe and recoverable actions. Experiments on Gym‑ANM and CityLearn benchmarks show SiLR consistently recovers all multi‑action episodes and outperforms scalar gates, while also providing a robust reward signal for policy learning.

By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou