arXiv AI

Repair Before Veto: Repair-Augmented Constraint Learning for Contextual Decisions

arXiv:2606. 02326v1 Announce Type: new Abstract: Hard constraints are usually treated as terminal vetoes: once a candidate violates a requirement, the learned rule rejects it and any repair is handled outside the decision semantics.

arXiv AI
2d ago

ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair

ContractRL introduces a contract-constrained sequential repair protocol for structured tool calls, modeling verifier-guided JSON repair as a bounded decision process. The policy observes candidate data, verifier feedback, JSON pointers, repair history, and budget, using a contract-derived action mask to filter invalid operations before a deterministic validator applies changes. Compared to Patch‑SFT and full regeneration, ContractRL achieves higher semantic success (0.9362 vs. 0.9076 and 0.9148) while generating fewer tokens (34.4 vs. 44.9 and 137.2), and policy optimization further improves success rates.

By Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Yina Sa, Daren Zha, Jun Xiao
arXiv AI
Sep 25

Auditability Is Not One Property: Rule Overlap, Behavioural Agreement, and Composition in Reinforcement Learning

The paper introduces a protocol for auditing and composing reinforcement‑learning policies using discrete behavioral rules, defining auditability through six testable predicates such as trace integrity and rule coverage. Experiments show that overlapping rule sets do not guarantee behavioral agreement, and that rule‑based fusion often fails to outperform value‑based composition, highlighting limitations in current description layers. The authors provide an evidence‑bounded audit framework and outline future directions for more robust skill composition.

By Liu Hung Ming
Hugging Face Trending Papers
Jun 19

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.