arXiv AI

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

arXiv:2607. 15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit.

arXiv AI
4d ago

Neuro-Symbolic Computer Use: Learning Reusable Policies for Reliable and Efficient Execution

The paper introduces neuro‑symbolic computer use, a method that learns reusable policies to execute recurring computer workflows efficiently. Instead of re‑planning each run, the learned policy encodes stable decisions (ordering, variables, loops, branches) into executable code while delegating observation‑dependent decisions to neural models. Using neuro‑symbolic policy iteration, the approach iteratively refines the policy from a single agent trajectory, diagnoses failures, and revises the code with a coding model, achieving superior Pass^3 scores and significant reductions in per‑run cost and latency on OSWorld‑Verified and ScienceBoard benchmarks.

By Hyewon Suh, Thanh Minh Nguyen, Chih-Lun Lee, Darrow Hartman, Lizhao Liu, Xin Eric Wang, Ang Li, Jiachen Yang
arXiv AI
Sep 11

Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward

The paper introduces the concept of proof‑carrying cognition, aiming to close the verification gap in language‑model reasoning by using reality‑settled rewards. It presents a theoretical framework linking verifier‑gold correlation to compute‑capability trade‑offs, demonstrates that unsound verifiers degrade under best‑of‑N selection while sound verifiers improve, and proposes a new benchmark metric, Soundness‑under‑Pressure, for evaluating reality‑settled reasoning systems.

By Eshwar Reddy M, Sourav Karmakar
arXiv AI
Aug 26

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

The paper introduces a Bayesian self‑escalation strategy for hierarchical large‑language‑model agents, allowing an agent to detect during its own reasoning that it is unlikely to succeed and hand control over to a stronger model. The authors formalise this as an optimal‑stopping problem over a learned competence posterior, derive a myopic escalation threshold, and prove that the optimal policy is a time‑varying threshold without assumptions on the raw signal. They provide theoretical guarantees—including a 1/√n regret decay with n calibration trajectories—and validate the approach in simulations and a real‑model code‑generation cascade, showing that the escalation frontier outperforms post‑hoc routing at equal cost. whyItMatters":"The study offers a principled, theoretically grounded method for agents to dynamically decide when to seek stronger models, potentially improving efficiency and reliability in hierarchical LLM systems."

By Nadeem Shaikh
arXiv Machine Learning
Aug 11

Finite Constant Frontiers and Auditable Regret Certificates for Average-Reward Reinforcement Learning

arXiv:2608. 07725v1 Announce Type: new Abstract: Average-reward reinforcement-learning regret is known up to logarithmic factors, but the numerical content of published guarantees is difficult to compare because probability mode, structural parameter, logarithmic normalization, prior information, and planning assumptions differ.

By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
arXiv AI
Sep 15

Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization

The paper introduces MACCHIATO, a training algorithm that builds a ReLU‑MLP from partial truth‑table data while simultaneously constructing an explicit Boolean circuit over AND, OR, and XOR gates that certifies the network’s computation. The method iteratively projects residuals onto low‑dimensional Boolean classes, compiles the resulting circuit into a ReLU‑MLP, and uses logic minimization and influence‑based variable selection to achieve a six‑layer network with provable truth‑table error bounds. Experiments on synthetic random‑junta tasks show that these certified networks outperform Adam‑trained MLPs in data‑sparse or projection‑aligned regimes and complete faster than flat ESPRESSO in certain settings.

By Hrad Ghoukasian, Anastasis Kratsios
arXiv AI
Jun 17

Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs

arXiv:2606. 17735v1 Announce Type: new Abstract: Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (LLMs), it often remains vulnerable to the autoregressive curse in long-horizon logical reasoning: small epistemic perturbations introduced early in generation can propagate irreversibly along the Markov decision process flow, triggering cascading failures that drive the reasoning trajectory toward collapse.

By Ziliang Wang, Kang An, Faqiang Qian, Jialu Cai, Cijun Ouyang, Yuhang Wang, Qibing Ren, Yichao Wu
arXiv Computation and Language
6d ago

Recursive Self-Improvement via On-Policy Distillation for Reasoning

The paper introduces a recursive self-improvement framework for language models that replaces an external teacher with a frozen copy of the student, enabling dynamic co-evolution (DCE) and self-refined concise learning (SRCL). DCE allows the privileged teacher to evolve alongside the student, while SRCL trains on shorter, verified rewrites to reduce verbosity. Experiments show that the combined DCE+SRCL approach outperforms traditional on‑policy self‑distillation across multiple model sizes and math benchmarks, achieving significant accuracy gains and shorter outputs.

By Shangjian Yin, Zehao Zhao, Kavosh Asadi, Rui Liu, Yuchen Lu, Shike Mei, Hang Cui, Luke Simon, Zhouxing Shi, Hamed Firooz
arXiv Machine Learning
Sep 11

Topological Necessities: Mechanism-Invariant Strategic Subgoals for Cross-Embodiment Goal-Conditioned Control

The paper introduces topological necessities—mechanism‑invariant subgoals derived from the topology of successful trajectories—used to guide long‑horizon goal‑conditioned reinforcement learning. By computing homology in dimensions 0 and 1 over a transport‑weighted carrier, the authors obtain an enumerable gate set that forms a recursive topological gate hierarchy. These certified gates transfer across different embodiments (e.g., from PointMaze to Ant and Humanoid) without retraining, achieving state‑of‑the‑art performance on several benchmark tasks.

By Hao Shi, Xi Li