arXiv AI By Chong Peng, Pin Qian, Su Wang, Yihang Chen, Varun Sah

BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL

Read the original on arXiv AI →

arXiv:2608. 02876v1 Announce Type: new Abstract: Tool-using agents do not merely consume observations: their actions determine what arrives next.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 17

DualSQL: Text-to-SQL with Multi-Agent Reinforcement Learning

DualSQL is a Text-to-SQL system that uses two agents sharing a single model backbone, enabling joint optimization via multi-agent reinforcement learning. The approach incorporates three database access tools for multi-step reasoning, rollout guardrails to stabilize training, and a new SQL correctness metric called robust execution match (REX). Trained on only 3,755 examples, DualSQL-4B reaches 68.0% execution accuracy on the BIRD dev set, while DualSQL-8B achieves 71.1%, surpassing prior state‑of‑the‑art single‑model solutions with 32B parameters.

By Shijie Chen, Yu Gan, Yeounoh Chung, Jiani Zhang, Quannan Li, Sravan Babu Bodapati, Cody J. Greer, Yu Su, Fatma Ozcan
arXiv AI
Sep 18

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

The paper investigates how agent harnesses—specifically planning guidance, execution organization, and completion verification—affect performance in retail and airline pilot tasks. By comparing fixed, task‑specific plans to shuffled policy text of equal length, the study finds that fixed plans improve success rates by about 7 percentage points, especially on complex tasks. A read‑only verifier rejects a majority of invalid episodes while incurring minimal cost, and its impact varies with the penalty for erroneous acceptance, often matching the full planning‑plus‑verification benefit at a lower cost.

By Yukun Zhang, Kemu Xu, Yishen Chen