arXiv AI By Xinyi Ling, Ye Liu, Reza Averly, Xia Ning

Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation

Read the original on arXiv AI →

arXiv:2604. 03924v2 Announce Type: replace-cross Abstract: Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 10

SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design

SocialRL is a multi-turn reinforcement learning framework that refines large language models’ social intelligence. It uses PPO to propagate delayed outcome rewards across turns, enabling long‑horizon planning, and introduces six process reward dimensions—such as goal advancement and relational attunement—to capture the goal‑relationship trade‑off. A reward model provides fine‑grained scoring and a stage‑aware weight schedule prioritizes relationship building early, goal pursuit mid‑way, and balanced closure later, yielding an average 9.2 percentage‑point improvement in goal achievement across multiple social‑dialogue benchmarks.

By Jianing Wang, Xintao Wang, Aili Chen, Jie Shi, Hongcheng Guo, Jun Gao, Wenxuan Zhao, Chengkun Lang, Yuanli Guo, Yanghua Xiao
arXiv AI
Jun 2

CA-BED: Conversation-Aware Bayesian Experimental Design

arXiv:2606. 01182v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at static reasoning tasks, yet their performance often degrades in interactive scenarios where information must be actively acquired through questioning.

By Daniel Arnould, Rashad Aziz, Zixuan Kang, Tanav Changal, Kevin Zhu, Sunishchal Dev, Gabriel Grand, Shreyas Sunil Kulkarni
arXiv AI
Aug 28

STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration

The paper introduces STEP, a State‑Aware Task Estimator and Planner that uses multi‑modal large language models to explicitly estimate system states and predict state transitions during task planning. By forecasting future states alongside actions, STEP reduces hallucinated actions and improves task‑convergent planning. In a simulated robot assembly task, STEP outperforms the state‑of‑the‑art by 32.8% in action executability and 14.8% in final‑state error.

By Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba
arXiv AI
Sep 10

An Auditable Symbolic-RAG-Generative AI Architecture for Goal-Oriented Conversation Orchestration

The paper introduces a Symbolic-RAG-Generative architecture called GRACE for goal‑oriented conversational systems. GRACE transforms business intent into a fixed objective set and uses a constrained policy to update the conversation state based on visitor‑authored evidence while ensuring visitor utility. Evaluation on real‑estate and professional‑cleaning dialogues shows high accuracy in state transitions, evidence precision/recall, and monotonicity.

By Ramon Gonzalez (Mentomy AI), Antonio Diaz (Mentomy AI)