arXiv Computation and Language By Shuqing Shi, Ziyan Wang, Milind Tambe, Yali Du

Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

Read the original on arXiv Computation and Language →

The paper introduces HARP, a framework for orchestrating heterogeneous agents with hidden preferences by maintaining numeric posterior beliefs updated via Bayes’ rule, rather than embedding beliefs in prompts. HARP achieves ∼O(√K) Bayesian regret and, with the HARP+ variant, adds a bonus for informative actions to keep inference active even when optimal actions are uninformative. Experiments across three problem settings show HARP+ outperforms other non‑oracle methods in scenarios where explicit joint inference is infeasible.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
6d ago

Robust Is Salient: An Informed Adversary Moves the Optimal Signal onto the Salience Pole

The paper investigates how an informed adversary can influence the optimal signal in a constrained signalling channel. It finds that the adversary‑robust optimum aligns with the salience pole on most items, differing only on a small subset where the salience‑to‑Bayes coordinate is undefined. The study demonstrates that as the adversary’s persuasion budget increases, the optimal signal shifts from a posterior‑maximizing to a margin‑maximizing strategy, and provides a diagnostic check for evaluating adversary‑awareness.

By Cris Huynh
arXiv AI
Jun 2

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

arXiv:2606. 01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO).

By Xiwen Chen, Wenhui Zhu, Jingjing Wang, Peijie Qiu, Zhipeng Wang, Huayu Li, ZhengXiao He, Xuanzhao Dong, Prayag Tiwari, Mingkun Xu, Yujian Xiong, Feng Luo, Abolfazl Razi, Brendan Hogan Rappazzo, Anderson Schneider, Yuriy Nevmyvaka
Hugging Face Trending Papers
Jun 19

Training the Orchestrator: A Supervised Approach to End-to-End PDDL Planning with LLM Agents

Translating natural-language planning intent into verified plans is a longstanding challenge: people communicate goals in language, while classical planners require formal PDDL specifications. Recent agentic frameworks bridge this gap by orchestrating a pool of specialized repair agents inside a verifier-checked refinement loop, but the orchestrator at the centre is itself a prompted frontier LLM, paying a frontier-LLM API call at every refinement step.