arXiv:2608. 11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules.
By Rahul Nair, Bastian Lipka, Elizabeth Daly
TripScore is a benchmark and evaluation framework for large language models (LLMs) in travel planning, built from real user logs and calibrated with 1,468 pairwise judgments from 203 travel experts. It uses a hierarchical feasibility gate for format and commonsense checks, and a unified point-wise reward that combines soft quality and preference fulfillment. Experiments show that reinforcement learning fine‑tuning, such as GRPO, consistently outperforms other methods when evaluated with TripScore.
By Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang
arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.
By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
arXiv:2608.30224v1 Announce Type: cross
Abstract: Large Language Models (LLMs) are increasingly used to annotate structured product data in e-commerce, but early deployment often begins as a cold-sta...
By Cheng Lyu, Jingyue Zhang, Vinny DeGenova, Mengwei Li, Yuanli Pei
In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations.
arXiv:2607. 14318v1 Announce Type: new Abstract: We introduce COAT (Counterfactual Optimal Action Tree), a framework for learning interpretable prescriptive policies from observational data.
By Youssef Drissi, Markus Ettl, Shivaram Subramanian, Wei Sun, Zack Xue
arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.
By Libin Qiu, Yuhang Ye, Zhirong Gao, Xide Zou, Junfu Chen, Ziming Gui, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Kun Zhao
arXiv:2606. 18519v1 Announce Type: cross Abstract: Though robotic systems are now being commercialized and deployed in various industries, many of these systems are highly specialized and often require an advanced skill set to operate and ensure they perform as instructed.
By Marcos Abel Zuzu\'arregui, Stefano Carpin
arXiv:2509. 21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience.
By Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
arXiv:2608.30399v1 Announce Type: cross
Abstract: Large language models (LLMs) exhibit strong semantic reasoning and open-ended generation abilities, but aligning these abilities with structured sequ...
By Yunqi Liu, Yang Zhang, Ruixing Zhang, Liangzhe Han, Yi Qiao, Tongyu Zhu, Leilei Sun
RuleWeaver is a benchmark construction framework designed to evaluate large language models’ ability to reason over complex, rule‑centered scenarios. It begins with corpus‑derived IF‑THEN meta rules, expands them into more intricate rules, and composes these into scenario‑based QA instances. The benchmark assesses not only final answer correctness but also process‑level metrics such as rubric‑based answer quality, rule recall, and rule precision, revealing that current LLMs achieve only about 50% of the maximum rubric score on these tasks.
By Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu
arXiv:2608. 12674v1 Announce Type: new Abstract: Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers.
By Ravi Teja Chunduri, Srikaran Reddy Boya, Deep Narayan Mishra, Ajay Kumar B, Karthik Kumaran, Pranay Kona