arXiv Computation and Language

Natural-Language Policies to Executable Decisions: An Interpretable Large Language Model Framework

The paper introduces a production‑grade large language model (LLM) pricing system for the tourism industry that separates structured extraction and policy selection from deterministic numeric pricing. Policies are compiled into interpretable condition trees, allowing new clauses and evolving rules to be added without code changes while maintaining auditability. Deployed across 12 business categories and 1,500 operators, the system handled 3,960 orders in six months, cutting the order‑management team from 15‑20 to 3 and reducing per‑order handling time from 10 minutes to under 2 minutes.

arXiv AI
Aug 13

Policy-as-logic for robust reasoning over rules

arXiv:2608. 11905v1 Announce Type: new Abstract: In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules.

By Rahul Nair, Bastian Lipka, Elizabeth Daly
arXiv AI
Sep 18

TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward

TripScore is a benchmark and evaluation framework for large language models (LLMs) in travel planning, built from real user logs and calibrated with 1,468 pairwise judgments from 203 travel experts. It uses a hierarchical feasibility gate for format and commonsense checks, and a unified point-wise reward that combines soft quality and preference fulfillment. Experiments show that reinforcement learning fine‑tuning, such as GRPO, consistently outperforms other methods when evaluated with TripScore.

By Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang
arXiv AI
Jun 2

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

arXiv:2606. 01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.

By Weiyi Chen, Shuaixiong Wang, Ziyun Gao, Kaichun Hu, Wangze Ni, Shimin Di, Chen Jason Zhang, Lei Chen
Hugging Face Trending Papers
Aug 12

Policy-as-logic for robust reasoning over rules

In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries must respect written policies or rules. We present a hybrid symbolic approach that expresses policies in formal logic and at inference time exploits the representation power of language models for fact extraction to ground predicates, and an answer set solver for reasoning such that responses are interpretable, auditable, and as we show, accurate and robust under input perturbations.

arXiv AI
Jun 17

Blueprint First, Model Second: A Framework for Deterministic LLM Workflow

arXiv:2508. 02721v2 Announce Type: replace-cross Abstract: While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational environments where procedural fidelity and predictable execution are strict requirements.

By Libin Qiu, Yuhang Ye, Zhirong Gao, Xide Zou, Junfu Chen, Ziming Gui, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Kun Zhao
arXiv AI
Jul 15

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents

arXiv:2509. 21842v2 Announce Type: replace Abstract: Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary generation, ensuring an enjoyable user experience.

By Yansong Ning, Rui Liu, Jun Wang, Kai Chen, Wei Li, Jun Fang, Kan Zheng, Naiqiang Tan, Hao Liu
arXiv Computation and Language
Aug 28

RuleWeaver: Benchmarking Rule-Centered Scenario Reasoning for Large Language Models

RuleWeaver is a benchmark construction framework designed to evaluate large language models’ ability to reason over complex, rule‑centered scenarios. It begins with corpus‑derived IF‑THEN meta rules, expands them into more intricate rules, and composes these into scenario‑based QA instances. The benchmark assesses not only final answer correctness but also process‑level metrics such as rubric‑based answer quality, rule recall, and rule precision, revealing that current LLMs achieve only about 50% of the maximum rubric score on these tasks.

By Bohan Yu, Shi-Yang Li, Pengfei Cao, Jun Zhao, Kang Liu