arXiv AI By Himel Dev, Madhusudan Basak, Tanmoy Sen, Paromita Shome, Bashima Islam

Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation

Read the original on arXiv AI →

arXiv:2607. 15562v1 Announce Type: cross Abstract: Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, item dependencies, and luggage limits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward

TripScore is a benchmark and evaluation framework for large language models (LLMs) in travel planning, built from real user logs and calibrated with 1,468 pairwise judgments from 203 travel experts. It uses a hierarchical feasibility gate for format and commonsense checks, and a unified point-wise reward that combines soft quality and preference fulfillment. Experiments show that reinforcement learning fine‑tuning, such as GRPO, consistently outperforms other methods when evaluated with TripScore.

By Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang
arXiv Computer Vision
Sep 7

First Things First: Teaching LLM-Based Agents to Prioritize Must-Haves before Nice-to-Haves

The paper introduces the problem of teaching multimodal large language models (MLLMs) to prioritize must‑have requirements over nice‑to‑have ones in autonomous agent tasks. It evaluates current MLLMs on 3,649 realistic service‑scenario problems and finds widespread failures, then proposes First Things First Reinforcement Learning (FTF‑rl) to explicitly optimize reasoning over multi‑priority user requirements, achieving significant performance gains and generalizing to logical and mathematical reasoning benchmarks.

By Tianjie Ju, Xinyue Xu, Wanxuan Sun, Lingxiao Diao, Gongshen Liu, Zhuosheng Zhang, Cheng Yang
arXiv AI
Aug 28

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 instances from a major Chinese travel platform, each with nearly 40 recorded behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leverages these trajectories, external retrieval tools, and internal memory, outperforming GPT‑4.1 and other baselines on the dataset.

By Zihao Cheng, Yingyu Shan, Hongru Wang, Zeming Liu, Xinyi Wang, Xiangrong Zhu, Yuhang Guo, Wei Lin, Yunhong Wang
Hugging Face Trending Papers
Aug 27

Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory

Behavior2Trip introduces a new task—Behavior‑Aware Travel Planning—where user preferences are inferred directly from past behavior trajectories rather than explicit instructions. The benchmark contains 11,400 Chinese travel‑planning instances, each with an average of 39.8 past behaviors across 14 attributes and 5 preference dimensions. A reinforcement‑learning agent, B2T‑Agent, leveraging behavior trajectories, external retrieval tools, and internal memory, outperforms strong baselines such as GPT‑4.1 on this challenging dataset.

arXiv AI
Jul 29

Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation

arXiv:2607. 25956v1 Announce Type: new Abstract: Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consistently matches heterogeneous instance-level regimes induced by demand concentration, inventory imbalance, replenishment scale, service constraints, and forecast volatility.

By Jintao Xu, Yingzheng Ma, Jiong Dong, Yongzhi Qi, Jianshen Zhang