arXiv AI By Linzhang Li, Yixin Dong, Guanjie Wang, Ziyi Xu, Alexander Jiang, Tianqi Chen

XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs

Read the original on arXiv AI →

arXiv:2601. 04426v4 Announce Type: replace Abstract: Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

ContextPipe: Database-Inspired Context Assembly for Long-Horizon Agents

ContextPipe is a database-inspired framework for assembling context in long-horizon large language model agents. It treats context assembly like relational query execution, using a five-phase pipeline—Plan, Bind, Optimize, Execute, Feedback—backed by a structured catalog, deterministic cache-aware optimizer, and EXPLAIN ANALYZE tracing. In a preliminary evaluation on the SWE-bench Pro Qutebrowser subset, ContextPipe reduced token volume by 31%, LLM calls by 23%, and response time by 9% compared to an append-only policy, though it lowered KV cache-hit ratio.

By Peng Xu, Zuyu Zhang, Yuze Sun, Feng Tian, Long Wang, Chen Zhang
arXiv AI
Jun 9

Harmonia: End-to-End RAG Serving Optimization

arXiv:2505. 07833v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) improves the reliability of large language models by integrating external knowledge, but serving RAG pipelines efficiently is challenging because requests traverse heterogeneous components spanning LLM inference, databases, and CPU-side processing.

By Saurabh Agarwal, Bodun Hu, Luis Pabon, Myungjin Lee, Jayanth Srinivasa, Aditya Akella
arXiv AI
Jul 22

Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments

arXiv:2607. 18357v1 Announce Type: cross Abstract: Large language models now write a growing share of the world's code, increasingly inside agents and serving systems that compile, execute, or dispatch generated code without line-by-line review.

By Shuoming Zhang, Ruiyuan Xu, Haofeng Li, Qiuchu Yu, Yangyu Zhang, Chunwei Xia, Xiaobing Feng, Chenxi Wang, Huimin Cui, Jiacheng Zhao