arXiv AI By Pawe\l{} Liskowski, Fuheng Zhao, Benjamin Han, Anupam Datta, Dimitris Tsirogiannis

Compositional Online Learning for Semantic Data Processing Systems

Read the original on arXiv AI →

The paper introduces compositional online learning for semantic data processing systems, addressing the high cost and latency of large language model (LLM) calls. It proposes a framework that combines lightweight online-learning components—such as memoization, per-call filter-ordering, and per-batch cascade-routing—within the LLM call boundary, allowing each component to make real-time decisions and update its models without exceeding the LLM round-trip time. A production case study in Cortex AISQL demonstrates that these components can reduce the per-row LLM cost by up to 8× compared to a baseline workload.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Inference-Native Zeroth-Order Optimization

The paper introduces Inference‑Native Zeroth‑Order (ZO) optimization, which redefines ZO as a query‑based process that can be executed directly by inference runtimes. By exposing ZO’s query semantics and using abstractions such as ProbePlan, factorized side states, and persistent subspace reuse, the method reduces state‑management cost and DRAM traffic dramatically. Experiments on large models (OPT‑13B, Qwen3‑8B) show that inference‑native steps are nearly identical to matched‑query controls while achieving significant memory savings and efficient batching.

By Zelin Li, Caiwen Ding
arXiv AI
Jul 9

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.

By Andrii Balashov, Olena Ponomarova
arXiv Computation and Language
Aug 31

Are These Modules Worth Their Cost? A Paradigm-Level Accuracy-Cost Analysis of In-context Learning Text-to-SQL

The paper evaluates 17 paradigm-level configurations of in‑context learning (ICL) text‑to‑SQL pipelines across five common modules, measuring each module’s marginal accuracy contribution and cost for four different backbone models. It finds that execution‑feedback refinement consistently improves accuracy at low cost, while other modules only help under specific backbone conditions. The study also shows that investing in a more elaborate pipeline for a mid‑tier backbone can be more cost‑effective than upgrading to a high‑capability model with a lean pipeline, providing a tiered, cost‑aware guideline that generalizes to additional backbones.

By Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang
arXiv AI
Sep 23

Graph Memory for LLM Agents: At What Cost? A Comparative Evaluation of Query, Ingest, and Update Performance Across Graph Database Engines

The paper evaluates seven graph database engines, including Corvic AI, on a synthetic biomedical property graph with 1.02 million nodes and 5.34 million rows. It benchmarks query latency, bulk‑ingest throughput, point‑update latency, and correctness across a twenty‑query workload that covers neighborhood lookups, bounded paths, set intersections, anti‑joins, aggregation, ranking, temporal filters, full scans, and relational joins. The study finds that no single engine is universally fastest; performance depends on query shape, and the largest cost difference arises from bulk‑ingest throughput, which varies by three orders of magnitude and dominates total cost for workloads with fewer than about 10⁵ queries per data refresh.

By Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach