arXiv:2603.21676v2 Announce Type: replace-cross
Abstract: Standard Transformers have a fixed computational depth, limiting their ability to generalize to tasks that require variable-depth reasoning....
By Hung-Hsuan Chen
The paper introduces Inference‑Native Zeroth‑Order (ZO) optimization, which redefines ZO as a query‑based process that can be executed directly by inference runtimes. By exposing ZO’s query semantics and using abstractions such as ProbePlan, factorized side states, and persistent subspace reuse, the method reduces state‑management cost and DRAM traffic dramatically. Experiments on large models (OPT‑13B, Qwen3‑8B) show that inference‑native steps are nearly identical to matched‑query controls while achieving significant memory savings and efficient batching.
By Zelin Li, Caiwen Ding
arXiv:2604. 00660v2 Announce Type: replace-cross Abstract: Modern data warehouses extend SQL with semantic operators that invoke large language models on each qualifying row, making per-row inference orders of magnitude more expensive than traditional SQL.
By Pawe{\l} Liskowski, Kyle Schmaus
arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.
By Andrii Balashov, Olena Ponomarova
The paper evaluates 17 paradigm-level configurations of in‑context learning (ICL) text‑to‑SQL pipelines across five common modules, measuring each module’s marginal accuracy contribution and cost for four different backbone models. It finds that execution‑feedback refinement consistently improves accuracy at low cost, while other modules only help under specific backbone conditions. The study also shows that investing in a more elaborate pipeline for a mid‑tier backbone can be more cost‑effective than upgrading to a high‑capability model with a lean pipeline, providing a tiered, cost‑aware guideline that generalizes to additional backbones.
By Jiayan Lin, Yujia Liu, Zijin Hong, Zheng Yuan, Yilin Xiao, Hao Chen, Qinggang Zhang, Xiao Huang, Feiran Huang
The paper evaluates seven graph database engines, including Corvic AI, on a synthetic biomedical property graph with 1.02 million nodes and 5.34 million rows. It benchmarks query latency, bulk‑ingest throughput, point‑update latency, and correctness across a twenty‑query workload that covers neighborhood lookups, bounded paths, set intersections, anti‑joins, aggregation, ranking, temporal filters, full scans, and relational joins. The study finds that no single engine is universally fastest; performance depends on query shape, and the largest cost difference arises from bulk‑ingest throughput, which varies by three orders of magnitude and dominates total cost for workloads with fewer than about 10⁵ queries per data refresh.
By Donald Nguyen, Gurbinder Gill, Hadi Ahmadi, Christopher J. Rossbach