arXiv AI

An Empirical Study of VLM Pipelines for Long-Document QA

The paper evaluates how different design choices—document feeding strategy, retrieval method, and execution mode—affect Vision‑Language Models (VLMs) on long‑document question answering. Experiments on two benchmarks show that a multi‑tool agent only outperforms static input when the VLM is large, that retrieval modality (image vs text) is more critical than the specific retriever, and that combining the best pipelines per question can significantly boost performance. The study highlights the trade‑offs between token efficiency, model size, and pipeline complexity for deploying VLMs on complex documents.

arXiv Computation and Language
Sep 22

BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models

arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.

By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed
arXiv AI
Jun 30

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

arXiv:2606. 28344v1 Announce Type: cross Abstract: Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting.

By Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min
arXiv AI
5d ago

TRACE: Trajectory Selection for Parallel Scaling of Search Agents

TRACE is a lightweight learned selector that ranks completed search trajectories by aggregating cross‑rollout evidence, preserving individual query and evidence occurrences while propagating information across shared content or document identity. Trained with answer‑level supervision over frozen text embeddings, TRACE selects an existing answer without additional search or autoregressive aggregation, and a single selector generalizes across rollout policies and agent backbones. Across six WebQA policies, six long‑horizon dataset‑backbone combinations, and multiple WebQA benchmarks, TRACE outperforms majority voting and generative aggregators, achieving higher accuracy and at least tenfold higher processing throughput.

By Qisheng Zhou, Zhen Xiong, Qiaoyu Tan
arXiv AI
Jul 8

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

arXiv:2607. 05399v1 Announce Type: cross Abstract: Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression techniques are difficult to compare because they were evaluated on different models, tasks, budgets, and serving stacks.

By Nikita Agrawal, Ruben Mayer
arXiv Machine Learning
1d ago

Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents

The paper introduces <flowreader>, a method that casts evidence selection for long multimodal documents as a minimum‑cost flow problem over a multimodal content graph. It uses spectral decomposition to identify latent query‑relevant aspects and allocates a fixed evidence budget proportionally to their spectral energy, ensuring aspect coverage without a language‑model planning call. On the VisDoMBench benchmark with Qwen3‑VL‑32B, <flowreader> achieves the highest macro accuracy (68.9%) and outperforms prior systems by 2.7 points, while using fewer content nodes and reader tokens.

By Ambuj Mehrish, Sebastiano Vascon