arXiv AI By Kenan E. Ak, Jay Mohta, Gwang Gook Lee, Yan Xu, Dimitrios Dimitriadis

An Empirical Study of VLM Pipelines for Long-Document QA

Read the original on arXiv AI →

The paper evaluates how different design choices—document feeding strategy, retrieval method, and execution mode—affect Vision‑Language Models (VLMs) on long‑document question answering. Experiments on two benchmarks show that a multi‑tool agent only outperforms static input when the VLM is large, that retrieval modality (image vs text) is more critical than the specific retriever, and that combining the best pipelines per question can significantly boost performance. The study highlights the trade‑offs between token efficiency, model size, and pipeline complexity for deploying VLMs on complex documents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

BudgetMem: Training-Free Selective Memory for Cost-Efficient Long-Context Processing in Language Models

arXiv:2511. 04919v3 Announce Type: replace Abstract: Processing long documents with large language models (LLMs) is expensive: a single query over a 100K-token document can cost from tens of cents to over a dollar in API fees, depending on the model, and memory grows linearly with context length.

By Chandra Vamsi Krishna Alla, Harish Naidu Gaddam, Manohar Kommi, Sheikh Nazib Ahmed