arXiv:2606. 24074v1 Announce Type: cross Abstract: Wang~\cite{Wang2026} introduced the Stochastic-Oracle Turing Machine (SOTM) framework and defined token complexity as the minimum expected cost of interacting with a stochastic oracle needed to attain a specified solution quality for a task.
By Jie Wang
arXiv:2606. 12647v1 Announce Type: cross Abstract: AI-augmented computing delegates natural language queries, code generation requests, and other open-ended tasks to a cluster of AI models that processes queries and generates responses.
By Jie Wang
arXiv:2609.35831v1 Announce Type: new
Abstract: Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented...
By Isaac Olufadewa, Miracle Adesina, Ezekiel Oladejo, Owen Adeniyi, Fadare Fadekemi, Olamide Oso, Uthman Babatunde, Matthew Olawoyin
arXiv:2607. 04281v1 Announce Type: cross Abstract: Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar queries, but most existing methods do not model the time-varying freshness of open-web evidence.
By Muhammad Mansoor, Tahir Ahmad, Yeo-Chan Yoon
arXiv:2606. 26836v1 Announce Type: new Abstract: Existing benchmarks typically report accuracy for a single model on a single run.
By Bradley Fowler, Ryan Smith, Daniel Thi Graviet, William Myers, Joshua Greaves, Narmeen Fatimah Oozeer, Ant\'ia Garc\'ia, Philip Quirke, Amirali Abdullah, Fazl Barez, Shriyash Kaustubh Upadhyay
arXiv:2609.38006v1 Announce Type: new
Abstract: Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable tha...
By Kenan Alkiek, Moontae Lee, David Jurgens, V. G. Vinod Vydiswaran