arXiv:2609.14500v1 Announce Type: new
Abstract: AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a hi...
By Seyed Morteza Emadi
The paper introduces a pragmatic information theory that unifies communication, control, and decision-making through the isoteleia mapping, which formalizes equifinality by treating distinct semantic paths that lead to the same optimal action as pragmatically equivalent. It establishes a three-tier hierarchy of syntactic, semantic, and pragmatic information, defines pragmatic entropy, mutual information, channel capacity, and rate-distortion, and proves coding theorems that generalize Shannon’s results. The authors also present pragmatic value and cost of information, a Lagrangian dual framework for cross-layer optimization, and a pragmatic efficiency bound that quantifies the maximum net utility for resource-constrained intelligent systems, extending the theory to continuous messages and dynamic settings.
By Kai Niu, Ping Zhang
arXiv:2602. 05463v2 Announce Type: replace-cross Abstract: Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption.
By Koichi Takahashi, Yusuke Hayashi
arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.
By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig
The paper presents an Edge-AI-driven decentralized task‑allocation framework for circular smart manufacturing. It combines a resource‑aware heuristic, a regression‑based Edge‑AI bid approximation, and a compact autoencoder‑regularized pairwise ranking model to evaluate tasks at the machine level. Simulation results show that the ranking method improves task completion, reduces tardiness and deadline misses, and lowers energy per completed task compared to the heuristic baseline.
By Mohammadhossein Ghahramani, Yan Qiao, Mengchu Zhou
arXiv:2602. 08261v2 Announce Type: replace Abstract: Auto-bidding systems strive to maximize marketing value while maintaining high compliance with efficiency constraints, such as Target Cost-Per-Action (CPA).
By Binglin Wu, Yingyi Zhang, Xianneng Li, Ruyue Deng, Chuan Yue, Weiru Zhang, Xiaoyi Zeng
The paper introduces a composite metric for quantizing small language models that balances information retention and throughput gains, using a normalized SQNR-based coefficient and roofline-based latency analysis. Profiling Gemma 3 1B shows that Feed‑Forward Network blocks and the embedding matrix are prime candidates for acceleration, with the metric enabling tuning of speed‑quality trade‑offs without actual execution. The authors demonstrate that their estimates predict accelerated speedup within about 4% error and allocate resources more effectively than evolutionary search or Shapley‑value methods.
By Artem Safronov
arXiv:2605. 18909v2 Announce Type: replace Abstract: Any system that models the world under finite representational capacity must compress; any compression entails a prior; and the prior is the system's bias.
By Ahmed Gamal Eldin
The paper investigates how much learned memory is required to leverage additional data in autoregressive prediction models. It introduces a predictive‑energy spectrum that jointly governs data and memory scaling, proving a minimax law that links the number of prediction blocks and the size of the learned state to this spectrum. The authors demonstrate that optimal bit allocation and masked query‑key attention mechanisms realize this law, and they provide experimental evidence across multiple pretrained‑model scales.
By Chiwun Yang, Xiaoyu Li
arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.
By Vasileios Sevetlidis
arXiv:2609.21523v1 Announce Type: new
Abstract: A system may be compressed before its downstream task is fully known. We ask how much retained state is then necessary and how much can be saved by lim...
By Ronald Katende
The paper argues that AI deployment performance depends on interactions among compression, compiler transformations, and serving policies rather than just model architecture. It introduces a three‑layer taxonomy—model‑level techniques, compiler transformations, and system policies—and frames deployment as a constrained multi‑objective optimization problem over accuracy, latency, throughput, memory footprint, and energy. The authors propose an evidence protocol for comparable benchmarking and synthesize data from edge and data‑center platforms to show that cross‑layer interactions drive deployment outcomes, concluding with a constraint‑aware selection procedure and open research problems.
By Tejinder Singh, John Pflueger, Jeebak Mitra, Robert Lincourt, Mitchell Markow, Bhavesh A. Patel