We’re releasing an analysis showing that since 2012, the amount of compute used in the largest AI training runs has been increasing exponentially with a 3. 4-month doubling time (by comparison, Moore’s Law had a 2-year doubling period)[^footnote-correction].
The hidden cost of asynchronous systems, how tiny CPU tasks quietly became our biggest bottleneck while scaling hundreds of LLM agents. The post Why Adding More AI Agents Made Our System Slower appeared first on Towards Data Science .
By Uri Peled
arXiv:2608. 26418v1 Announce Type: cross Abstract: Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months.
By Architect Labs
OpenAI’s new custom inference chip, Jalapeño, achieves industry-leading speed and efficiency in AI inference. It delivers faster, more power‑efficient performance with higher throughput and lower latency for modern models. The chip represents a significant hardware advancement for deploying AI workloads.
arXiv:2511. 07885v5 Announce Type: replace-cross Abstract: Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure.
By Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, J. Wes Griffin, Herumb Shandilya, Adrian Gamarra Lafuente, Medhya Goel, Rebecca Joseph, Shlok Natarajan, Etash Kumar Guha, Shang Zhu, Ben Athiwaratkun, John Hennessy, Azalia Mirhoseini, Christopher R\'e
arXiv:2608. 03682v1 Announce Type: new Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment.
By Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Tianyue Zhang, Weikai Xie, Xiyuan Tan, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Ziqi Guo
The article titled "The Work Now Within Reach" discusses how increasingly capable and affordable AI technologies can broaden the range of tasks that individuals and businesses can perform. It highlights the potential for AI to enhance productivity and enable more efficient, cost-effective growth. The piece emphasizes the expanding opportunities for leveraging AI to achieve greater work outcomes.
arXiv:2606. 07632v1 Announce Type: new Abstract: Proper accounting of the energy requirements and environmental impact of artificial intelligence (AI) systems is necessary for researchers, developers, policy makers, and users to assess the barriers to building systems at scale.
By Jared Fernandez, Clara Na, Yonatan Bisk, Constantine Samaras, Emma Strubell
arXiv:2609.24274v1 Announce Type: cross
Abstract: Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between polic...
By Khanh D. Nguyen, Hoang M. Truong, An T. Le
arXiv:2505. 01458v2 Announce Type: replace-cross Abstract: Navigation and manipulation are core capabilities in Embodied AI, but training agents to perform them directly in the real world is costly, time-consuming, and unsafe.
By Lik Hang Kenny Wong, Xueyang Kang, Kaixin Bai, Jianwei Zhang
AI’s next frontier isn’t just about capability—it’s about who gets to use it. Our mission to put AI in the hands of as many people as possible is what drives us.
The paper introduces rMuscle, a real‑time Vision‑Language‑Action inference framework that mimics human muscle memory to accelerate robotic decision making. By exploiting repeated task similarity, rMuscle uses a dual‑phase cache: a Context Cache reuses visual‑token outputs and an Action Cache reuses neuron activation patterns, reducing computation and weight accesses. Experiments on RTX 4090 and Jetson Thor show 1.29–1.42× speedups on LIBERO, RoboTwin, and physical manipulation tasks while preserving success rates on real robots.
By Kaijun Zhou, Zhiyang Li, Le Chen, Jinyu Gu