arXiv AI By Sherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler

Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration

Read the original on arXiv AI →

Maia 200 is a software‑defined dataflow system that delivers high performance AI acceleration, achieving 10,145 Tflop/s in FP4 and 5,072 Tflop/s in FP8 within a 750 W TDP and 7 TB/s HBM bandwidth. It exemplifies a new class of Software Defined Locally Accessed Dataflow Architectures (SDLA), which program dataflow engines to orchestrate specialized memories and data‑movement engines, shifting focus from thread‑centric to data‑movement‑centric design. The system offers significant cost and energy savings while supporting massive parallelism for AI inference workloads, positioning it as a compelling solution for next‑generation high‑performance computing.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 5

Advanced AI Service Provisioning in O-RAN through LLM Engine Integration

arXiv:2605. 23809v2 Announce Type: replace-cross Abstract: The Open Radio Access Network (O-RAN) architecture allows AI to be embedded directly into the RAN through modular xApps and rApps, yet creating these applications collecting data, training models, writing code, and deploying them safely remains slow and largely manual.

By Seyed Bagher Hashemi Natanzi, Pranshav Gajjar, Bo Tang, Vijay K. Shah
arXiv Machine Learning
Sep 14

AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training

AsyncFlow is an asynchronous streaming reinforcement learning framework designed to improve the post‑training phase of large language models. It introduces a distributed data storage and transfer module that enables panoramic data management and fine‑grained scheduling, allowing automated pipeline overlapping and dynamic load balancing. The framework also employs an asynchronous producer‑consumer workflow to reduce computational idleness by deferring parameter updates within staleness thresholds, and it is architecturally decoupled from training and inference engines, providing modular, customizable user interfaces. Experiments show an average throughput improvement of 1.59× over the state‑of‑the‑art baseline.

By Zhenyu Han, Ansheng You, Haibo Wang, Kui Luo, Guang Yang, Wenqi Shi, Menglong Chen, Sicheng Zhang, Zeshun Lan, Chunshi Deng, Huazhong Ji, Wenjie Liu, Yu Huang, Yixiang Zhang, Chenyi Pan, Jing Wang, Xin Huang, Chunsheng Li, Jianping Wu