arXiv:2606. 08476v1 Announce Type: cross Abstract: Context parallelism (CP) is essential for training large-scale, long-context language models, as it partitions sequences to reduce memory overhead.
By Zheng Wang, Eric Liu, Linan Jiang, Zhongkai Yu, Zaifeng Pan, Yue Guan, Yuke Wang, Yufei Ding
Developers can now build fast speech-to-speech experiences into their applications
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
Introducing GPT-5. 3-Codex-Spark—our first real-time coding model.
Flama is an open‑source Python framework that unifies the development and deployment of production‑ready web APIs, machine‑learning services, and large‑language‑model (LLM) applications. Built on ASGI, it offers an async‑first, type‑driven programming model with seven subsystems—including dependency injection, a pluggable schema layer, automatic CRUD generation, a portable binary model format, a multi‑backend LLM server, a Rust‑accelerated core, and a Model Context Protocol module. The framework also provides built‑in JWT authentication, pagination, background tasks, WebSocket and streaming support, OpenAPI generation, and a CLI for running, packaging, and inspecting models.
By Jos\'e A. Perdiguero L\'opez, Miguel A. Dur\'an-Olivencia
arXiv:2607. 06202v1 Announce Type: cross Abstract: The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth.
By Yipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng, Mingfan Li, Yuyang Yang, Guanhua Li, Yuquan Zhang, Yimeng Xu, Zhongzhe Hu, Zhiyuan Huang, Qihang Duan, Junsong Wang, Wenkai Ling, Baochuan Yang, Xianzhi Yu, Han Bao, Yijie Chen, Guihai Chen
arXiv:2607. 26464v1 Announce Type: cross Abstract: Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs).
By Zekun Ren, Hongzhao Tan, Jiaen Yee, Kedar Hippalgaonkar
arXiv:2609.13814v1 Announce Type: new
Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...
By Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou, Jincenzi Wu, Xujie Song, Kexin Shi, Zihang Liu, Pengcheng Zhu, Jiayi Zhou, Baoyue Zhang, Changhao Zhang, Zitong Wang, Jinhong Wang, Tong Niu, Jingjing Liu, Junan Lin, Haolin He, Hengshuo Chu, Yuhui Chen, Jian Liu, Yuge Huang, Junliang Xing, Yuntao Wang, Weiqiang Wang, Chun Yu, Yuanchun Shi
The paper introduces Jarvis, an offline, edge‑deployable voice assistant designed for autonomous racecars. It combines speech recognition, synthesis, and a lightweight text‑to‑command classifier fine‑tuned from the Mistral 7B model to provide high‑level behavioral commands. Experiments show 97.63 % intent recognition accuracy with an average latency of 1.39 s, outperforming larger online‑hosted models and enabling quick response times for time‑critical driving tasks.
By Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz