The paper introduces the Tri‑Metric Router, a deterministic, training‑free policy that chooses among Raw, Neural, and Lexical pipelines for retrieval‑augmented generation on commodity GPUs. It uses three CPU‑side signals—spatial complexity, syntactic density, and type‑token ratio—to balance VRAM headroom and latency, calibrated on LongBench qasper. The method eliminates out‑of‑memory failures and improves alignment and F1 scores compared to always‑on lexical compression without extra VRAM or training costs.
By Saipraveen Vabbilisetty, Ajay Kumar Boddepalli, Deep Narayan Mishra, Shashank Kapadia, Haoan Wang, Anupriya Sharma
arXiv:2602. 04101v2 Announce Type: replace Abstract: We present Interfaze, a native hybrid model that fuses task-specific deep neural networks (CNNs and DNNs) directly into a transformer decoder through a shared embedding space.
By Harsha Vardhan Khurdula, Vineet Agarwal, Yoeven D Khemlani
IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.
By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao
The Transformer Accelerator (TFA) is a synthesizable, parameterizable INT8 memory‑to‑memory engine designed for transformer inference and machine translation. It features a one‑time‑multiplexed datapath that handles prompt processing and autoregressive generation, and implements key operations such as matrix multiplication, softmax, RMSNorm, and elementwise functions through eight 512‑bit macro‑op descriptors. In extensive verification, TFA achieved zero mismatches across 25 tests and 34 constrained‑random runs, matched floating‑point references on multiple translation tasks, and delivered a 20× speedup over a 22‑thread CPU while projecting significant energy reductions in larger designs.
By Shashank
GEPARD is a streaming text‑to‑speech model that uses a standard large language model backbone to generate speech autoregressively, decoding audio with an FSQ‑based neural codec. It streams audio chunk‑by‑chunk as text arrives, achieving a real‑time factor of about 0.067 and an aggregate speedup of roughly 204× on a single GPU with 256 concurrent streams. The design keeps all complex auxiliary mechanisms outside the decode loop, enabling deployment with a standard LLM engine (vLLM) without kernel modifications.
By Denis Pavlov, Ulanbek Abdurazakov, Nursultan Bakashov
The paper introduces the Offline AI Modules workstream, which provides a voice‑first offline architecture, a low‑cost hardware reference bill of materials, and a reproducible quantization and benchmarking pipeline for instruction‑tuned language models in the 2‑5B parameter range. It evaluates three models across four quantization formats on two hardware tiers—NVIDIA Jetson Orin NX and Raspberry Pi5—measuring deployment metrics and multilingual quality. The key result is that Q4_K_M quantization offers the best size‑to‑quality trade‑off, enabling high decode throughput and strong topic classification accuracy while staying within memory limits on both tiers.
By Sunday Afariogun, Odunolaoluwa Jenrola, Zeinab Nezami
The paper introduces Numbat, a self‑contained machine‑learning stack implemented entirely in Zig with no external runtime dependencies. It covers tensor computation, automatic differentiation, neural‑network modules, mixed precision, multi‑GPU training, data loading, and monitoring, and exposes a stable C ABI with over 1,400 entry points and bindings for six languages. The authors verify the stack against a reference implementation at multiple levels, uncovering ten silent recipe divergences, and demonstrate its practical capability by training a 25.9M‑parameter YOLOv8m detector on COCO 2017, achieving a competitive mAP score and matching single‑GPU performance.
By Thang Tran (CloudKites AI Lab, New South Wales, Australia), Lan Dang (Monash Business School, Monash University, Victoria, Australia)
XLOG is a CUDA‑native logic programming engine that fuses neural perception with deterministic Datalog, probabilistic inference, and epistemic world views. It offers multiple reasoning modes that share device data planes, with host‑orchestrated Datalog and exact inference, and resident recursive and Monte Carlo cores that avoid host‑device transfers. The system supports end‑to‑end gradients through GPU knowledge compilation, achieves significant speedups in MNIST‑addition training and join operations, and demonstrates competitive accuracy on several benchmark tasks.
By Levi Dubrovin, Nikita Pospelov, Kirill Sabitov
arXiv:2606. 26453v1 Announce Type: new Abstract: We present KernelPro, a closed-loop multi-agent system that automatically generates, profiles, and iteratively optimizes GPU kernel code by integrating large language model (LLM) code generation with hardware profiler feedback and pluggable bottleneck detection tools.
By Jiading Gai, Shuai Zhang, Kaj Bostrom, Jin Huang, Vihang Patil, Haoyang Fang, Bernie Wang, Huzefa Rangwala, George Karypis
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
By Adwaid Suresh, Aparna A, Harshini V M, Jona Delcy C A, Killi Uma Maheswara Rao, Ram Charan Golla, Surendra Vendra
arXiv:2606. 04025v1 Announce Type: cross Abstract: Dominant programming paradigms inherit an execution model optimised for a bygone era of a single human mind instructing a local machine, leaving contemporary systems burdened with historical path dependencies.
By Philip Sheldrake, Dirk Scheffler
arXiv:2608. 12004v1 Announce Type: cross Abstract: In modern AI frameworks, GPU kernels are key to overall system performance.
By Jinjun Huang, Zhongzhen Wen, Tongtong Xu, Meng Yan, Xin Xia, Zhongxin Liu