Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive
Related stories
Exploring simple optimizations for SDXL
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models
The paper reports the development of 2‑bit microkernels for CPUs and mixed‑precision 2‑bit kernels for Intel Xe2 GPUs, achieving near‑roofline performance. Integrated into LLM inference pipelines, these kernels deliver up to 7× speedup over 16‑bit inference on CPUs and 6.7× on GPUs, surpassing the current state‑of‑the‑art bitnet.cpp runtime by 2.2×. The work demonstrates that ultra‑low‑bit LLM models can be deployed efficiently, offering significant gains in latency, memory, throughput, and energy consumption.
It\^o maps for any-step SDEs
arXiv:2606. 11156v1 Announce Type: cross Abstract: Recent one-step generative models accelerate sampling by learning deterministic flow maps of the underlying dynamics.
RoofLang: Enabling AI-Driven Architecting of LLM Inference Systems
arXiv:2609.12551v2 Announce Type: replace-cross Abstract: AI is beginning to make substantive contributions to LLM inference optimization. Existing AI optimizations are predominantly profiling-based....
EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems
arXiv:2607. 09063v1 Announce Type: new Abstract: Edge devices are increasingly utilized for deploying deep learning applications on embedded systems.
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
arXiv:2511. 21759v2 Announce Type: replace-cross Abstract: Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential for parallel decoding.
Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge
arXiv:2606. 19964v1 Announce Type: new Abstract: Tsetlin Machine (TM) is a logic-based machine learning approach that relies on simple bitwise operations and finite-state automata, which makes it attractive for edge AI deployments.
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
arXiv:2607. 20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs.
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
arXiv:2607. 06114v1 Announce Type: cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs).