Accelerating SD Turbo and SDXL Turbo Inference with ONNX Runtime and Olive
Related stories
Exploring simple optimizations for SDXL
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
It\^o maps for any-step SDEs
arXiv:2606. 11156v1 Announce Type: cross Abstract: Recent one-step generative models accelerate sampling by learning deterministic flow maps of the underlying dynamics.
EvoLP: Self-Evolving Latency Predictor for Model Compression in Real-Time Edge Systems
arXiv:2607. 09063v1 Announce Type: new Abstract: Edge devices are increasingly utilized for deploying deep learning applications on embedded systems.
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.
Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models
arXiv:2511. 21759v2 Announce Type: replace-cross Abstract: Diffusion-based large language models (dLLMs) have recently gained significant attention for their exceptional performance and inherent potential for parallel decoding.
Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge
arXiv:2606. 19964v1 Announce Type: new Abstract: Tsetlin Machine (TM) is a logic-based machine learning approach that relies on simple bitwise operations and finite-state automata, which makes it attractive for edge AI deployments.
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
arXiv:2607. 20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs.
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
arXiv:2607. 06114v1 Announce Type: cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (NFEs).
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
arXiv:2608. 15018v1 Announce Type: new Abstract: Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints.
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
arXiv:2606. 17104v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge.