← Back to all news
arXiv Machine Learning September 2, 2026 By Teng-Ruei Chen

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 17

The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inference

arXiv:2608. 13756v1 Announce Type: new Abstract: Two GPU kernels implementing the same scaled INT8 GEMM interface are usually treated as interchangeable.

By Teng-Ruei Chen
llmsefficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 19

The Correctness Illusion in LLM-Generated GPU Kernels

arXiv:2606. 20128v1 Announce Type: cross Abstract: Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks.

By Dipankar Sarkar
llmsbenchmarks
More like this →
arXiv AI
Aug 7

Runtime Observability for Heterogeneous Attention Memory

arXiv:2608. 05863v1 Announce Type: new Abstract: Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory in a different form, and each fails differently under compression.

By Fanzhe Wei, Li Liu, Ziyang Wang, Chenyu Wang
efficiency
More like this →
arXiv Machine Learning
Aug 11

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

arXiv:2608. 08910v1 Announce Type: cross Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales.

By Matteo Grella
llmsbenchmarks
More like this →
arXiv AI
Jul 17

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

arXiv:2607. 14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.

By A. C. Opus, J. Q. Lu
ragmultimodalbenchmarks
More like this →
arXiv Machine Learning
Aug 7

Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening

arXiv:2608. 05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.

By Seon Ho Kim, Ui Jeong Jeon, Su Hyeon Kim, Min Tae Hwang
fine-tuningefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea