arXiv:2609.39223v2 Announce Type: new
Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...
By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
arXiv:2607. 29134v1 Announce Type: cross Abstract: Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulation into SQL workloads (primarily join-and-aggregate tensor contractions).
By Andrei Ilinescu, Aadi Patwardhan, Rihan Hai
Cross-Backend QIEO is a runtime core for quantum-inspired evolutionary optimization that unifies execution across OpenMP5, CUDA, HIP, and multiple high-level languages. It compiles a single C++ implementation per hardware target and dispatches to CPU, multi-core, NVIDIA, or AMD backends at runtime, adapting kernels to each device’s memory hierarchy. The framework is validated with real-world bindings: a Python neural‑network hyperparameter optimizer achieving 88.60 % MNIST accuracy, a MATLAB wind‑farm layout optimizer matching particle swarm results and outperforming genetic algorithms, and a Julia package that reduces mean SSE by 2.1× in Lotka–Volterra parameter estimation.
By Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh, Abhishek Chopra
arXiv:2605. 30358v2 Announce Type: replace Abstract: Quantum computing remains in the Noisy Intermediate-Scale Quantum (NISQ) era, with performance constrained by noise.
By Zhenxiao Fu, Lei Jiang, Fan Chen
The paper introduces MQSS-Selector, a reinforcement‑learning guided pass selection system for an MLIR compilation pipeline aimed at unified High Performance Computing‑Quantum Computing (HPCQC) infrastructures. It addresses the challenges of Noisy Intermediate‑Scale Quantum (NISQ) devices by integrating device selection, compiler‑pass optimization, and job queue scheduling into a single learning‑based framework. The selector can simultaneously optimize multiple objectives—fidelity, compilation time, and scheduling latency—while adapting to circuit characteristics and device conditions.
By Andre Youssefi (Leibniz Supercomputing Centre), Erc\"ument Kaya (Leibniz Supercomputing Centre, Technical University of Munich), Minh Chung (Leibniz Supercomputing Centre), Jorge Echavarria (Munich Quantum Valley), Laura B. Schulz (Argonne National Laboratory), Martin Schulz (Leibniz Supercomputing Centre, Technical University of Munich)
The study evaluates the portability of INT8 post‑training quantization across seven hardware platforms, including CPUs, GPUs, and vendor NPUs, by keeping the ONNX model and quantization scales constant. It finds that INT8 performance and output consistency vary significantly: CPU dot‑product instructions determine speedup, identical INT8 outputs only occur when integer kernels match, and vendor NPUs require their own quantization pipelines. The authors also show that edge‑NPU latency is dominated by data transfer rather than compute and provide scripts and reports for reproducibility.
By Yuyeong Shin