arXiv Machine Learning

VQCSim: When Does Compile-Once Statevector Simulation Beat Generic Quantum Frameworks?

arXiv:2607. 11985v1 Announce Type: cross Abstract: Hybrid quantum-classical machine learning workflows repeatedly evaluate many small parametrized circuits during training and model exploration.

arXiv Machine Learning
2d ago

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

arXiv:2609.39223v2 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive p...

By Weili Xu, Jisen Li, Yuqing Jian, Chenxi Li, Zhizhou Sha, Yifan Yu, Qingyang Wu, Chenfeng Xu, Zhongzhu Zhou, Tianyi Zhang, Ben Athiwaratkun
arXiv Computation and Language
6d ago

Cross-Backend QIEO: Universal Runtime Portability across OpenMP5, CUDA, HIP, and Multi-Language Interfaces

Cross-Backend QIEO is a runtime core for quantum-inspired evolutionary optimization that unifies execution across OpenMP5, CUDA, HIP, and multiple high-level languages. It compiles a single C++ implementation per hardware target and dispatches to CPU, multi-core, NVIDIA, or AMD backends at runtime, adapting kernels to each device’s memory hierarchy. The framework is validated with real-world bindings: a Python neural‑network hyperparameter optimizer achieving 88.60 % MNIST accuracy, a MATLAB wind‑farm layout optimizer matching particle swarm results and outperforming genetic algorithms, and a Julia package that reduces mean SSE by 2.1× in Lotka–Volterra parameter estimation.

By Aman Mittal, Ferdin Sagai Don Bosco, Kasturi Venkata Srikanth, Abhishek Singh, Aditya Singh, Abhishek Chopra
arXiv Machine Learning
Sep 25

MQSS-Selector: RL-Guided Pass Selection for an MLIR Compilation Pipeline

The paper introduces MQSS-Selector, a reinforcement‑learning guided pass selection system for an MLIR compilation pipeline aimed at unified High Performance Computing‑Quantum Computing (HPCQC) infrastructures. It addresses the challenges of Noisy Intermediate‑Scale Quantum (NISQ) devices by integrating device selection, compiler‑pass optimization, and job queue scheduling into a single learning‑based framework. The selector can simultaneously optimize multiple objectives—fidelity, compilation time, and scheduling latency—while adapting to circuit characteristics and device conditions.

By Andre Youssefi (Leibniz Supercomputing Centre), Erc\"ument Kaya (Leibniz Supercomputing Centre, Technical University of Munich), Minh Chung (Leibniz Supercomputing Centre), Jorge Echavarria (Munich Quantum Valley), Laura B. Schulz (Argonne National Laboratory), Martin Schulz (Leibniz Supercomputing Centre, Technical University of Munich)
arXiv Machine Learning
Sep 16

Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators

The study evaluates the portability of INT8 post‑training quantization across seven hardware platforms, including CPUs, GPUs, and vendor NPUs, by keeping the ONNX model and quantization scales constant. It finds that INT8 performance and output consistency vary significantly: CPU dot‑product instructions determine speedup, identical INT8 outputs only occur when integer kernels match, and vendor NPUs require their own quantization pipelines. The authors also show that edge‑NPU latency is dominated by data transfer rather than compute and provide scripts and reports for reproducibility.

By Yuyeong Shin
arXiv Machine Learning
Sep 3

Inference-Native Zeroth-Order Optimization

The paper introduces Inference‑Native Zeroth‑Order (ZO) optimization, which redefines ZO as a query‑based process that can be executed directly by inference runtimes. By exposing ZO’s query semantics and using abstractions such as ProbePlan, factorized side states, and persistent subspace reuse, the method reduces state‑management cost and DRAM traffic dramatically. Experiments on large models (OPT‑13B, Qwen3‑8B) show that inference‑native steps are nearly identical to matched‑query controls while achieving significant memory savings and efficient batching.

By Zelin Li, Caiwen Ding
arXiv Machine Learning
Sep 7

Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks

The paper investigates how postprocessing routines in quantum neural network software can cause significant data loss when run on large quantum hardware. In a case study of Qiskit’s “SamplerQNN”, a filter that assumes measurement bit‑strings are in virtual qubit space removed 85–99.6% of valid shots on IBM backends, leading to unnormalised probability vectors and distorted predictions. This loss caused inference accuracy to drop from 0.94 to 0.39 and compressed training loss signals by 22–27×, severely reducing optimizer sensitivity. The authors implemented a layout‑based marginalisation fix that was merged into the library to make “SamplerQNN” forward‑compatible with current and future hardware.

By Soraya V. Panambalom, Edoardo Altamura, Nick Chancellor, Jonte R. Hance
Hugging Face Trending Papers
Sep 4

Impact of Data Loss in Postprocessing on Training and Inference of Quantum Neural Networks

The paper investigates how postprocessing routines in quantum neural network software can inadvertently discard a large portion of valid measurement data when run on real quantum hardware. In a case study of Qiskit’s “SamplerQNN”, a filter that assumes virtual qubit space caused 85–99.6% of measurement shots to be lost on IBM backends, leading to unnormalised probability vectors, degraded inference accuracy (from 0.94 to 0.39), and a 22–27× compression of the training loss signal. The authors provide a layout‑based marginalisation fix that has been merged into the library to ensure forward‑compatibility with current and future hardware.

arXiv Machine Learning
Jul 17

PolyQ: Codesigning End-to-End Quantization Framework for Scalable Edge CPU LLM Inference

arXiv:2607. 14618v1 Announce Type: new Abstract: CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs.

By Hyunwoo Oh, Suyeon Jang, Hanning Chen, KyungIn Nam, Sanggeon Yun, Ryozo Masukawa, Mohsen Imani
arXiv Machine Learning
Jun 11

Family-Aware Residual Architecture for Predicting Quantum Circuit Simulation Performance

arXiv:2606. 11620v1 Announce Type: cross Abstract: Approximate tensor-network simulators enable classical simulation of quantum circuits beyond the reach of exact methods, but selecting optimal approximation parameters -- such as bond dimension thresholds -- remains a costly trial-and-error process.

By Honjar Xing, Yehong Jiang, Xianbang Wang, Zehua Wang, Zhicheng Jiang