Hugging Face Trending Papers

QC-Stark: A Multi-Task Benchmark Revealing Capability Dissociations in LLMs Evaluated on Quantum Computing Tasks

Read the original on Hugging Face Trending Papers →

QC-Stark is a benchmark that evaluates large language models on 11 quantum computing tasks, including circuit construction, debugging, compilation, error correction, and simulation. It comprises 2,750 evaluations across 10 models, 5 difficulty levels, and 5 seeds, revealing that overall model rankings can hide significant per-task differences, with Spearman correlation being insignificant for 4 of the 11 tasks. The benchmark’s measurement quality is validated by a 2‑parameter Item Response Theory model, prompt sensitivity analysis confirms ranking robustness, and all tasks are auto‑verifiable via execution, eliminating the need for manual evaluation. The code and data are publicly available on Hugging Face.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 18

QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.

By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv AI
Aug 11

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing

arXiv:2608. 07743v1 Announce Type: new Abstract: Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must preserve the task, respect access and output models, expose required promises, and remain within a defensible complexity scope.

By Yijing Zuo, Zhe Fu, Zihan Nie, Zhihui Zhu, Haohan Wang
arXiv AI
Sep 16

QART: A Quantum-Classical Hybrid Architecture for Long-Horizon Reasoning -- Exploring a Conditional Path toward Quantum Scaling

QART is a quantum‑classical hybrid architecture that augments a language model with quantum encoding, CIM‑based QUBO optimization, and quantum decoding to improve long‑horizon reasoning. The authors claim that, under certain assumptions, QART can maintain a non‑zero probability of recovering an optimal reasoning path while traditional autoregressive LLMs see their acceptance probability drop to zero as cumulative risk grows. Experiments on six benchmarks with three backbone models show that QART outperforms the baselines in 14 of 15 pairings, with relative gains up to 84.0% on SciCode.

By Lehao Lin, Yuheng Cheng, Guolong Liu, Yao Li, Xuning Tan, Xiyuan Zhou, Ruixi Zou, Shi Wang, Huan Zhao, Wenxuan Liu, Haifeng Wu, Junhua Zhao
arXiv AI
Jul 24

PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation

arXiv:2605. 25572v2 Announce Type: replace-cross Abstract: The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based code assistants: general-purpose models hallucinate PennyLane-specific gate names, misplace device configurations, and produce structurally invalid circuits when faced with specialized quantum coding challenges.

By Minghao Shao, Nouhaila Innan, Hariharan Janardhanan, Muhammad Kashif, Alberto Marchisio, Muhammad Shafique