arXiv AI

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

arXiv:2607. 17219v1 Announce Type: cross Abstract: Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free prediction of the projective quantum question-order model.

arXiv Machine Learning
Sep 18

QEncodeBench: Can Large Language Models Encode Classical Problems into Verified Quantum Oracles?

QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.

By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv Computation and Language
Aug 24

Prompt-Model Interaction Reaches the Fixed Points: A deterministic, task-free structural readout -- and the factorizations of it that failed

The paper demonstrates that a prompt’s influence is not inherent to the prompt itself but depends on the model, as prompts optimized for one model degrade on another and rankings shift under neutral reformatting. By examining a task‑free structural readout—specifically the fixed‑point behavior of a short‑window argmax map—the authors show that nine tokens of conditioning can move the fixed‑point fraction across most of its range, altering structural classes and model rankings, while instruction tuning has no effect. Attempts to explain this phenomenon through prefix length, content type, bidirectionality, or attention‑sink dominance all fail, indicating that the prompt‑model pair is the fundamental unit of explanation. whyItMatters":"The study reveals that prompt effectiveness is model‑specific and that simple structural readouts can capture this interaction, challenging assumptions about prompt generality and guiding future prompt‑engineering efforts."

By Nicol\'as Vera Z\'u\~niga
arXiv AI
Sep 24

Are Stated Reasoning Steps Causally Load-Bearing?

The study investigates whether the reasoning steps a language model writes are causally responsible for its answers. Using a causal intervention method on the activation stream, the authors find that for Qwen3-4B, about 77% of stated steps are causally load‑bearing, while behavioral tests overestimate this by roughly 11 percentage points. The faithfulness of reasoning decreases with model size and depth of reasoning, especially for the smaller Qwen3-1.7B.

By Abhiram Bhupatiraju, Rayan Nyaupane