arXiv AI

Coding Agents for Coding Theory

The authors report that a large‑scale experiment using a language‑model coding agent over five weeks produced new lower bounds for DNA‑barcode‑style codes. By restricting searches to codes with a prescribed symmetry, the agent improved the best known code of length 6 and minimum edit distance 3 from 114 to 120 words, and similarly raised lower bounds for lengths 6–9 and distances 3–6. The study also documents failures and the limitations of the verification protocol, noting that intermediate results were never rechecked and could lead to erroneous conclusions.

arXiv AI
Aug 11

Multi-agent discovery of practical quantum LDPC codes

arXiv:2608. 08996v1 Announce Type: cross Abstract: Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful finite-length instances remains a challenging design problem because code performance must be optimized while satisfying practical constraints.

By Dongheng Qian, Tianyi Li
arXiv AI
Sep 25

Operator Packages, Proposer Strength, and Construction-Family Plateaus in Office-Scale Verified Search

The paper reports on a large‑scale verified search experiment using a 30B language model on a laptop, evaluating three operator packages—schematic notebooks, named obstacles, and behavioural repulsion—in a factorial design across nine construction problems. Results show that the full composition of operators closes the seed‑to‑record gap more effectively than any single component, increases construction‑hash diversity, and that memory plus repulsion consistently avoids collapse. A frontier proposer achieves similar gains in far fewer samples, but the search ultimately stalls near a plateau where the reference family is adopted and optimized only when provided as code.

By Roberto I. Ono Filho
Hugging Face Trending Papers
Jul 29

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% defense success rate. We show it can be breached by composing two attacks that are individually harmless against it: an established code-completion encoding and an established best-of-N search, neither of which exceeds 4.

arXiv Machine Learning
Sep 18

Measurement Under Selection: Decoy-Calibrated Failure Audits for Language Models

The paper introduces Janus, a method for validating error patterns in language models by comparing error rates across predefined yes/no properties and using shuffled decoy labels to set significance thresholds. Janus requires that a pattern’s error difference surpasses the decoy-derived threshold and is replicated on held‑out data before reporting. Experiments on a controlled code‑finding task confirm several meaningful error patterns, while on MuSiQue and LongBench v2 Janus reports no confirmed patterns for the tested properties, contrasting with standard shuffling tests that sometimes confirm patterns.

By Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
Hugging Face Trending Papers
Jun 1

Evolutionary Discovery of Bivariate Bicycle Codes with LLM-Guided Search

Quantum LDPC code discovery requires searching large algebraic design spaces while reliably certifying the parameters and equivalence classes of any candidates found. We introduce an LLM-guided evolutionary workflow in which language models mutate Python programs that generate bivariate-bicycle and perturbed bivariate-bicycle code ansätze.