arXiv AI

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

arXiv:2608. 08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased.

arXiv Machine Learning
Sep 18

The Complexity Kink: A Prompt-Side Structural Complexity Index for Code-Generation Reliability

The paper introduces a six‑dimension prompt‑side structural‑complexity index to assess code‑generation reliability before a model generates output. Using 5,000 Python prompts and 21 large language models, the authors find that pass rates exhibit a non‑monotonic breakpoint around a composite score of 13.75, with task‑type and construction‑frame adjustments shifting this threshold. The study also reports high inter‑rater reliability (ICC = 0.872) and demonstrates that the index can predict failure likelihood without relying on output correctness.

By Michael Hernandez, Tian Zhao