arXiv AI By Peiying Zhu, Sidi Chang

When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation

Read the original on arXiv AI →

The paper investigates how marketplace guardrails affect welfare in language‑model agent simulations of hotel transactions. It finds that initial reports of large welfare gains disappear when controlling for offer schemas and buyer choice, and that guardrails mainly redistribute rather than increase welfare unless sellers are explicitly forced to produce inefficient bundles. The authors propose a construct‑validity framework to flag invalid or inconclusive policy claims before they are reported.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting

The paper examines how an agent’s probability report is evaluated twice—once by a strictly proper scoring rule and again by an approval rule that determines a decision. It shows that when the approval rule is welfare‑maximizing, it cannot be affine, yet the resulting distortion is predictable and can be mitigated by a reserve report that neutralizes the cost of pretending to be the marginal type. A Lipschitz rule with a single kink achieves first‑best welfare, while smooth rules cannot, and the key constraint is the steepness of the rule rather than its smoothness.

By Lauri Lov\'en, Sasu Tarkoma