Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces inexpensive, scalable methods for evaluating language model behavior across different vendors and releases. By running identical public stimuli on a cross‑vendor panel and analyzing transcripts via exact match, LLM‑coded codebooks, or instrumented environments, the authors can quantify model responses at a cost of a few dollars per model. Applying these tools to four years of releases reveals patterns of convergence, resistance, positional stability, and compliance that vary by generation, lab, and harness.
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
arXiv:2607. 10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model "value dispositions," and a concentration/extremity index over repeated draws is read as how sharply a model commits.
arXiv:2609.13936v1 Announce Type: new Abstract: Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This repo...
arXiv:2607. 18310v1 Announce Type: cross Abstract: Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent.
arXiv:2608. 06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory.