Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces inexpensive, scalable methods for evaluating language model behavior across different vendors and releases. By running identical public stimuli on a cross‑vendor panel and analyzing transcripts via exact match, LLM‑coded codebooks, or instrumented environments, the authors can quantify model responses at a cost of a few dollars per model. Applying these tools to four years of releases reveals patterns of convergence, resistance, positional stability, and compliance that vary by generation, lab, and harness.
arXiv:2602. 16111v2 Announce Type: replace-cross Abstract: Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set guardrails in A/B experiments.
The study evaluates how well large language models (LLMs) can annotate missing metadata in PubChem’s ~2 million bioassays, focusing on BioAssay Ontology (BAO) assay format and detection method fields. It finds that 36 % of assays lack an assay format, 89 % lack a BioAssay type, and over 99.9 % lack any BAO‑mapped format or detection technology term, highlighting a critical sparsity in metadata. Seven open‑source and proprietary LLMs achieve recall ≥0.96 for biochemical and cell‑based assay formats, with similar performance for detection technology, though disagreements rise for under‑represented classes and often stem from inconsistencies in silver labels rather than LLM errors. A qualitative test with an industrial curator shows LLM‑generated evidence can prompt revisions of existing labels, indicating that LLMs can flag potentially mislabeled assays.
The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.
arXiv:2606. 31478v1 Announce Type: new Abstract: Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail.
arXiv:2601. 14429v2 Announce Type: replace-cross Abstract: Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated.