arXiv Computation and Language

Bye Bye Perspective API: Lessons for Building and Governing Measurement Infrastructure

arXiv AI
Sep 25

Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases

The paper introduces inexpensive, scalable methods for evaluating language model behavior across different vendors and releases. By running identical public stimuli on a cross‑vendor panel and analyzing transcripts via exact match, LLM‑coded codebooks, or instrumented environments, the authors can quantify model responses at a cost of a few dollars per model. Applying these tools to four years of releases reveals patterns of convergence, resistance, positional stability, and compliance that vary by generation, lab, and harness.

By Tapan Parikh
arXiv AI
4d ago

Can LLMs Reliably Annotate Bioassay Metadata to Improve Data Readiness?

The study evaluates how well large language models (LLMs) can annotate missing metadata in PubChem’s ~2 million bioassays, focusing on BioAssay Ontology (BAO) assay format and detection method fields. It finds that 36 % of assays lack an assay format, 89 % lack a BioAssay type, and over 99.9 % lack any BAO‑mapped format or detection technology term, highlighting a critical sparsity in metadata. Seven open‑source and proprietary LLMs achieve recall ≥0.96 for biochemical and cell‑based assay formats, with similar performance for detection technology, though disagreements rise for under‑represented classes and often stem from inconsistencies in silver labels rather than LLM errors. A qualitative test with an industrial curator shows LLM‑generated evidence can prompt revisions of existing labels, indicating that LLMs can flag potentially mislabeled assays.

By Laura van Weesep, Riccardo Tedoldi, Jens Sj\"olund, Hossein Azizpour, Susanne Winiwarter, Ola Engkvist, Jon Paul Janet, Samuel Genheden, Juan Viguera Diez
arXiv AI
Aug 25

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

The paper argues that large language models (LLMs) are evaluated too narrowly, focusing on isolated technical metrics rather than holistic, developmental, and societal aspects. It proposes a diagnostic ontology that links evaluation dimensions to the LLM training pipeline, turning evaluation into a root‑cause analysis tool. The authors introduce an anthropomorphic framework—IQ, PQ, EQ, and VQ—to assess LLM capabilities, operationalize it with a modular architecture, and validate it through meta‑analysis of over 200 benchmarks, outlining key challenges and future directions.

By Jun Wang, Ninglun Gu, Kailai Zhang, Pengyong Li, Yelun Bao, Jin Yang, Xu Yin, Liwei Liu, Zijiao Zhang, Yihuan Liu, Gary G. Yen, Junchi Yan
arXiv AI
Jul 29

Measuring the State of Open Science in Transportation Using Large Language Models

arXiv:2601. 14429v2 Announce Type: replace-cross Abstract: Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated.

By Junyi Ji, Ruth Lu, Linda Belkessa, Liming Wang, Silvia Varotto, Yongqi Dong, Nicolas Saunier, Mostafa Ameli, Gregory S. Macfarlane, Bahman Madadi, Cathy Wu
arXiv AI
Aug 6

ORCA-bench: How Ready Are Language Model Agents for Oncall?

arXiv:2607. 28545v2 Announce Type: replace-cross Abstract: Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.

By Albert Gong, Kyuseong Choi, Abhineet Agarwal, Jason Schechner, Ryan Huang, Raj Agrawal, Anish Agarwal, Raaz Dwivedi
arXiv AI
Jun 9

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.

By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
arXiv Computation and Language
Sep 18

BurnRiSc: Toward Non-Invasive Burnout Screening in Open Source from Public Repository Signals

BurnRiSc is a framework that uses 14 behavioral and linguistic signals derived from GitHub activity to compute a monthly Burnout Risk Score (BRS) for open‑source contributors. The scores are based on the Oldenburg Burnout Inventory’s exhaustion and disengagement dimensions and are weighted using labeled cases. In a preliminary study of 68 contributors across ten repositories, sustained BRS elevation predicted 6 of 10 disclosed burnout cases 6–15 months in advance, and 10 of 10 when considering peak BRS as a second criterion.

By Timofey Sanko, Yuan Tian, Mariam Guizani