Goodhart’s law famously says: “When a measure becomes a target, it ceases to be a good measure. ” Although originally from economics, it’s something we have to grapple with at OpenAI when figuring out how to optimize objectives that are difficult or costly to measure.
Learn how OpenAI evaluates political bias in ChatGPT through new real-world testing methods that improve objectivity and reduce bias.
Assistant Professor Bailey Flanigan has arrived at complex computational methods for helping democracy thrive.
By Michaela Jarvis | MIT Laboratory for Information and Decision Systems
arXiv:2608. 16016v1 Announce Type: cross Abstract: Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of conventional assessments that rely on single-point submissions and product-only grading.
By Rajan Kadel, Bellal Hossain, Samar Shailendra, Bushra Naeem
arXiv:2608. 00961v2 Announce Type: replace-cross Abstract: AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction.
By Donna M Bye, Levin Kuhlmann
The paper argues that computational text‑based ideal point estimation (CT‑IPE) methods should be viewed as configurable measurement pipelines rather than fixed estimators. It presents a large‑scale comparative experiment involving 17 CT‑IPE algorithms, 5,537 runs, and about 4.25 million left‑right position estimates, and describes shared infrastructure that enables joint execution of these heterogeneous methods. Sensitivity analyses reveal that most algorithms exhibit low hyperparameter sensitivity (ICC < .10), with any remaining sensitivity concentrated in a few key researcher choices such as the language or embedding model, seed keyword lists, and number of topics.
By Patrick Parschan
A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al.
arXiv:2606. 00040v1 Announce Type: cross Abstract: As Generative AI (GenAI) becomes integral to education, fostering GenAI literacy is critical.
By Angxuan Chen, Jiyou Jia
LLM arenas turn pairwise human preferences into model rankings. Those preferences may reflect how an answer is presented as well as what it says. We take a stylometric approach to 137,293 decisive Fre...
arXiv:2606. 04152v1 Announce Type: new Abstract: Large language models are reshaping research practice while quietly eroding researchers epistemic accountability.
By Clarisse de Souza, Gabriel Barbosa, Simone Diniz Junqueira Barbosa, B\'arbara Betts, Renato Cerqueira, Juliana Jansen Ferreira