OpenAI Blog

Interpretable machine learning through teaching

We’ve designed a method that encourages AIs to teach each other with examples that also make sense to humans.

arXiv Computer Vision
Sep 7

From Interpretability Methods to Interpretable Models

The paper argues that explainable AI for computer vision has focused too much on developing interpretability methods rather than assessing how interpretable the models themselves are. It proposes a shift toward model-centric evaluation, using existing tools to compare what different models represent and compute, and emphasizes the need to measure whether humans can truly understand these models. The authors review the current toolbox, survey limited model comparison work, draw parallels to systems neuroscience, and outline a future agenda for model-focused XAI.

By Julien Colin, Nuria Oliver, Thomas Serre
arXiv Computation and Language
Sep 1

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.

By Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce
arXiv AI
Sep 25

Beyond Simple Input-Output Assessment Tasks: Leveraging Automated Programming Assessment for Non-Trivial Courses

The article discusses how machine learning exercises can be designed for automated assessment tools, framing them as deterministic input-output tasks. It emphasizes that this approach does not create a new grading system but enables existing platforms (e.g., VPL for Moodle, Codeforces, MOJ) to support AI education more effectively. The authors argue that integrating theory with practice through such exercises can foster dynamic, interactive AI courses.

By Artur Jordao