Interpretable and pedagogical examples
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
We’ve designed a method that encourages AIs to teach each other with examples that also make sense to humans.
arXiv:2606. 12289v1 Announce Type: cross Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling their computations.
The paper presents ESSE, a self‑explanation tutor that uses a large language model to give immediate feedback on students’ line‑by‑line explanations of introductory programming worked examples. It evaluates the LLM’s judgments against a domain expert and a crowd of non‑experts, finding that the model is reliable enough to serve as the tutor’s assessment engine. In an introductory Java course, the tutor’s feedback encourages students to persist, improves the completeness and conceptual depth of their explanations, and shows evidence of learning.
arXiv:2606. 26523v1 Announce Type: new Abstract: We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tools of mechanistic interpretability.
arXiv:2606. 26228v1 Announce Type: cross Abstract: We review the concepts of interpretability and explainability as they apply to machine learning in physics.