Hugging Face Blog
Nov 23, 2022
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
An expert in machine learning, statistics, and computation, Rakhlin succeeds Professor Ankur Moitra.
arXiv:2408. 02379v2 Announce Type: replace-cross Abstract: Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming regulation such as the EU AI Act.
arXiv:2607. 09668v1 Announce Type: new Abstract: Ground truth datasets play a fundamental role as reference values in the training and evaluation of machine learning models.