Argilla 2.4: Easily Build Fine-Tuning and Evaluation Datasets on the Hub — No Code Required
Related stories
Data is better together: Enabling communities to collectively build better datasets together using Argilla and Hugging Face Spaces
MLLM-DataEngine: Closing the Loop of Multimodal Instruction Tuning Data Generation
arXiv:2607. 15299v1 Announce Type: cross Abstract: In this paper, we propose MLLM-DataEngine, a novel closed-loop system that bridges data generation, model training, and evaluation.
ALF: An Active Learning Framework for Scientific Discovery
ALF is a modular active learning framework designed to streamline the entire data acquisition process for scientific discovery. It offers a single API that supports both offline benchmarking against existing datasets and online deployment with an oracle for real‑world candidate acquisition. The framework is open‑source and available on GitHub.
DuckDB: analyze 50,000+ datasets stored on the Hugging Face Hub
Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets
Welcome RL Environments to the hub
nanoVLM: The simplest repository to train your VLM in pure PyTorch
Scikit-fingerprints: Python library for scikit-learn compatible molecular fingerprints and chemoinformatics
arXiv:2608. 02027v1 Announce Type: new Abstract: We present scikit-fingerprints, a comprehensive, fully scikit-learn compatible library for molecular machine learning in Python, based on RDKit.
TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins
arXiv:2606. 17660v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hyperparameter choices, and na\"ive runs can even degrade model performance.
Equally Good, Yet Different: Benchmarking Rashomon sets in AutoML packages
arXiv:2609.36970v1 Announce Type: new Abstract: The Rashomon effect describes the existence of multiple near-optimal models that achieve comparable performance while offering fundamentally different...
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
arXiv:2607. 00710v1 Announce Type: cross Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones.