`LeRobotDataset:v3.0`: Bringing large-scale datasets to `lerobot`
Related stories
LeRobot goes to driving school: World’s largest open-source self-driving dataset
LeRobot v0.6.0: Imagine, Evaluate, Improve
SmolVLA: Efficient Vision-Language-Action Model trained on Lerobot Community Data
LeRobot v0.4.0: Supercharging OSS Robot Learning
Bringing Robotics AI to Embedded Platforms: Dataset Recording, VLA Fine‑Tuning, and On‑Device Optimizations
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of real-world development. Such environments typically involve both complex code and large-scale data (i.
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics
arXiv:2607. 04146v1 Announce Type: cross Abstract: This work establishes that trigger-word data poisoning of vision language action models is practical, while at the same time the open-source robotics ecosystem holds trust assumptions about community contributions.
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
arXiv:2606. 15300v1 Announce Type: new Abstract: Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of real-world development.
VersaDB: A High-Performance AI Storage Database for Unifying Mutimodal Datasets
arXiv:2608.22795v1 Announce Type: new Abstract: The AI field has been rapidly developing, leading to the emergence of a large number of AI training datasets of various types. These datasets contain d...
SciDER: Scientific Data-centric End-to-end Researcher
arXiv:2603. 01421v3 Announce Type: replace Abstract: While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data.
NeMo Data Designer: An Extensible Framework for Multimodal Synthetic Data Generation
NeMo Data Designer (NDD) is an open‑source framework for generating multimodal synthetic data. It uses a declarative configuration format that lets users define dataset columns—text, code, structured outputs, images, embeddings, and statistical samplers—to steer diversity. The system supports a preview‑and‑revision workflow, dependency resolution, and retry logic, and can be extended via plugins. Case studies demonstrate its use for structured, agentic, multimodal, and domain‑specialized tasks, including datasets for Nemotron model development and enterprise deployments.