Announcing New Dataset Search Features
Related stories
On the missing data layer and a potential solution
arXiv:2608. 02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer.
SearchGPT is a prototype of new AI search features
We’re testing SearchGPT, a temporary prototype of new search features that give you fast and timely answers with clear and relevant sources.
Invent a Dataset: Measuring dataset generation abilities with zero seed
Invent-A-Dataset is a prompt‑based system that generates realistic, large‑scale datasets from a description, targeting the zero‑data regime where no initial data exists. The authors benchmarked it against five leading model APIs across eight task types and up to 20,000 samples, finding that it outperforms competitors with 17% higher quality and 19% more diverse samples. The diversity advantage grows with dataset size, leading to better downstream training performance and consistently higher rankings for fine‑tuned models.
Introducing the Data Measurements Tool: an Interactive Tool for Looking at Datasets
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
arXiv:2607. 05970v1 Announce Type: cross Abstract: Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems.
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
arXiv:2607. 00710v1 Announce Type: cross Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones.
Introducing RTEB: A New Standard for Retrieval Evaluation
Introducing Search Toolkit
Search Toolkit is a composable framework for building production search pipelines for AI applications.
Search Strategies for Optimal Classification and Regression Trees
arXiv:2607. 28170v1 Announce Type: new Abstract: Optimal decision trees (ODTs) are compact, interpretable machine learning models that globally optimize a given objective, but their scalability remains challenging.
Bringing Agentic Search to Earth Observation Data Discovery
arXiv:2607. 02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony.