CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 29677v1 Announce Type: new Abstract: Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata.
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
EviStreams is a live, open‑source, no‑code web platform that enables systematic review teams to control AI‑assisted data extraction at three stages: program design, field specification, and extracted predictions. Reviewers use a form builder to define typed fields, run extraction on PDFs, inspect AI‑generated values with supporting passages, and perform blinded dual review with adjudication to produce an auditable consensus export. An evaluation across four clinical corpora and three model families shows that extraction quality depends more on field specification than on the model choice.
arXiv:2606. 14516v1 Announce Type: new Abstract: AI evaluations are widely used for testing and understanding progress.
arXiv:2605. 28787v2 Announce Type: replace-cross Abstract: In the era of autonomous agents, machine-actionable data is critical for data-driven workflows.
arXiv:2606. 06462v1 Announce Type: new Abstract: Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance.