arXiv Machine Learning

Large language model-enabled automated data extraction for concrete materials informatics

arXiv:2604. 22938v2 Announce Type: replace-cross Abstract: The promise of data-driven materials discovery remains constrained by the scarcity of large, high-quality, and accessible experimental datasets.

arXiv AI
Sep 16

Extracting ontology-compliant knowledge from scientific text describing irradiated materials using large language models

arXiv:2609.17291v1 Announce Type: new Abstract: The quest for new materials increasingly relies on predictive models and comprehensive simulations that span scales from atomic to macroscopic levels....

By Marco Luca Sbodio, Marcos Mart\'inez Galindo, Vanessa Lopez, Blanca Biel, Pablo Canca, Pedro Delgado, Jes\'us I. Mendieta-Moreno, Raphael Tack, Maria J. Caturla
arXiv Machine Learning
Sep 3

When Literature Data Mislead Artificial Intelligence in Materials Discovery

The article examines how scientific literature, often used as a data source for AI in materials science, can contain hidden inaccuracies such as text-figure mismatches, ambiguous axis labels, unit inconsistencies, and missing measurement context. By tracing solid electrolyte conductivity values from original papers to curated datasets, the authors uncover recurrent errors that are numerically plausible yet hard to detect, leading to significant label noise in AI models. A cross-database example demonstrates that ambiguous reporting can cause a 100‑fold error in conductivity values, underscoring the need for traceable reporting, rigorous curation, and validation practices in AI-driven discovery.

By Qian Wang, Ying Li, Ryuhei Sato, Hidemi Kato, Shin-ichi Orimo, Hao Li, Eric Jianfeng Cheng
arXiv AI
Aug 24

MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature

MatMMExtract is an open‑source pipeline that disassembles compound scientific figures into individual sub‑panels and generates structured, grounded image‑text pairs using a large language model guided by a materials science taxonomy. Applied to 14,810 open‑access articles, it produced 391,606 panel‑level pairs with sub‑captions, a two‑level visualisation category (19 classes, 100+ subtypes), and scientific summaries. The project also introduces MaterialScope, a 2,811‑figure detection dataset, and demonstrates that Gemini 3.1 Flash Lite yields high‑quality annotations with low hallucination, while a dual‑encoder baseline outperforms zero‑shot CLIP on the resulting MatSciFig dataset.

By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
arXiv Machine Learning
Sep 21

LLMs as Feature Engineers for Text-and-Tabular Prediction

The paper presents an iterative framework that uses large language models (LLMs) to automatically extract interpretable, schema‑bound categorical features from unstructured text for use in tabular prediction models. A generator LLM proposes semantic definitions, an extractor LLM materializes the features, and a downstream tabular model evaluates their predictive performance, with error‑driven natural‑language feedback guiding the search. Across three public datasets, the error‑driven loop speeds up feature discovery up to three times and the resulting features outperform any subset when combined with TF‑IDF and dense embeddings, while also providing instance‑level interpretability through SHAP importance rankings and a semantic audit trail.

By Merwan Barlier, Blaz Skrlj
arXiv Machine Learning
4d ago

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...

By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko