arXiv AI By Tanjin He, Aikaterini Vriza, Logan Ward, Xu Huang, Yiming Chen, Anubhav Jain, Gerbrand Ceder, Rajeev S. Assary, Ian T. Foster, Maria K. Y. Chan

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

Read the original on arXiv AI →

arXiv:2607. 23886v1 Announce Type: cross Abstract: X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Prototype-guided transfer of sparse literature knowledge for electrolyte additive discovery

ProtoMI is a literature‑driven framework that learns structural priors from 126 reported boron‑containing electrolyte additives and applies them to screen 179,977 unlabeled candidates. Using graph contrastive learning, it identifies seven interpretable prototypes and adapts them through semi‑supervised contrastive learning, achieving enrichment factors of 9.2–45.6 while evaluating less than 2% of the candidate space. The method led to the discovery of four commercially accessible additives, including TNDB, which improves high‑temperature LiFePO4||graphite cycling by 34.93% and forms protective interphases that suppress solvent decomposition and Fe deposition.

By Weixiang Hong, Hongting Du, Jiayue Tang, Ruifeng Tan, Yangjian Quan, Jia Li, Jiaqiang Huang
arXiv Machine Learning
Jul 10

MatBind: A Shared Embedding Space for Multimodal Materials Characterization

arXiv:2607. 08470v1 Announce Type: new Abstract: Fully characterizing a crystalline material requires integrating heterogeneous data sources -- atomic structures, diffraction patterns, electronic density of states, and natural language -- each of which captures a different facet of the same physical object.

By Le Yang (Institute for Advanced Simulations), Anoop K. Chandran (J\"ulich Supercomputing Centre, Forschungszentrum J\"ulich), Jona \"Ostreicher (Institute of Nanotechnology, Karlsruhe Institute of Technology), Evgenii Sovetkin (J\"ulich Supercomputing Centre, Forschungszentrum J\"ulich), Adrian Mirza (Helmholtz-Zentrum Berlin f\"ur Materialien und Energie, Helmholtz Institute for Polymers in Energy Applications Jena), Sebastien Bompas (Institute for Advanced Simulations), Bashir Kazimi (Institute for Advanced Simulations), Pascal Friederich (Institute of Nanotechnology, Karlsruhe Institute of Technology), Stefan Kesselheim (J\"ulich Supercomputing Centre, Forschungszentrum J\"ulich, 1. Physikalisches Institut, University of Cologne), Kevin Maik Jablonka (Helmholtz Institute for Polymers in Energy Applications Jena, Center for Energy and Environmental Chemistry Jena, Friedrich Schiller University Jena), Stefan Sandfeld (Institute for Advanced Simulations, Faculty 5 - Georesources and Materials Engineering, RWTH Aachen University)
arXiv Machine Learning
Jul 3

IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

arXiv:2607. 01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electrochemical diagnostics, second-life analytics, and battery safety research.

By Sakthi Prabhu Gunasekar, Prasanna Kumar Rangarajan
arXiv Machine Learning
Sep 3

When Literature Data Mislead Artificial Intelligence in Materials Discovery

The article examines how scientific literature, often used as a data source for AI in materials science, can contain hidden inaccuracies such as text-figure mismatches, ambiguous axis labels, unit inconsistencies, and missing measurement context. By tracing solid electrolyte conductivity values from original papers to curated datasets, the authors uncover recurrent errors that are numerically plausible yet hard to detect, leading to significant label noise in AI models. A cross-database example demonstrates that ambiguous reporting can cause a 100‑fold error in conductivity values, underscoring the need for traceable reporting, rigorous curation, and validation practices in AI-driven discovery.

By Qian Wang, Ying Li, Ryuhei Sato, Hidemi Kato, Shin-ichi Orimo, Hao Li, Eric Jianfeng Cheng
arXiv Machine Learning
Sep 14

Accelerating battery research with an interoperable interface between FINALES and Kadi4Mat

The paper presents a methodological framework that links the FINALES experiment orchestration system with the Kadi4Mat research data management ecosystem to create interoperable, automated workflows for battery research. By coordinating experiment planning, execution, data handling, and analysis across distributed sites, the framework enables reproducible, end‑to‑end experimental pipelines. The authors demonstrate its utility by studying sodium‑ion coin cell formation, using a Gaussian process model to map formation parameters to electrochemical performance and identify promising experimental regions.

By Giovanna Tosato (Karlsruhe Institute of Technology), Leon Merker (Karlsruhe Institute of Technology, Helmholtz Institute Ulm, Technical University of Munich), Monika Vogler (Technical University of Munich), Michael Selzer (Karlsruhe Institute of Technology), Arnd Koeppe (Karlsruhe Institute of Technology)