arXiv:2606. 29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale.
By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
MatMMExtract is an open‑source pipeline that disassembles compound scientific figures into individual sub‑panels and generates structured, grounded image‑text pairs using a large language model guided by a materials science taxonomy. Applied to 14,810 open‑access articles, it produced 391,606 panel‑level pairs with sub‑captions, a two‑level visualisation category (19 classes, 100+ subtypes), and scientific summaries. The project also introduces MaterialScope, a 2,811‑figure detection dataset, and demonstrates that Gemini 3.1 Flash Lite yields high‑quality annotations with low hallucination, while a dual‑encoder baseline outperforms zero‑shot CLIP on the resulting MatSciFig dataset.
By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
The paper investigates using manufacturer catalogue photography to bootstrap a computer‑vision system for recognizing carbide rotary burrs, a task that traditionally relies on manual quality checks. It shows that while frozen feature extractors struggle to separate key attributes like head shape and tooth profile, metric learning can cluster catalogue images almost perfectly, yet only about half of this performance transfers to real field photographs. The study finds that simple domain‑sensitivity reductions—such as converting images to grayscale and applying Hungarian assignment based on order sheets—yield the largest gains, suggesting catalogue images are a useful cold‑start source rather than a ready‑for‑deployment training set.
By Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju, Nicolas Pyschny, Felix Hackel\"oer, Florian Zwanzig
arXiv:2604. 26633v2 Announce Type: replace-cross Abstract: Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, and collecting balanced training sets is slow and costly.
By Paul Julius K\"uhn, Mika Pommeranz, Arjan Kuijper, Saptarshi Neil Sinha
The paper investigates using catalogue photographs to bootstrap a computer‑vision system for recognizing rotary milling tools (carbide burrs) in industrial settings. It shows that standard frozen feature extractors fail to separate key attributes, while metric learning yields excellent clustering on catalogue images but only half the accuracy on real field photos. The study finds that simple domain‑sensitivity reductions—grayscale conversion and order‑sheet‑constrained retrieval—yield the largest transfer gains, positioning catalogue photography as a useful cold‑start rather than a ready‑to‑deploy training domain.
By Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju, Nicolas Pyschny, Felix Hackel\"oer, Florian Zwanzig
JEPAMatch introduces a new semi‑supervised learning framework that replaces traditional output‑thresholding with explicit geometric shaping of latent representations. By combining the FlexMatch loss with a latent‑space regularization inspired by LeJEPA, the method encourages isotropic Gaussian structure in the embedding space, mitigating class imbalance and noisy pseudo‑labels. Experiments on CIFAR‑100, STL‑10, and Tiny‑ImageNet show consistent performance gains and faster convergence compared to existing FixMatch‑based baselines.
By Ali Aghababaei-Harandi, Aude Sportisse, Massih-Reza Amini