arXiv:2606. 29667v1 Announce Type: cross Abstract: The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale.
By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
MatMMExtract is an open‑source pipeline that disassembles compound scientific figures into individual sub‑panels and generates structured, grounded image‑text pairs using a large language model guided by a materials science taxonomy. Applied to 14,810 open‑access articles, it produced 391,606 panel‑level pairs with sub‑captions, a two‑level visualisation category (19 classes, 100+ subtypes), and scientific summaries. The project also introduces MaterialScope, a 2,811‑figure detection dataset, and demonstrates that Gemini 3.1 Flash Lite yields high‑quality annotations with low hallucination, while a dual‑encoder baseline outperforms zero‑shot CLIP on the resulting MatSciFig dataset.
By Subham Ghosh, Shubham Tiwari, Mohammad Ibrahim, Abhishek Tewari
The paper investigates using manufacturer catalogue photography to bootstrap a computer‑vision system for recognizing carbide rotary burrs, a task that traditionally relies on manual quality checks. It shows that while frozen feature extractors struggle to separate key attributes like head shape and tooth profile, metric learning can cluster catalogue images almost perfectly, yet only about half of this performance transfers to real field photographs. The study finds that simple domain‑sensitivity reductions—such as converting images to grayscale and applying Hungarian assignment based on order sheets—yield the largest gains, suggesting catalogue images are a useful cold‑start source rather than a ready‑for‑deployment training set.
By Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju, Nicolas Pyschny, Felix Hackel\"oer, Florian Zwanzig
arXiv:2604. 26633v2 Announce Type: replace-cross Abstract: Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, and collecting balanced training sets is slow and costly.
By Paul Julius K\"uhn, Mika Pommeranz, Arjan Kuijper, Saptarshi Neil Sinha
The paper investigates using catalogue photographs to bootstrap a computer‑vision system for recognizing rotary milling tools (carbide burrs) in industrial settings. It shows that standard frozen feature extractors fail to separate key attributes, while metric learning yields excellent clustering on catalogue images but only half the accuracy on real field photos. The study finds that simple domain‑sensitivity reductions—grayscale conversion and order‑sheet‑constrained retrieval—yield the largest transfer gains, positioning catalogue photography as a useful cold‑start rather than a ready‑to‑deploy training domain.
By Abilash Philip Madavath, Chandra Yuvesh Aubeeluck, Augustin Raju, Nicolas Pyschny, Felix Hackel\"oer, Florian Zwanzig
JEPAMatch introduces a new semi‑supervised learning framework that replaces traditional output‑thresholding with explicit geometric shaping of latent representations. By combining the FlexMatch loss with a latent‑space regularization inspired by LeJEPA, the method encourages isotropic Gaussian structure in the embedding space, mitigating class imbalance and noisy pseudo‑labels. Experiments on CIFAR‑100, STL‑10, and Tiny‑ImageNet show consistent performance gains and faster convergence compared to existing FixMatch‑based baselines.
By Ali Aghababaei-Harandi, Aude Sportisse, Massih-Reza Amini
This paper presents the IEEE International Conference on Multimedia and Expo (ICME) 2026 Grand Challenge on Cross-Scenario Defect Detection and Fine-Grained Severity Grading for High-Precision Manufacturing. The challenge is motivated by two key limitations of existing industrial defect inspection systems: (1) current deep learning-based methods often suffer significant performance degradation when deployed in unseen production scenarios, and (2) most benchmarks neglect severity-aware assessment, which is critical for risk control and yield optimization.
The paper introduces an AI-assisted pipeline for retrieving goldsmith marks from silverware images, combining mark localization with metric‑learning fine‑tuning on three backbone models (ResNet‑50, ViT‑S/16, and DINOv2 ViT‑S/14). Systematic tests of cropping strategies show that manual cropping and metric‑learning fine‑tuning yield the best performance, with DINOv2 ViT‑S/14 achieving 62.63% mAP and 73.74% Top‑1 accuracy. The authors release a manually annotated dataset, code, and a public web interface to support reproducibility and adoption in digital humanities.
By Atmik Tiwari, Vincent Christlein, Mark Fichtner, Freya Gohlke, Birgit Sch\"ubel, Theresa Witting, Heike Zech, Mathias Zinnen
arXiv:2607. 28695v1 Announce Type: cross Abstract: Here is the plain text version optimized for arXiv's submission form.
By Aryuemaan Kumar Chowdhury
arXiv:2508.21424v3 Announce Type: replace
Abstract: Deep learning models have achieved state-of-the-art performance in many computer vision tasks. However, in real-world scenarios, novel classes that...
By Lucas Rakotoarivony
arXiv:2608.31052v1 Announce Type: cross
Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
By Keith G. Mills, Evan B. Sanders, Gregory J. Matthews, Juliet K. Brophy
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber