The paper introduces an AI-assisted pipeline for retrieving goldsmith marks from silverware images, combining mark localization with metric‑learning fine‑tuning on three backbone models (ResNet‑50, ViT‑S/16, and DINOv2 ViT‑S/14). Systematic tests of cropping strategies show that manual cropping and metric‑learning fine‑tuning yield the best performance, with DINOv2 ViT‑S/14 achieving 62.63% mAP and 73.74% Top‑1 accuracy. The authors release a manually annotated dataset, code, and a public web interface to support reproducibility and adoption in digital humanities.
By Atmik Tiwari, Vincent Christlein, Mark Fichtner, Freya Gohlke, Birgit Sch\"ubel, Theresa Witting, Heike Zech, Mathias Zinnen
Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists. Computer vision offers a promising avenue for decipherment but requires large, densely annotated datasets.
arXiv:2607. 09104v1 Announce Type: cross Abstract: While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming.
By Camila Piscioneri Magalh\~aes, Lucas Pascotti Valem
The paper evaluates whether incorporating the hierarchical structure of Tironian notes can improve automatic recognition of this complex Latin shorthand system. Experiments compare flat classifiers (ResNet18, ConvNeXt, Swin, ViT) with hierarchy‑aware models (HD‑CNN and routing approaches) on handwritten and manuscript samples, with and without few‑shot adaptation. Results show that hierarchical models outperform flat ones when no adaptation is applied, but flat models surpass them after few‑shot adaptation, indicating that hierarchy can aid recognition under non‑adapted conditions.
The paper introduces Biomedica, an open-source dataset sourced from PubMed Central that includes over 6 million scientific articles and 24 million image‑text pairs, along with 27 metadata fields and expert human annotations. To facilitate use, the authors provide scalable streaming and search APIs via a web server. They demonstrate the dataset’s value by training embedding models, chat‑style models, and retrieval‑augmented chat agents, all of which outperform previous open systems in their categories.
By Alejandro Lozano, Min Woo Sun, James Burgess, Jeffrey J. Nirschl, Christopher Polzak, Yuhui Zhang, Liangyu Chen, Jeffrey Gu, Ivan Lopez, Josiah Aklilu, Anita Rau, Austin Wolfgang Katzer, Collin Chiu, Orr Zohar, Xiaohan Wang, Alfred Seunghoon Song, Chiang Chia-Chun, Robert Tibshirani, Serena Yeung-Levy
arXiv:2607. 04262v1 Announce Type: new Abstract: Convolutional Neural Network (CNN) and Vision Transformer (ViT) for image classification exploit a dense grid of pixels containing redundant information.
By Sarabeshwar Balaji, Shubham Mohanty, Akash Anil
arXiv:2509. 25289v4 Announce Type: replace-cross Abstract: Identifying an effective clustering algorithm for a given dataset remains a fundamental unsupervised learning issue.
By Mohammadreza Bakhtyari, Bogdan Mazoure, Renato Cordeiro de Amorim, Guillaume Rabusseau, Vladimir Makarenkov
arXiv:2512. 11982v2 Announce Type: replace-cross Abstract: Finding scientifically interesting phenomena through slow manual labeling campaigns severely limits our ability to explore the billions of galaxy images produced by telescopes.
By Nolan Koblischke, Liam Parker, Francois Lanusse, Jo Bovy, Irina Espejo, Shirley Ho
arXiv:2607. 00975v1 Announce Type: cross Abstract: Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis.
By Tong Shao, Hongshun Ling, Li Zhang, Jinjing Wu, Junke Wang, Yuan Gao, Fang Wang
With the emergence of various pre-trained vision and language models, computer vision is shifting from narrow-domain to open-domain recognition. The construction of a more powerful yet general keypoint detection (GKD) model to support diverse tasks has become increasingly important in the field.
arXiv:2507. 19702v1 Announce Type: cross Abstract: Identifying influential nodes in complex networks is a critical task with a wide range of applications across different domains.
By Mohammed A. Ramadhan, Abdulhakeem O. Mohammed
arXiv:2607. 17582v1 Announce Type: new Abstract: Approximate Nearest Neighbor Search (ANNS) plays a pivotal role in modern deep learning pipelines.
By Zheqi Shen, Jingbo Su, Zijin Wan, Yan Gu, Yihan Sun