arXiv Machine Learning

A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes

arXiv Computer Vision
Aug 25

SketchFlow: Zero-Shot Vector Sketch Generation via GMM Prior Flow in CLIP Latent Space

SketchFlow is a new generative framework for creating high‑quality vector sketches from text prompts. It uses a Gaussian Mixture Model prior in the CLIP latent space and an Optimal Transport Conditional Flow Matching model to map this prior to sketch features, which are then decoded by a Hybrid Diffusion Decoder combining 1D U‑Net and Transformer architectures. The approach achieves superior visual quality and human‑like drawing styles, and supports zero‑shot synthesis for unseen concepts and smooth semantic interpolation.

By Jin Zhou, Hongliang Yang, Pengfei Xu, Hui Huang
arXiv Computer Vision
Sep 2

Beyond Landmark Extraction: A Framework for Robust Geometric Feature Construction in Structured Image Classification

The paper argues that in structured image classification, the key question is what information a classifier should receive before making a prediction, rather than which algorithm performs best. It proposes a systematic framework for constructing landmark-derived representations—such as coordinate, distance, angle, and hybrid features—and evaluates them on static hand gesture recognition. Experiments show that hybrid representations, which combine complementary geometric components, outperform raw coordinate features and other single-type representations, highlighting the importance of thoughtful feature construction.

By Saravana Mauree, Sakshi Arya
Hugging Face Trending Papers
Aug 10

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions.

arXiv Computer Vision
Sep 17

CALIPER: Metric-Grounded Model-Free Recognition of Visually Similar Industrial Parts

CALIPER is a model‑free RGB‑D framework that performs fine‑grained recognition of visually similar industrial parts by combining support‑based appearance matching with metric size evidence. Each class is onboarded from a single turntable RGB‑D video and a few labeled real images, enabling 3D reconstruction for appearance support and depth‑aligned size profiling. At inference, a YOLOv8n‑seg model localizes parts, a frozen DINOv2 backbone with an episodically trained embedding head matches support, and margin‑conditioned metric fusion selectively uses size evidence for ambiguous cases, achieving high accuracy on 18 parts and robust enrollment of unseen screws without retraining.

By Alankrit Gupta, Chenxi Tao, Seung-Kyum Choi
arXiv AI
Jun 11

LASA: A Weak Supervision Method for Open-Vocabulary Scene Sketch Semantic Segmentation

arXiv:2606. 11837v1 Announce Type: cross Abstract: Open-vocabulary scene sketch semantic segmentation aims to assign dense semantic labels to sparse line drawings based on flexible category vocabularies specified at inference time, without relying on pixel-level annotations during training.

By Liwen Yi, Xianlin Zhang, Yue Zhang, Yue Ming, Xueming Li