A Lightweight Convolutional Neural Network for Real-Time Recognition of Hand-Drawn Geometric Shapes
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
SketchFlow is a new generative framework for creating high‑quality vector sketches from text prompts. It uses a Gaussian Mixture Model prior in the CLIP latent space and an Optimal Transport Conditional Flow Matching model to map this prior to sketch features, which are then decoded by a Hybrid Diffusion Decoder combining 1D U‑Net and Transformer architectures. The approach achieves superior visual quality and human‑like drawing styles, and supports zero‑shot synthesis for unseen concepts and smooth semantic interpolation.
arXiv:2607.05568v2 Announce Type: replace-cross Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
The paper argues that in structured image classification, the key question is what information a classifier should receive before making a prediction, rather than which algorithm performs best. It proposes a systematic framework for constructing landmark-derived representations—such as coordinate, distance, angle, and hybrid features—and evaluates them on static hand gesture recognition. Experiments show that hybrid representations, which combine complementary geometric components, outperform raw coordinate features and other single-type representations, highlighting the importance of thoughtful feature construction.
Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers. They convert tabular data into image representations, mapping each feature at a fixed pixel location derived from a dimensionality-reduction method (e.
Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions.
arXiv:2604.22875v3 Announce Type: replace-cross Abstract: When answering questions about images, humans naturally point, label, and draw to explain their reasoning. In contrast, modern vision-languag...