arXiv Computer Vision

Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation

arXiv Computation and Language
Sep 25

Multimodal Thinking with Renderable Programs

The paper introduces SVGLM, a framework that integrates scalable vector graphics (SVG) primitives into vision‑language models to enable image generation within reasoning tasks. By treating SVG both as image descriptions and text instructions, SVGLM offers a compact and interpretable method for connecting text and image reasoning. The authors provide a curated SVG‑based image editing dataset and demonstrate strong SVG generation and image‑aware reasoning performance on a mathematical benchmark.

By Sunli Chen, Ding Zhong, Ziqiao Ma, Jiaxin Liu, Zeyuan Yang, Hao Zhang, Lie Lu, Joyce Chai, Chuang Gan
arXiv AI
Jun 12

A Mathematical Forum Platform for Collaborative Problem Solving and Dataset Generation for AI Reasoning

arXiv:2606. 12976v1 Announce Type: new Abstract: Sharing mathematical content in online forums remains a significant friction point for students and educators: writing raw LATEX is error-prone, standalone optical character recognition tools require platform switching, and current forum software offers no integrated path from a photograph of a formula to a rendered post.

By Akbar Erkinov, Nurmukhammad Abdurasulov
arXiv Computer Vision
4d ago

Back2Struct: Making Structured Images Editable Again

Back2Struct is a system that converts structured images—such as diagrams, charts, and flowcharts—into editable vector graphics code (SVG/XML). By predicting semantically rich, object-level SVG code rather than low-level pixel vectorization, it allows the generated graphics to be imported into tools like PowerPoint for easy editing, restyling, and reuse. The model is trained with supervised fine‑tuning and reward‑based learning that enforces syntactic validity, concise length, and visual fidelity to the input, leading to higher accuracy, editability, and user alignment compared to baselines.

By Pengyu Yan, Yixin Wu, Yunjie Tian, David Doermann
arXiv Machine Learning
Aug 28

Chart2SVG: Editable SVG Generation from Raster Chart Images

Chart2SVG is a multimodal large language model that transforms static raster chart images into editable SVGs enriched with semantic structure. By embedding chart‑specific semantic tokens into a vision‑language framework and training on the Beagle+ dataset of 33K distilled chart samples, the model captures both geometric primitives and their functional roles. The resulting SVGs are visually accurate and structurally consistent, and the accompanying Chart Structure Graph (CSG) exposes visual dependencies for interactive exploration, chart repurposing, and layout reuse.

By Jinning Cui, Lu Chen, Haoyan Shi, Yue He, Chenglong Wang, Mengyu Zhou, Weidong Huang, Yunhai Wang
arXiv AI
Sep 2

VectorGym: A Multi-Task Benchmark for SVG Code Generation, Sketching and Editing

arXiv:2603.29852v2 Announce Type: replace-cross Abstract: We introduce VectorGym, a comprehensive benchmark suite for Scalable Vector Graphics (SVG) that spans generation from text and sketches, comp...

By Joan Rodriguez, Haotian Zhang, Abhay Puri, Haoran Dai, Tianyang Zhang, Meng Lin, Rishav Pramanik, Xiaoqing Xie, Marco Terral Rodriguez, Darsh Kaushik, Aly Shariff, Perouz Taslakian, Spandana Gella, Sai Rajeswar, David Vazquez, Christopher Pal, Marco Pedersoli