DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 28406v1 Announce Type: new Abstract: Text-to-image and multimodal generative models are increasingly used to produce scientific figures such as mechanism diagrams, experimental-design schematics, conceptual frameworks, and graphical abstracts.
We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scientific image tasks require not only high-fidelity synthesis, but also robust understanding of scientific semantics, structural relations, domain knowledge, and task intent.
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery. Inspired by Peirce's Semiotic Triad, we introduce Scientific Image Reasoning (SciIR), a comprehensive resource for training and evaluation of scientific image generation.
The paper introduces a capability‑centric data infrastructure for generalist image generation, integrating task‑specific supervision with a curriculum that aligns with the dependencies among generative capabilities. It employs three interoperable data engines—text‑image grounding, inter‑image transformation, and image‑knowledge association—alongside caption experts to harmonize text‑to‑image and editing supervision. The system curates massive corpora (440M T2I images, 120M editing pairs, 27M image‑entity pairs) and trains multimodal diffusion models (3B and 6B parameters) from scratch, achieving broad visual coverage and versatile rendering as shown by CPI‑Bench and qualitative tests.
arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
arXiv:2601. 04390v2 Announce Type: replace Abstract: High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create.