Towards Robust Sequential Decomposition for Complex Image Editing
arXiv:2605. 09233v2 Announce Type: replace-cross Abstract: Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions.
arXiv:2509. 24900v2 Announce Type: replace-cross Abstract: The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data.
arXiv:2605. 09233v2 Announce Type: replace-cross Abstract: Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions.
arXiv:2607. 13125v1 Announce Type: cross Abstract: We introduce Boogu-Image-0.
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion.
arXiv:2607. 13125v2 Announce Type: replace-cross Abstract: We introduce Boogu-Image-0.
arXiv:2607. 08201v1 Announce Type: cross Abstract: Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity.
arXiv:2605. 18714v2 Announce Type: replace-cross Abstract: Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture.
arXiv:2510. 08532v2 Announce Type: replace-cross Abstract: Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language.
arXiv:2607. 05465v1 Announce Type: cross Abstract: Complex image creation and editing often require more than a single generation or editing model.
arXiv:2608. 15006v1 Announce Type: cross Abstract: Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasoning.
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they cannot achieve interleaved generation (text-image sequence), which has crucial applications in visual narratives, guidance, and embodied manipulation.
arXiv:2606. 05058v1 Announce Type: cross Abstract: Computer-Aided Design (CAD) underpins modern engineering and manufacturing by enabling the creation of precise, editable 3D models.
arXiv:2606. 00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle.