Leveraging Color Naming for Image Enhancement
arXiv:2607. 08185v1 Announce Type: cross Abstract: Enhancing images to make them visually appealing is a persistent challenge in computer vision.
arXiv:2607. 08185v1 Announce Type: cross Abstract: Enhancing images to make them visually appealing is a persistent challenge in computer vision.
Paint-Anything introduces a unified hex-prompt interface that allows users to specify any 24‑bit hex color for both image generation and editing. The method trains on a new Paint‑500K dataset created from real images with object grounding, perceptual color labeling, and editing‑pair synthesis, and supplements this with pure‑color anchors to address shadow‑induced color inaccuracies. Evaluated on the newly proposed Any Color Benchmark (ACBench), Paint‑Anything achieves significant improvements over the base FLUX.2‑4B model, boosting T2I and editing scores by 85.3 % and 28.3 % respectively, and outperforms competing methods on the CompColor metric.
arXiv:2609.23345v1 Announce Type: new Abstract: As generative AI becomes increasingly used in anime-style image creation, distinguishing human-drawn, AI-inpainted, and text-to-image images is importa...
arXiv:2606. 00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle.
Image style is a highly abstract, human-constructed concept shaped by a range of visual factors and intrinsically entangled with content, yet a unified and explicit definition of image style remains l...
arXiv:2607. 15482v1 Announce Type: new Abstract: The increasing complexity of state-of-the-art machine learning models has made their behavior progressively harder to interpret, spurring rapid advancements in the field of eXplainable Artificial Intelligence (XAI).
MegaStyle++ introduces a hierarchical definition of image style, ranging from overall style identity to fine‑grained visual attributes, to provide a more structured, transferable, and interpretable representation. Using this definition, the authors refined the MegaStyle annotation pipeline and released MegaStyle++‑8M, a dataset with 150K style identities, 1M fine‑grained prompts, and 8M stylized images. Analyses show that the hierarchical approach expands style diversity and semantic breadth while accurately capturing the intrinsic visual style of reference images.
arXiv:2608. 20334v1 Announce Type: new Abstract: We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing.
SenseNova-U1.5 is an 8B‑MoT native unified multimodal model that can understand, reason about, and generate visual content without using an encoder or VAE. It improves visual fidelity and text rendering through spatially coherent patch reconstruction, large‑scale training on curated generation and editing data, and native resolutions up to 4K. Post‑training, specialized experts for visual aesthetics, bilingual text rendering, infographic generation, and image editing are optimized and distilled into a multi‑expert framework, yielding advances in image fidelity, complex composition, multi‑reference editing, and instruction following.
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a relatively small visual generator can be pushed through systematic training engineering under a constrained computational budget.
arXiv:2609.21169v1 Announce Type: cross Abstract: Independent control of tonescale regions (e.g., shadows, highlights) is essential for painters, photographers and cinematographers to bring 2D images...
arXiv:2605.20941v2 Announce Type: replace Abstract: Existing neural painting methods are target-driven: given a reference image, strokes are optimized to reconstruct it, fixing the outcome before pai...