arXiv:2608.22329v1 Announce Type: cross
Abstract: Emotion-aware artistic image generation requires a model to satisfy semantic content, artistic style, and target emotion simultaneously. The key chal...
By Qianqian Tang, Jiayi Gao, Ting Lei, Yang Liu
arXiv:2606. 09846v1 Announce Type: cross Abstract: Visual art remains largely inaccessible to blind and low-vision (BLV) audiences due to brief or absent alt-text, which rarely conveys the sensory, spatial, or emotional qualities of an artwork.
By Vignesh Nagarajan
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.
By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le
Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions.
Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice, multi-shot composition is critical for downstream creative workflows: commercial posters often require multiple crops with different emphases (e.
The paper introduces Chat-Edit-3D++ (CE3D++), an interactive 3D and 4D scene editing system that uses a Hash-Atlas network to separate 2D editing from 3D reconstruction. CE3D++ employs a large language model to accept arbitrary textual input, interpret user intent, and autonomously invoke appropriate visual models, enabling multi‑round dialogue and diverse editing effects. The approach is extended to monocular 4D scenes by adding motion constraints and a trajectory dataset, allowing a smaller LLM to schedule up to 30 visual tools accurately.
By Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai, Wenrui Ding, Yi Yang, Shuchang Zhou, Ming-Hsuan Yang
arXiv:2502. 06819v2 Announce Type: replace Abstract: This paper presents a framework for generating 3D indoor scenes from text prompts.
By Yao Wei, Matteo Toso, Pietro Morerio, Changjae Oh, Michael Ying Yang, Alessio Del Bue
The paper introduces LEGO, a benchmark dataset pairing user text descriptions with human‑annotated fine‑grained constraints and reference 3D scenes, and LEGO‑Eval, an evaluation framework that decomposes descriptions into atomic constraints and verifies each using grounding and spatial reasoning tools. It demonstrates that LEGO‑Eval detects misalignment more accurately than existing methods and that current 3D scene synthesis approaches achieve at most a 10% success rate on this benchmark.
By Minseok Kang, Dongwook Choi, Gyeom Hwangbo, Seungwon Lim, Kai Tzu-iunn Ong, Jinyoung Yeo
This paper presents an overview of the inaugural PortraitCraft Challenge, held as one of the official competitions at CVPR 2026. The challenge focuses on portrait composition understanding and generation, aiming to advance AI research in portrait aesthetics analysis and controllable image synthesis.
Instruction-driven editing of 3D human motion requires precise spatiotemporal localization, rich semantic grounding, and strict preservation of unmodified content. Existing methods either resort to training-free adaptation of generative models or rely solely on triplet supervision; however, adaptation often yields suboptimal control, and manually curated triplet datasets remain severely limited in scale and semantic diversity.
arXiv:2603.11421v2 Announce Type: replace
Abstract: Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Impl...
By Songlin Yang, Zhe Wang, Xuyi Yang, Songchun Zhang, Xianghao Kong, Taiyi Wu, Xiaotong Zhao, Ran Zhang, Alan Zhao, Anyi Rao
SceneReGen is a new framework for reconstructing 3D scenes from a single image by generating and assembling complete object meshes within a shared observation‑aligned scene frame. It uses selective pose factorization to encode each object’s observed orientation directly into the generated mesh, while estimating translation and scale from instance‑level and global scene cues. Evaluated on the 3D‑FUTURE dataset, SceneReGen outperforms existing methods on scene‑level metrics and shows strong performance on object‑level metrics, demonstrating its effectiveness in autonomous‑driving and embodied‑AI scenarios.
By Zefan Tian, Yuteng Ye, Yiheng Zhang, Yuhang Yang, Xueqiang Lv, Shizhou Zhang, Le Liu, Di Xu