Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots from one scene. In practice, multi-shot composition is critical for downstream creative workflows: commercial posters often require multiple crops with different emphases (e.
The paper introduces a new approach to aesthetic image cropping by modeling human preference as a continuous, multi-peaked field rather than relying on discrete, grid‑based annotations. It presents the Continuous Preference Field (CPF) that reconstructs a dense preference landscape from sparse labels, and uses this to train a VLM‑based cropping model (CPIC) that achieves state‑of‑the‑art accuracy and strong out‑of‑domain generalization. Additionally, the authors propose CPICD, a recalibrated benchmark that corrects grid‑bound artifacts in existing datasets, providing a more reliable evaluation framework.
By Ziqing Zhang, Xiao Liu, Kai Liu, Jianze Li, Weihang Zhang, Linghe Kong, Yulun Zhang
CS-CLIP is a vision‑language model that improves compositional reasoning by using scene graphs to identify compositional elements and create structured negative examples through selective masking. The approach retains only the most contradictory negatives, encouraging the model to depend on compositional structure instead of surface cues. CS-CLIP achieves state‑of‑the‑art performance on compositional reasoning benchmarks while maintaining strong cross‑modal retrieval and downstream visual reasoning capabilities with fewer training samples.
By SeongJun Jeong, Minjoon Jung, Woo Suk Choi, Youwon Jang, Byoung-Tak Zhang
arXiv:2609.14657v1 Announce Type: cross
Abstract: While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental c...
By Sehwan Park, Taehoon Kim, Geonhee Han, Dohyun Kim, Seung Wook Kim, Paul Hongsuck Seo
arXiv:2405. 03650v4 Announce Type: replace-cross Abstract: We study Generated Contents Enrichment (GCE), a conditional image-generation task in which a sparse scene description is first enriched through an explicit scene representation and then rendered into semantically richer visual content.
By Mahdi Naseri, Jiayan Qiu, Zhou Wang
PANORAMA introduces a new panoptic grounded captioning framework that jointly generates detailed image captions and associates each phrase with precise pixel-level masks. The authors create PanoCaps, a human‑annotated benchmark with dense captions and near‑complete pixel coverage, and propose a phrase‑mask matching protocol with a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations, learns to select appropriate masks, and achieves state‑of‑the‑art grounding performance on PanoCaps and other pixel‑level tasks.
By Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid