arXiv AI

Personalized Image Generation with Reasoning and Reflection

arXiv AI
Aug 24

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

PhotoBench is a new benchmark built from authentic personal photo albums that moves beyond simple visual matching to focus on personalized, intent-driven retrieval. It incorporates a multi-source profiling framework that combines visual semantics, spatial‑temporal metadata, social identity, and temporal events to generate complex queries reflecting users’ life trajectories. Evaluation on PhotoBench reveals two key limitations: a modality gap where unified embedding models fail on non‑visual constraints, and a source fusion paradox where agentic systems struggle with tool orchestration.

By Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
arXiv AI
Jun 19

VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions

arXiv:2606. 19627v1 Announce Type: cross Abstract: The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds.

By Katya Mirylenka, Egor Malykh, Mahdyar Ravanbakhsh, Michael Gygli, Marco-Andrea Buchmann, Andrew Dzhoha, Svitlana Borzenko, Francesca Catino, Mohamed Gaafar, Maarten Versteegh, Thomas Kober, Dario d'Andrea, Ellie Langhans
arXiv Computation and Language
Sep 18

To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

The paper introduces ReaLMem, a benchmark built from authentic multi‑year personal visual archives with first‑person annotations, designed to evaluate AI systems on factual recall, persona inference, and predictive personalization. It also proposes ChronoProfiler, a temporal‑weighting module that calculates stability scores for user attributes to resolve preference conflicts and enhance personalized decision making. Experiments with multimodal large language models and memory systems show that predictive personalization remains the hardest task, highlight performance gaps, and demonstrate that temporally informed representations significantly improve personalization.

By Wenqi Zhou, Zhuorui Yu, Kaiao Wen, Hao Zheng, Xinyi Zheng, Peiran Wu, Enmin Zhou, Chi-Hao Wu, Junxiao Shen
arXiv Computer Vision
Sep 7

SEAL: Semantic-aware Single-image Sticker Personalization with a Large-scale Sticker-tag Dataset

The paper introduces SEAL, a plug‑and‑play module that enhances single‑image sticker personalization in diffusion models by adding a semantic‑guided spatial attention loss, a split‑merge token strategy, and structure‑aware layer restriction. SEAL integrates without altering the U‑Net backbone and addresses overfitting issues such as visual entanglement and structural rigidity. Alongside SEAL, the authors release StickerBench, a large sticker dataset with six attribute tags to enable systematic evaluation of identity preservation and contextual controllability.

By Changhyun Roh, Yonghyun Jeong, Jonghyun Lee, Chanho Eom, Jihyong Oh
arXiv Computer Vision
Aug 27

VGA-BenchV2: An Expanded Unified Benchmark and Multi-Model Framework for Evaluating Video Aesthetics and Generation Quality

VGA‑BenchV2 is an expanded, human‑aligned benchmark and optimization framework that jointly evaluates video generation quality and aesthetic value. It builds on the original VGA‑Bench taxonomy, adding 52 sub‑dimensions and 1,016 curated prompts to generate over 60,000 videos from 12 mainstream models. The benchmark significantly enlarges human supervision with 36,000 task‑level annotations and introduces a hybrid evaluator (VAQA‑Net, VTag‑Net, VGQA‑Net) that aligns well with human judgments and can be used as a reward model for reinforcement‑learning fine‑tuning.

By Longteng Jiang, DanDan Zheng, Qianqian Qiao, Heng Huang, Huaye Wang, Yihang Bo, Bao Peng, Jingdong Chen, Jun Zhou, Xin Jin