arXiv AI By Siyi Xie, Xuanke Shi, Jinsheng Quan, Haoran Tang, Zukai Chen, Lei Yang, Quan Wang

TransPhy: Visual In-Context Learning for Physically Grounded Image Editing

Read the original on arXiv AI →

TransPhy is a framework for visually in-context learning that focuses on physically grounded image editing. It introduces PhysVICL-74, a dataset of 74 transformation rules and 5,240 source–target pairs, and evaluates models on novel-instance transfer and unseen-rule generalization. The method predicts the demonstrated rule and a query-specific target-state description, then synthesizes the target image using token-wise mixture-of-experts guided by localized transition cues, improving rule adherence, query consistency, and generalization over existing methods.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 17

VicEdit: Learning to Edit Videos from Visual In-Context Examples

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair.

arXiv Computer Vision
2d ago

Reference-free Human-Object Interaction Editing

arXiv:2503.09130v2 Announce Type: replace-cross Abstract: This paper presents InteractEdit, a novel framework for reference-free Human-Object Interaction (HOI) editing that tackles the challenging ta...

By Jiun Tian Hoe, Weipeng Hu, Wei Zhou, Chao Xie, Ziwei Wang, Xudong Jiang, Yap-Peng Tan, Chee Seng Chan