arXiv Computer Vision

The MODA General Attribute Suite: A Four-Track Evaluation Benchmark for Fashion Attribute Extraction

arXiv Computer Vision
Sep 3

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

MMTryon is a multi‑modal, multi‑reference virtual try‑on framework that generates high‑quality compositional try‑on results using text instructions and multiple garment images. It addresses three overlooked problems: supporting multiple try‑on items, allowing dressing style specification via text, and eliminating reliance on segmentation models by using a parsing‑free garment encoder and a scalable data generation pipeline. Experiments on high‑resolution benchmarks and in‑the‑wild test sets show MMTryon outperforms state‑of‑the‑art methods qualitatively and quantitatively.

By Xujie Zhang, Ente Lin, Michael Kampffmeyer, Zhenyu Xie, Jiang Li, Ting Liu, Xiaochao Qu, Luoqi Liu, Xiaodan Liang
arXiv Machine Learning
Jun 2

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

arXiv:2601. 22276v2 Announce Type: replace Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sustainable data marketplaces.

By Mingyu Lu, Soham Gadgil, Chris Lin, Chanwoo Kim, Su-In Lee
arXiv AI
Sep 4

SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking

SHELF is a Python system that creates synthetic, controlled benchmark data for evaluating large language models on bibliographic tasks such as classification, clustering, retrieval, pair classification, and instruction retrieval. It generates 62,899 model-written documents based on Library of Congress vocabularies and compares methods like TF, TF‑IDF, BM25, popular encoders, and zero‑shot decoders, reporting performance metrics such as 0.8887 for subject classification and 0.2605 for genre‑form classification. The tool also allows independent variation of bibliographic facets and can produce unseen documents beyond a model’s training cutoff, with results indicating that model rankings transfer more reliably than absolute scores when compared to other benchmarks.

By Michael J. Bommarito II
arXiv Computation and Language
Aug 27

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

The paper introduces a scalable product‑linking system that uses a retrieve‑then‑match cascade. First, a lightweight text cross‑encoder auto‑resolves the majority of merchant‑catalog product pairs with high precision, while an agentic multimodal vision‑language model handles the remaining ambiguous cases by inspecting images and performing web searches. This approach balances computational cost and accuracy, improving overall link coverage from 68% to 77% without requiring fine‑tuning of the agent.

By Jian Wang, Steven Xu, Sanjyot Thete, Maryam Barouti, Tom Tang, Elaine Wu, Charu Sareen, Kyle MacDonald