arXiv AI By Wenjie Zheng, Haoji Hu, Jiali Lu, Xingze Zou, Jing Wang

D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation

Read the original on arXiv AI →

The paper introduces D3S2, a diffusion‑guided dataset distillation framework tailored for semantic segmentation. It tackles long‑tailed class imbalance, pixel‑wise alignment, and high computational cost by first selecting class‑balanced masks and then synthesizing images with a pretrained diffusion model conditioned on those masks. Guided diffusion sampling further refines the data with segmentation‑consistency and class‑wise feature matching losses, achieving superior performance at a 1% compression rate on ADE20K and COCO‑Stuff.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 25

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D semantic segmentation dataset featuring 20,000 image pairs and 160 fine‑grained categories, surpassing existing benchmarks like NYUv2 and SUN RGB‑D in both scale and semantic diversity. The dataset provides high‑fidelity annotations obtained through rigorous re‑evaluation and correction of prior labels, ensuring a clean ground‑truth foundation. Additionally, the authors introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art performance across evaluated benchmarks, demonstrating the value of high‑quality multimodal information.

By Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan
arXiv Machine Learning
Aug 28

A Framework for Low-Effort Training Data Generation for Urban Semantic Segmentation

The paper introduces a framework that adapts a diffusion model to a target urban domain using only imperfect pseudo‑labels, enabling the generation of high‑fidelity, target‑aligned images from semantic maps of any synthetic dataset. By filtering poor generations, correcting image‑label misalignments, and standardising semantics, the method transforms low‑effort synthetic data into competitive real‑domain training sets. Experiments on five synthetic and two real datasets show up to +8.0 %pt mIoU improvement over state‑of‑the‑art translation methods, demonstrating that rapidly constructed synthetic datasets can match the performance of high‑effort, manually designed ones.

By Damjan Kal\v{s}an, Denis Zavadski, Tim K\"uchler, Haebom Lee, Stefan Roth, Carsten Rother
arXiv Computer Vision
Sep 4

PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models

PoseDreamer is a new pipeline that uses diffusion models to generate large‑scale synthetic datasets for 3D human mesh estimation, providing 3D mesh annotations that remain aligned with the generated images. The system incorporates controllable image generation, Direct Preference Optimization for control alignment, curriculum‑based hard sample mining, and multi‑stage quality filtering to produce over 500,000 high‑quality samples with a 76% improvement in image‑quality metrics over traditional rendering‑based datasets. Models trained on PoseDreamer match or surpass those trained on real‑world or conventional synthetic data, and combining PoseDreamer with synthetic datasets yields better performance than mixing real and synthetic data alone.

By Lorenza Prospero, Orest Kupyn, Ostap Viniavskyi, Jo\~ao F. Henriques, Christian Rupprecht