arXiv Computer Vision

CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance

arXiv Computer Vision
Sep 11

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.

By Anindya Mondal, Sauradip Nag, Anjan Dutta
Hugging Face Trending Papers
Jun 22

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generation without any benchmark-specific training required. Our model is built on existing 3B-parameter unified foundation model and is adapted for object localization tasks using three key innovations: density-aware adaptive zooming with objectness maps for spatial grounding; a boundary-aware count policy via GRPO to eliminate crop-boundary errors; and a cycle-consistent GRPO strategy where the understanding branch self-critiques generated outputs, closing the understanding-generation gap without any external annotations.

arXiv AI
Jul 21

Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

arXiv:2607. 16409v1 Announce Type: cross Abstract: Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they still struggle to follow complex spatial instructions and logical constraints in controllable image generation.

By Junhao Liu, Jian-Wei Zhang, Tao Huang, Miles Yang, Zhao Zhong, Liefeng Bo
arXiv Computer Vision
Sep 17

DualCount: Structurally Consistent Density and Point Modeling for Zero-Shot Object Counting

DualCount introduces an instance-aware dual-decoder framework that couples density and point representations for zero‑shot object counting. By treating density estimation as a structured mass allocation over latent object instances, it applies two geometric constraints—per‑instance mass conservation and center‑of‑mass alignment—to enforce instance‑level consistency. Experiments on FSC‑147, PUCPR+, and CARPK demonstrate that this approach consistently reduces counting error and achieves new state‑of‑the‑art performance.

By Xuan Cuong Ngo
arXiv Computer Vision
Aug 24

Exploring the Performance Frontier of Compact Unified Image Generation Models

Swift-Image is a compact unified model that performs text-to-image generation, single-image editing, and multi-image editing using a 6B parameter DiT architecture. It employs a progressive training pipeline, parallel expert reinforcement learning, and multi-teacher distillation to balance diverse objectives, while a Prompt Enhancer decouples high-level reasoning from pixel-level rendering. After training, structural pruning and few-step distillation produce efficient 3B and accelerated variants that maintain near‑lossless performance and improve editing efficiency.

By Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu, Hao Yan, Zhengze Xu, Yuhang Yu, Yongchao Du, Xingjian Wang, Jun Zheng, Qinye Zhou, Yaqi Cai, Zhengrui Chen, Chao Lin, Yefeng Shen, Yuan Wang, Zhengtao Wu, Ge Wu, Xiaoli Xu, Denghui Yang, Huayu Zhang, Mingzhou Zhang, Mengting Chen
arXiv AI
Aug 7

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

arXiv:2608. 06161v1 Announce Type: new Abstract: Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often optimize perceptual realism without reliably satisfying task-critical functional constraints.

By Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel
arXiv AI
Aug 7

Depth-Guided Video Object Counting in Crowded Scenes

arXiv:2608. 06236v1 Announce Type: cross Abstract: Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category based on given text or visual prompts.

By Yuanjing Xu, Xinyan Liu, Weidong Chen, Zixuan Zou, Linhao Zhang, Zhuangzhe Meng, Antoni B. Chan, Weigang Zhang
arXiv Computer Vision
Aug 21

Swift-Image: Exploring the Performance Frontier of Compact Unified Image Generation Models

arXiv:2608. 20334v1 Announce Type: new Abstract: We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing.

By Taihang Hu, Zhao Wang, Zuan Gao, Tao Liu, Hao Yan, Zhengze Xu, Yuhang Yu, Yongchao Du, Xingjian Wang, Jun Zheng, Qinye Zhou, Zhengrui Chen, Chao Lin, Yefeng Shen, Zhengtao Wu, Ge Wu, Xiaoli Xu, Denghui Yang, Huayu Zhang, Mingzhou Zhang, Mengting Chen