arXiv Computer Vision By Sandeep Wadhwa, Mayank Vatsa, Richa Singh, Parrva Chirag Shah, Prakhar Galriya

NumBench: Diagnosing Counting Failures in Text-to-Image Models

Read the original on arXiv Computer Vision →

NumBench is a large benchmark of 640,000 text‑to‑image prompts that tests how well models count objects, covering 1,600 categories and counts from 1 to 100. It uses a factorial design to vary composition, spatial guidance, and appearance while balancing counts, and introduces a process model that predicts a near‑quadratic collision deficit at low occupancy. The authors also propose the Confidence‑Weighted Numeric Precision Score for scalable evaluation and find that all tested systems perform poorly above 50 objects, with count range, layout, and composition having the largest effects.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 11

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

ABACUS is a 3‑billion‑parameter vision‑language model that simultaneously tackles object counting, crowd counting, referring‑expression counting, and count‑faithful image generation. It introduces density‑aware adaptive zooming with an objectness map, a boundary‑aware count policy trained via GRPO to avoid over‑ or under‑counting at crop edges, and a cycle‑consistent GRPO strategy that scores generated images for count accuracy and aesthetic quality without external critics. The model sets new state‑of‑the‑art performance on seven benchmarks, outperforming both specialized and larger generalist models.

By Anindya Mondal, Sauradip Nag, Anjan Dutta
arXiv Computer Vision
Aug 26

Object Counting Across Modalities: Taxonomies, Benchmarks, Applications, and Open Challenges

The paper reviews the evolution of object‑counting techniques from class‑specific density regression to open‑vocabulary, foundation‑model‑backed counters that can handle visual and textual prompts. It highlights that current evaluation relies on a few saturated benchmarks, leading to models exploiting statistical regularities rather than true generalization. The authors propose a five‑axis taxonomy and audit the literature across domains such as microscopy, remote sensing, crowd counting, and agriculture, identifying six structural contradictions and outlining a roadmap for robust, multimodal evaluation protocols.

By Joana Konadu Owusu, Shivanand Venkanna Sheshappanavar