Count Anything AI Model Counts Objects Across Six Domains

This new open-source model uses a dual-counter architecture to handle diverse images, from satellite views to histopathology slides.

By Central
Researchers at Tsinghua University developed Count Anything, a unified model for counting across six visual domains.
Highlights
  • Count Anything uses two counters: a region-level sparse counter for large objects and a pixel-level dense counter for small packed objects.
  • The model achieves roughly nine errors per queried category per image, outperforming alternatives like CountGD and CLIP-Count by more than double.
  • The CLOC dataset, with 220,000 images across 619 categories, supports training and evaluation for text-guided counting.

Large language models can describe images, interpret charts, and extract text from photos, making multimodality a standard expectation for modern AI systems. Yet one deceptively simple task remains persistently difficult: reliably counting objects in an image. Getting those counts right carries real-world weight, whether a clinician reads a scan, a farmer estimates crop yields, or a city planner analyzes traffic patterns. Until now, each of these counting tasks has required its own specialized system. Count Anything, a new AI model developed by researchers at Tsinghua University and collaborating institutions, aims to change that by counting objects across six fundamentally different visual domains.

What the Count Anything Model Does

Count Anything is a text-guided object counting system designed to handle images as diverse as crowded streets, satellite views of vehicles, histopathology slides, microscopic cell cultures, agricultural fields, and bacterial colony plates. A user provides a text prompt specifying the object type to count, and the model returns a point set marking every detected instance in the image. The goal is a single unified architecture that works reliably where previous systems required domain-specific tuning.

Two Counters Working in Tandem

The core innovation is a dual-counter architecture. One component, a region-level sparse counter, specializes in large, clearly visible objects and outputs bounding boxes around each detection. The other, a pixel-level dense counter, handles small, tightly packed objects by placing a single dot on each target. Both predictions are merged through a complementary fusion process. A simple deduplication rule prevents double-counting: when both counters flag the same object, only the prediction with higher confidence is retained.

The architecture builds on Meta’s SAM3, a pretrained model capable of joint image and text understanding. Rather than retraining the entire backbone, the researchers attach lightweight adapter modules tailored to the counting task, preserving the base model’s general visual knowledge while adding domain-agnostic counting capability.

The CLOC Dataset: Six Domains in One Collection

Training a model to generalize across such varied imagery required a matching dataset. The researchers assembled and cleaned existing public datasets, reconciling conflicting labels, and released the result as CLOC, described as the largest dataset for text-guided counting to date. CLOC contains approximately 220,000 images spanning 619 object categories and 15 million labeled instances across six domains: general scene photography, remote sensing and drone imagery, histopathology, cellular microscopy, agriculture (including wheat ears), and microbiology.

Benchmark Performance and Comparisons

In the team’s reported evaluations, Count Anything substantially outperforms existing systems including CountGD, CLIP-Count, and Grounding DINO. On average, the model miscounts by roughly nine objects per queried category per image. The best competing model misses by more than twice that margin. In pure crowd-counting tasks, Count Anything remains competitive but does not surpass the most specialized dedicated systems.

The researchers acknowledge remaining limitations. Ambiguous or highly specialized terminology can cause the model to miss objects or misclassify them. In extremely dense scenes with heavy occlusion, distinguishing whether two predictions refer to the same object or two distinct ones becomes difficult.

Context: AI’s Persistent Blind Spot in Basic Visual Tasks

The difficulty of fundamental visual tasks for AI was recently highlighted by the BabyVision benchmark, where most frontier models scored below the average three-year-old. Even top performers like Gemini 3 Pro barely reached 50 percent accuracy, while adults scored above 94 percent. The gap was most pronounced in counting occluded 3D blocks, where the best model managed only 20.5 percent while humans solved the task without a single error.

Who Should Try This Now

Count Anything is available as open-source code on GitHub, making it immediately accessible to researchers and developers working on any of the six supported domains. Practitioners in medical imaging, precision agriculture, remote sensing, and microbiology can test the model against their own datasets. For teams currently maintaining separate counting pipelines for different image types, unifying around a single cross-domain model represents a practical opportunity to reduce engineering overhead while maintaining competitive accuracy. Visit the Count Anything repository to explore the model and the CLOC dataset directly.

Share This Article