Basketball Shot Counting with YOLOv12 and SAM3
2026 · Personal computer vision experiment
A two-stage video analysis pipeline that uses a compact YOLOv12 detector to prompt SAM3, then converts stable segmentation masks into made- and missed-shot counts. The project demonstrates the practical accuracy and compute tradeoff between lightweight detection and GPU-assisted foundation-model tracking.
System design
I labeled a small, purpose-built dataset for three classes—ball, hoop, and player—and split its 50 images into 36 training, 8 validation, and 6 test images. A YOLOv12n detector trained on those frames supplies bounding-box prompts for SAM3 rather than carrying the full tracking task by itself.
SAM3 propagates pixel-level masks for the prompted objects across the video. Downstream geometric and temporal logic uses the ball and hoop masks to identify shot attempts and classify makes and misses. The clip below isolates this tracking stage and shows the quality of the masks that drive the counter.
Lightweight inference versus stable tracking
The YOLOv12n checkpoint is approximately 5.5 MB and is the practical choice when latency, memory, or edge deployment matters most. In the YOLO-only shot counter below, however, intermittent and lower-quality ball detections produce noisier trajectories and less reliable event logic.
Using YOLO only for initialization and SAM3 for mask propagation produces much cleaner object geometry and more accurate counts, but the heavier model has a substantially larger compute requirement. I ran this version on an NVIDIA RTX 5070; that improves output quality, but makes it unsuitable for many low-power or CPU-only systems. The pair of demos makes that engineering tradeoff visible rather than treating model accuracy as the only design criterion.
YOLOv12 object-detection baseline
This shorter clip shows the detector and tracker without the shot-counting layer. It provides a direct view of the box detections used by the lightweight baseline and as prompts for the SAM3 pipeline.
YOLOv12n training diagnostics
Training stopped after 113 epochs. On the eight-image validation split, the run reached a peak mAP@50 of 0.789 and mAP@50–95 of 0.652. The normalized confusion matrix also shows the central limitation: hoop and player recall reached 0.80, while ball recall was 0.50. These results should be interpreted in the context of the very small dataset, but they explain why a lightweight detector alone was brittle on the moving ball.