Transformer-based polygon detection for real-time parking analytics
A high-precision, transformer-based vision system for automatic detection of parking slot polygons and real-time slot occupancy classification using a camera-only setup.
This project proposes a high-precision, transformer-based vision system for automatic detection of parking slot polygons and real-time slot occupancy classification using a camera-only setup. Unlike traditional parking detection models that rely on bounding boxes or region heuristics, our framework models each parking slot as a multi-vertex polygon, allowing far more accurate geometric understanding. The core of the system is a SwinTransformer-based Mask2Former segmentation model, which provides robust polygonal masks under lighting variations, occlusions, tilted camera angles, and heterogeneous vehicle shapes. A second-stage classifier termed Dynamic Gap Analysis performs occupancy estimation by analyzing the geometric and appearance features inside each predicted slot. The system achieves production-grade accuracy, supports real-time inference, and is designed for deployment in smart parking lots, residential complexes, and urban mobility systems.
PyTorch, Detectron2, Swin Transformer, Mask2Former, OpenCV, ONNX Runtime, CVAT (polygon annotation), Python, FastAPI, Redis Queue, Docker, TensorRT, NVIDIA GPU