YOLOv5-based metro gate vaulting detection system — end-to-end from model training to RK3588 edge deployment, enabling real-time identification and alerting of violations in metro scenarios.
Metro gate vaulting (or crawling under) is one of the most common violations in urban rail transit, causing both fare losses and serious safety hazards. Traditional approaches rely on manual monitoring, which is inefficient and prone to missed detections. This project uses a YOLOv5 object detection model trained specifically to identify gate vaulting behavior, deployed on a Rockchip RK3588 edge device for end-to-end real-time inference from video stream input to violation alerts.
Visual detection in metro gate scenarios faces multiple challenges:
Customized modifications based on the YOLOv5s architecture. First, we lightweight the Backbone by replacing standard convolutions in C3 modules with Depthwise Separable Convolutions, reducing parameters from 7.2M to 4.1M while maintaining feature extraction capability.
For the dataset, we collected 8,000+ frames from metro station surveillance videos, annotating three target classes: normal_pass, climb_over, and crawl_under. We used Mosaic + MixUp + random cropping augmentation, and introduced optical flow features from adjacent frames as auxiliary input to capture the temporal characteristics of vaulting actions.
Rockchip RK3588 is one of the most powerful domestic edge AI SoCs, integrating quad-core Cortex-A76 + quad-core Cortex-A55 CPU, Mali-G610 GPU, and a dedicated 6 TOPS NPU (Neural Processing Unit). Its NPU supports INT8/INT16/FP16 mixed-precision inference, making it ideal for deploying quantized YOLO models.
Our development board features 8GB LPDDR5 memory and a dedicated NPU accelerator. Compared to GPU solutions (like Jetson Orin Nano), RK3588 offers lower power consumption (typically 5-8W vs 15W) and better cost-effectiveness, making it ideal for large-scale metro station deployments. The 6 TOPS NPU is more than sufficient for YOLOv5s-level models, achieving 60+ FPS in real-world testing.
Deploying from PyTorch model to RK3588 NPU requires the following key steps:
rknn-toolkit2 to convert ONNX model to RKNN format with INT8 quantization (500-image calibration set)rknn-lite2 runtime with NPU-accelerated inferenceWhen deploying on RK3588, we performed multi-dimensional inference optimization:
core_mask=RKNN.NPU_CORE_0_1_2 for maximum throughputThe system achieves 62 FPS inference speed on RK3588 with end-to-end latency (including pre/post-processing) under 20ms, fully meeting real-time metro monitoring requirements. The INT8 quantized model is only 3.8MB, running stably at 6W power consumption, suitable for large-scale deployment across station gates.