Research Paper:
GSP-YOLO: An Efficient Tire Detection Algorithm for Truck Scale Scenarios
Qiaole Zhu, Quan Liang, Nanlin Kuang, and Jinpeng Zhang
School of Computer Science and Mathematics, Fujian University of Technology
No.69 Xuefu South Road, Shangjie Town, Minhou County, Fuzhou, Fujian 350118, China
Corresponding author
Truck scale systems are often located in complex industrial environments, and existing inspection methods based on the appearance of an entire vehicle are difficult to accurately locate trucks in space-constrained scenarios. This study proposes GSP-YOLO, a tire detection algorithm for truck scale scenarios based on an improved YOLOv8n, that assists in locating trucks on the scale by detecting tires and enhances the detection performance in industrial environments. GSP-YOLO incorporates a global-to-local spatial aggregation module into the neck structure to improve the perception of tires on small scales. A shared detail-enhanced convolutional detection head is designed in the detection head to enhance its ability to recognize complex features while reducing the computational cost. The model also replaces conventional convolution with poly-scale convolution to enhance object recognition capability and reduce feature loss. Additionally, the wise-intersection over union loss function is employed as the bounding box regression loss to suppress competition among high-quality anchors and reduce the impact of low-quality samples, thereby improving detection performance. The experimental results demonstrated that GSP-YOLO achieves an mAP50 of 89.3% on the dataset, representing a 2.0% improvement over the baseline model, and an increase of 0.9% in mAP50–95. This model significantly enhances the detection capability within truck scale systems and ensures reliable performance in complex industrial environments.
- [1] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” 2014 IEEE Conf. on Computer Vision and Pattern Recognition, pp. 580-587, 2014. https://doi.org/10.1109/CVPR.2014.81
- [2] R. Girshick, “Fast R-CNN,” 2015 IEEE Int. Conf. on Computer Vision, pp. 1440-1448, 2015. https://doi.org/10.1109/ICCV.2015.169
- [3] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.39, No.6, pp. 1137-1149, 2017. https://doi.org/10.1109/TPAMI.2016.2577031
- [4] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” 2017 IEEE Int. Conf. on Computer Vision, pp. 2980-2988, 2017. https://doi.org/10.1109/ICCV.2017.322
- [5] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” 2017 IEEE Int. Conf. on Computer Vision, pp. 2999-3007, 2017. https://doi.org/10.1109/ICCV.2017.324
- [6] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” 2016 IEEE Conf. on Computer Vision and Pattern Recognition, pp. 779-788, 2016. https://doi.org/10.1109/CVPR.2016.91
- [7] J. Redmon and A. Farhadi, “YOLO9000: Better, faster, stronger,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition, pp. 6517-6525, 2017. https://doi.org/10.1109/CVPR.2017.690
- [8] J. Redmon and A. Farhadi, “YOLOv3: An incremental improvement,” arXiv:1804.02767, 2018. https://doi.org/10.48550/arXiv.1804.02767
- [9] A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal speed and accuracy of object detection,” arXiv:2004.10934, 2020. https://doi.org/10.48550/arXiv.2004.10934
- [10] C. Li et al., “YOLOv6: A single-stage object detection framework for industrial applications,” arXiv:2209.02976, 2022. https://doi.org/10.48550/arXiv.2209.02976
- [11] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,” 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition, pp. 7464-7475, 2023. https://doi.org/10.1109/CVPR52729.2023.00721
- [12] W. Liu et al., “SSD: Single shot multibox detector,” Proc. of the 14th European Conf. on Computer Vision, Part 1, pp. 21-37, 2016. https://doi.org/10.1007/978-3-319-46448-0_2
- [13] F. Tang et al., “DuAT: Dual-aggregation transformer network for medical image segmentation,” Proc. of the 6th Chinese Conf. on Pattern Recognition and Computer Vision, Part 5, pp. 343-356, 2023. https://doi.org/10.1007/978-981-99-8469-5_27
- [14] D. Li, A. Yao, and Q. Chen, “PSConv: Squeezing feature pyramid into one compact poly-scale convolutional layer,” Proc. of the 16th European Conf. on Computer Vision, Part 21, pp. 615-632, 2020. https://doi.org/10.1007/978-3-030-58589-1_37
- [15] Z. Tong, Y. Chen, Z. Xu, and R. Yu, “Wise-IoU: Bounding box regression loss with dynamic focusing mechanism,” arXiv:2301.10051, 2023. https://doi.org/10.48550/arXiv.2301.10051
- [16] K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.37, No.9, pp. 1904-1916, 2015. https://doi.org/10.1109/TPAMI.2015.2389824
- [17] T.-Y. Lin et al., “Feature pyramid networks for object detection,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition, pp. 936-944, 2017. https://doi.org/10.1109/CVPR.2017.106
- [18] S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition, pp. 8759-8768, 2018. https://doi.org/10.1109/CVPR.2018.00913
- [19] X.-Y. Li, Z.-X. Wei, and Y.-L. Yang, “Research on the application of HOG and SVM for wheel recognition in weigh-in-motion,” J. of Highway and Transportation Research and Development (English Edition), Vol.15, No.3, pp. 102-110, 2021. https://doi.org/10.1061/JHTRCQ.0000793
- [20] E. J. OBrien, C. C. Caprani, S. Blacoe, D. Guo, and A. Malekjafarian, “Detection of vehicle wheels from images using a pseudo-wavelet filter for analysis of congested traffic,” IET Image Processing, Vol.12, No.12, pp. 2222-2228, 2018. https://doi.org/10.1049/iet-ipr.2018.5369
- [21] Y.-S. Ruan, I.-C. Chang, and H.-Y. Yeh, “Vehicle detection based on wheel part detection,” 2017 IEEE Int. Conf. on Consumer Electronics – Taiwan, pp. 187-188, 2017. https://doi.org/10.1109/ICCE-China.2017.7991058
- [22] S. Ghanem and R. A. Kerekes, “Robust wheel detection for vehicle re-identification,” Sensors, Vol.23, No.1, Article No.393, 2023. https://doi.org/10.3390/s23010393
- [23] M. Shenoda, “Lighting and rotation invariant real-time vehicle wheel detector based on YOLOv5,” arXiv:2305.17785, 2023. https://doi.org/10.48550/arXiv.2305.17785
- [24] Z. Chen, Z. He, and Z.-M. Lu, “DEA-Net: Single image dehazing based on detail-enhanced convolution and content-guided attention,” IEEE Trans. on Image Processing, Vol.33, pp. 1002-1015, 2024. https://doi.org/10.1109/TIP.2024.3354108
- [25] Y. Wu and K. He, “Group normalization,” Proc. of the 15th European Conf. on Computer Vision, Part 13, pp. 3-19, 2018. https://doi.org/10.1007/978-3-030-01261-8_1
- [26] Z. Zheng et al., “Enhancing geometric factors in model learning and inference for object detection and instance segmentation,” IEEE Trans. on Cybernetics, Vol.52, No.8, pp. 8574-8586, 2022. https://doi.org/10.1109/TCYB.2021.3095305
- [27] X. Dai et al., “Dynamic head: Unifying object detection heads with attentions,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition, pp. 7369-7378, 2021. https://doi.org/10.1109/CVPR46437.2021.00729
- [28] Z. Yu et al., “YOLO-FaceV2: A scale and occlusion aware face detector,” Pattern Recognition, Vol.155, Article No.110714, 2024. https://doi.org/10.1016/j.patcog.2024.110714
- [29] Z. Zheng et al., “Distance-IoU loss: Faster and better learning for bounding box regression,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.34, No.7, pp. 12993-13000, 2020. https://doi.org/10.1609/aaai.v34i07.6999
- [30] Z. Gevorgyan, “SIoU loss: More powerful learning for bounding box regression,” arXiv:2205.12740, 2022. https://doi.org/10.48550/arXiv.2205.12740
- [31] H. Rezatofighi et al., “Generalized intersection over union: A metric and a loss for bounding box regression,” 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition, pp. 658-666, 2019. https://doi.org/10.1109/CVPR.2019.00075
- [32] R. L. Draelos and L. Carin, “Use HiResCAM instead of Grad-CAM for faithful explanations of convolutional neural networks,” arXiv:2011.08891, 2020. https://doi.org/10.48550/arXiv.2011.08891
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.