Research Paper:
YOLO-Pooling: Exploring the Potential of Pooling Operations in Object Detection
Xuebin Yue, Mengkui Hao, Yao Yao, Yangyang Wang, and Yan Wang
School of Automation and Electrical Engineering, Zhongyuan University of Technology
No.1 Huaihe Road, Longhu Town, Xinzheng, Zhengzhou, Henan 451191, China
Corresponding author
Pooling operations play a crucial role in object detection by reducing feature map dimensions, enhancing position invariance, enabling multi-scale detection capabilities, and lowering computation overhead, which is particularly important in industrial inspection scenarios such as metal surface defect detection, concrete crack detection, hot-rolled strip inspection, and PCB defect analysis where complex backgrounds and subtle defects demand high accuracy and robustness. However, existing architectures (e.g., FPN/PANet, U-Net with attention) rely on single-type pooling or simple multi-scale concatenation, leading to information loss or inefficient feature aggregation. To address these limitations, we aim to explore the potential of pooling operations in the context of object detection tasks. To this end, we propose two modules based on pooling operations. The first is the dimension reduction pooling (DRP) module, which differs from traditional single pooling by combining 3×3 max pooling and average pooling during dimension reduction (to preserve both salient features and global distribution) and compressing channels via 1×1 convolution (to avoid redundancy), providing a richer feature representation. The second is the multi-scale feature aggregation (MFA) module, which innovatively integrates CSP structure, multi-step weighted feature aggregation, and stereoscopic attention (parallel channel-spatial attention)—distinct from U-Net’s symmetric mapping and FPN/PANet’s simple pathway concatenation. It seamlessly integrates coarse semantic information with fine-grained semantic information through top-down and bottom-up pathways, employs a stereoscopic attention mechanism to enhance feature representation, expands perception scope, and improves generalization capabilities. Based on these two modules, the YOLO-Pooling is proposed, an object detection model that progressively refines deep and shallow semantic features. The proposed method is evaluated on six public datasets: GC10-DET, Crack, Barcodes, NEU-DET, PCB, and a subset of COCO, and the mAP of the method is 67.74%, 86.11%, 97.78%, 73.27%, 96.56%, and 6.49%, respectively, significantly higher than the state-of-the-art detection methods. Experimental results demonstrate that the proposed DRP + MFA design outperforms existing multi-scale aggregation and pooling-based architectures by solving the trade-off between pooling-induced information loss and model accuracy, fundamentally improving object localization and detection accuracy, and maintains efficient inference speed. This efficiency stems from precise FLOPs control, optimized memory access patterns, and operator fusion, enabling higher fps than many mainstream YOLO models despite additional modules (DRP and MFA).
YOLO-pooling overall framework
- [1] Y. Zhou, “A YOLO-NL Object Detector for Real-Time Detection,” Expert Systems with Applications, Vol.238, Article No.122256, 2024. https://doi.org/10.1016/j.eswa.2023.122256
- [2] Y. Cai, T. Luan, H. Gao, H. Wang, L. Chen, Y. Li, M. A. Sotelo, and Z. Li, “YOLOv4-5D: An Effective and Efficient Object Detector for Autonomous Driving,” IEEE Trans. on Instrumentation and Measurement, Vol.70, Article No.4503613, 2021. https://doi.org/10.1109/TIM.2021.3065438
- [3] Z. Wang, Y. Li, Y. Liu, and F. Meng, “Improved Object Detection via Large Kernel Attention,” Expert Systems with Applications, Vol.240, Article No.122507, 2024. https://doi.org/10.1016/j.eswa.2023.122507
- [4] R. Xu, X. Zhao, F. Liu, and B. Tao, “High-Precision Monocular Vision Guided Robotic Assembly Based on Local Pose Invariance,” IEEE Trans. on Instrumentation and Measurement, Vol.72, Article No.5031312, 2023. https://doi.org/10.1109/TIM.2023.3328699
- [5] X. Yue, H. Li, M. Shimizu, S. Kawamura, and L. Meng, “YOLO-GD: A Deep Learning-Based Object Detection Algorithm for Empty-Dish Recycling Robots,” Machines, Vol.10, No.5, Article No.294, 2022. https://doi.org/10.3390/machines10050294
- [6] W. Zhou, Y. Zhu, J. Lei, J. Wan, and L. Yu, “APNet: Adversarial Learning Assistance and Perceived Importance Fusion Network for All-Day RGB-T Salient Object Detection,” IEEE Trans. on Emerging Topics in Computational Intelligence, Vol.6, No.4, pp. 957-968, 2022. https://doi.org/10.1109/TETCI.2021.3118043
- [7] R. Cong, W. Song, J. Lei, G. Yue, Y. Zhao, and S. Kwong, “PSNet: Parallel Symmetric Network for Video Salient Object Detection,” IEEE Trans. on Emerging Topics in Computational Intelligence, Vol.7, No.2, pp. 402-414, 2023. https://doi.org/10.1109/TETCI.2022.3220250
- [8] Y. Ge, Z. Li, X. Yue, H. Li, Q. Li, and L. Meng, “IoT-Based Automatic Deep Learning Model Generation and the Application on Empty-Dish Recycling Robots,” Internet of Things, Vol.25, Article No.101047, 2024. https://doi.org/10.1016/j.iot.2023.101047
- [9] X. Yue, H. Li, Y. Fujikawa, and L. Meng, “Dynamic Dataset Augmentation for Deep Learning-Based Oracle Bone Inscriptions Recognition,” J. Comput. Cult. Herit., Vol.15, No.4, Article No.76, 2022. https://doi.org/10.1145/3532868
- [10] X. Yue, Z. Wang, R. Ishibashi, H. Kaneko, and L. Meng, “An Unsupervised Automatic Organization Method for Professor Shirakawa’s Hand-Notated Documents of Oracle Bone Inscriptions,” Int. J. on Document Analysis and Recognition (IJDAR), Vol.27, pp. 583-601, 2024.
- [11] J.-J. Liu, Q. Hou, Z.-A. Liu, and M.-M. Cheng, “PoolNet: Exploring the Potential of Pooling for Salient Object Detection,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, No.1, pp. 887-904, 2023. https://doi.org/10.1109/TPAMI.2021.3140168
- [12] S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path Aggregation Network for Instance Segmentation,” 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition, pp. 8759-8768, 2018. https://doi.org/10.1109/cvpr.2018.00913
- [13] R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation,” 2014 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 580-587, 2014. https://doi.org/10.1109/cvpr.2014.81
- [14] S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” arXiv preprint, arXiv:1506.01497, 2015. http://arxiv.org/abs/1506.01497
- [15] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 2999-3007, 2017. https://doi.org/10.1109/iccv.2017.324
- [16] J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” arXiv preprint, arXiv:1804.02767, 2018. http://arxiv.org/abs/1804.02767
- [17] A. Bochkovskiy, C.-Y. Wang, and H. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection,” arXiv preprint, arXiv:2004.10934, 2020. https://arxiv.org/abs/2004.10934
- [18] Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun, “YOLOX: Exceeding YOLO Series in 2021,” arXiv preprint, arXiv:2107.08430, 2021. https://arxiv.org/abs/2107.08430
- [19] C.-Y. Wang, A. Bochkovskiy, and H. M. Liao, “YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,” arXiv preprint, arXiv:2207.02696, 2022. https://doi.org/10.48550/arXiv.2207.02696
- [20] I. Robinson, P. Robicheaux, M. Popov, D. Ramanan, and N. Peri, “RF-DETR: Neural Architecture Search for Real-Time Detection Transformers,” arXiv preprint, arXiv:2511.09554, 2025. https://doi.org/10.48550/arXiv.2511.09554
- [21] Z. Liao, Y. Zhao, X. Shan, Y. Yan, C. Liu, L. Lu, X. Ji, and J. Chen, “RT-DETRv4: Painlessly Furthering Real-Time Object Detection With Vision Foundation Models,” arXiv preprint, arXiv:2510.25257, 2025. https://doi.org/10.48550/arXiv.2510.25257
- [22] L. Deng, Y. Tan, and S. Chen, “RAFNet: Rotation-Aware Anchor-Free Framework for Geospatial Object Detection,” Computer Vision and Image Understanding, Vol.257, Article No.104373, 2025. https://doi.org/10.1016/j.cviu.2025.104373
- [23] X. Yue, H. Li, and L. Meng, “An Ultralightweight Object Detection Network for Empty-Dish Recycling Robots,” IEEE Trans. on Instrumentation and Measurement, Vol.72, Article No.2505612, 2023. https://doi.org/10.1109/TIM.2023.3241078
- [24] W.-Y. Hsu and W.-Y. Lin, “Ratio-and-Scale-Aware YOLO for Pedestrian Detection,” IEEE Trans. on Image Processing, Vol.30, pp. 934-947, 2021. https://doi.org/10.1109/TIP.2020.3039574
- [25] Y. Hui, J. Wang, and B. Li, “WSA-YOLO: Weak-Supervised and Adaptive Object Detection in the Low-Light Environment for YOLOV7,” IEEE Trans. on Instrumentation and Measurement, Vol.73, Article No.2507012, 2024. https://doi.org/10.1109/TIM.2024.3350120
- [26] J. Tang, X. Hu, S. Jeon, and W. Chen, “Light-YOLO: A lightweight Detection Algorithm Based on Multi-Scale Feature Enhancement for Infrared Small Ship Target,” Complex & Intelligent Systems, Vol.11, Article No.130, 2025. https://doi.org/10.1007/s40747-024-01726-3
- [27] Q. Hou, L. Zhang, M.-M. Cheng, and J. Feng, “Strip Pooling: Rethinking Spatial Pooling for Scene Parsing,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 4002-4011, 2020. https://doi.org/10.1109/cvpr42600.2020.00406
- [28] Z. Gao, L. Wang, and G. Wu, “LIP: Local Importance-Based Pooling,” 2019 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 3354-3363, 2019. https://doi.org/10.1109/iccv.2019.00345
- [29] Y.-H. Wu, Y. Liu, X. Zhan, and M.-M. Cheng, “P2T: Pyramid Pooling Transformer for Scene Understanding,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, No.11, pp. 12760-12771, 2023. https://doi.org/10.1109/TPAMI.2022.3202765
- [30] T.-Y. Lin, P. Dollar, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature Pyramid Networks for Object Detection,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 936-944, 2017. https://doi.org/10.1109/cvpr.2017.106
- [31] C. Guo, B. Fan, Q. Zhang, S. Xiang, and C. Pan, “AugFPN: Improving Multi-Scale Feature Learning for Object Detection,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 12592-12601, 2020. https://doi.org/10.1109/cvpr42600.2020.01261
- [32] Z. Chen, H. Ji, Y. Zhang, Z. Zhu, and Y. Li, “High-Resolution Feature Pyramid Network for Small Object Detection on Drone View,” IEEE Trans. on Circuits and Systems for Video Technology, Vol.34, No.1, pp. 475-489, 2024. https://doi.org/10.1109/TCSVT.2023.3286896
- [33] G. Zhou, L. Yu, E. Gao, and Y. Lu, “DMEFF-Net: A Dynamic Multiscale Enhanced Feature Fusion Model for Small Object Detection in Remote Sensing Images,” IEEE J. of Selected Topics in Applied Earth Observations and Remote Sensing, Vol.19, pp. 1006-1022, 2025. https://doi.org/10.1109/JSTARS.2025.3634977
- [34] M. A. Jahin, S. Soudeep, M. F. Mridha, N. Fahad, and M. J. Hossen, “DyCAF-Net: Dynamic Class-Aware Fusion Network,” 2025 IEEE 12th Int. Conf. on Data Science and Advanced Analytics (DSAA), 2025. https://doi.org/10.1109/dsaa65442.2025.11247981
- [35] Z. Guo, Y. Wang, and N. Chen, “LAR-TSDETR: A Lightweight Adaptive Robust Traffic Sign Detection Transformer with Multi-Scale Fusion for Real-Time Recognition Under Challenging Conditions,” Measurement Science and Technology, Vol.36, No.9, Article No.096129, 2025. https://doi.org/10.1088/1361-6501/ae050a
- [36] J. Lian, Z. Wan, M. Gao, and J. Chen, “CFMD: Dynamic Cross-Layer Feature Fusion for Salient Object Detection,” Int. Conf. on Intelligent Computing, pp. 91-102, 2025. https://doi.org/10.1007/978-981-96-9901-8_8
- [37] S. Liu and J. Li, “EC-PFN: A Multiscale Woven Fusion Network for Industrial Product Surface Defect Detection,” Complex & Intelligent Systems, Vol.11, Article No.59, 2025. https://doi.org/10.1007/s40747-024-01699-3
- [38] C. Peng, X. Li, and Y. Wang, “TD-YOLOA: An Efficient YOLO Network With Attention Mechanism for Tire Defect Detection,” IEEE Trans. on Instrumentation and Measurement, Vol.72, Article No.3529111, 2023. https://doi.org/10.1109/TIM.2023.3312753
- [39] J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation Networks,” 2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7132-7141, 2018. https://doi.org/10.1109/cvpr.2018.00745
- [40] Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, “ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 11531-11539, 2020. https://doi.org/10.1109/cvpr42600.2020.01155
- [41] S. Woo, J. Park, J. Lee, and I. S. Kweon, “CBAM: Convolutional Block Attention Module,” Proc. of the European Conf. on Computer Vision (ECCV), 2018.
- [42] Q. Hou, D. Zhou, and J. Feng, “Coordinate Attention for Efficient Mobile Network Design,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 13708-13717,2021. https://doi.org/10.1109/cvpr46437.2021.01350
- [43] L. Kang, Z. Lu, L. Meng, and Z. Gao, “YOLO-FA: Type-1 Fuzzy Attention Based Yolo Detector for Vehicle Detection,” Expert Systems with Applications, Vol.237, Article No.121209, 2024. https://doi.org/10.1016/j.eswa.2023.121209
- [44] W. Wang, S. Zhao, J. Shen, S. C. H. Hoi, and A. Borji, “Salient Object Detection with Pyramid Attention and Salient Edges,” 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1448-1457, 2019. https://doi.org/10.1109/cvpr.2019.00154
- [45] M.-H. Guo, Z.-N. Liu, T.-J. Mu, and S.-M. Hu, “Beyond Self-Attention: External Attention Using Two Linear Layers for Visual Tasks,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, No.5, pp. 5436-5447, 2023. https://doi.org/10.1109/TPAMI.2022.3211006
- [46] S. Zhang, Y. Fu, X. Zhao, J. Fang, Y. Liu, X. Wang, B. Zhang, and J. Yu, “Sequence-to-Point Learning Based on Spatio-Temporal Attention Fusion Network for Non-Intrusive Load Monitoring,” Complex & Intelligent Systems, Vol.11, Article No.171, 2025. https://doi.org/10.1007/s40747-025-01803-1
- [47] K. Wu, Y. Xu, and J. Zhang, “Lightweight Multi-Scale Dynamic Feature Focusing Network Integrating Spatial Channel Attention Mechanism for Autonomous Driving Object Detection,” Digital Signal Processing, Vol.170, Article No.105770, 2025. https://doi.org/10.1016/j.dsp.2025.105770
- [48] M. E. Aghili, H. Ghassemian, and M. Imani, “YOLO-PICO: Lightweight Object Recognition in Remote Sensing Images Using Expansion Attention Modules,” Pattern Recognition, Vol.176, Article No.113114, 2026. https://doi.org/10.1016/j.patcog.2026.113114
- [49] J. Yang, X. Yue, and L. Wu, “A Collaborative Multi-Attention Network for Real-Time Small Object Detection in UAV Imagery,” Scientific Reports, Vol.16, Article No.5852, 2026. https://doi.org/10.1038/s41598-026-36440-2
- [50] N. Anwar, G.-A. Bilodeau, and W. Bouachir, “Dual-Stream Attention With Multi-Modal Queries for Object Detection in Transportation Applications,” arXiv preprint, arXiv:2508.04868, 2025. https://doi.org/10.48550/arXiv.2508.04868
- [51] X. Yue and L. Meng, “YOLO-MSA: A Multiscale Stereoscopic Attention Network for Empty-Dish Recycling Robots,” IEEE Trans. on Instrumentation and Measurement, Vol.72, Article No.2528014, 2023. https://doi.org/10.1109/TIM.2023.3315355
- [52] K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” arXiv preprint, arXiv:1409.1556, 2014. https://doi.org/10.48550/arXiv.1409.1556
- [53] K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” 2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 770-778, 2016. https://doi.org/10.1109/cvpr.2016.90
- [54] G. Huang, Z. Liu, L. V. D. Maaten, and K. Q. Weinberger, “Densely Connected Convolutional Networks,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 2261-2269, 2017. https://doi.org/10.1109/cvpr.2017.243
- [55] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” 2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 2818-2826, 2016. https://doi.org/10.1109/cvpr.2016.308
- [56] F. Chollet, “Xception: Deep Learning With Depthwise Separable Convolutions,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1800-1807, 2017. https://doi.org/10.1109/cvpr.2017.195
- [57] A. Howard, M. Sandler, B. Chen, W. Wang, L.-C. Chen, M. Tan, G. Chu, V. Vasudevan, Y. Zhu, R. Pang, H. Adam, and Q. Le, “Searching for MobileNetV3,” 2019 IEEE/CVF Int. Conf. on Computer Vision (ICCV), pp. 1314-1324, 2019. https://doi.org/10.1109/iccv.2019.00140
- [58] K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, “GhostNet: More Features From Cheap Operations,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1577-1586, 2020. https://doi.org/10.1109/cvpr42600.2020.00165
- [59] M. Tan, R. Pang, and Q. V. Le, “EfficientDet: Scalable and Efficient Object Detection,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 10778-10787, 2020. https://doi.org/10.1109/cvpr42600.2020.01079
- [60] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C. Fu, and A. C. Berg, “SSD: Single Shot MultiBox Detector,” 2016 European Conf. on Computer Vision (ECCV), pp. 21-37, 2016. https://doi.org/10.1007/978-3-319-46448-0_2
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.