Research Paper:
YOLO-GF: A Lightweight Detector for Character-Level Recognition in Medicine Packaging
Shiyong Geng*1,, Jintao Chen*2, Kuozhan Wang*3, Hengyi Li*3
, Xuebin Yue*3
, and Lin Meng*4

*1School of Integrated Circuits, Zhongyuan University of Technology
No.1 Huaihe Road, Longhu Town, Xinzheng, Zhengzhou, Henan 451191, China
Corresponding author
*2The First Affiliated Hospital of Henan University of CM
No.19 Renmin Road, Jinshui District, Zhengzhou, Henan 450000, China
*3School of Automation and Electrical Engineering, Zhongyuan University of Technology
No.41 Zhongyuan Middle Road, Zhongyuan District, Zhengzhou, Henan 450007, China
*4College of Science and Engineering, Ritsumeikan University
1-1-1 Nojihigashi, Kusatsu, Shiga 525-8577, Japan
The aging global population has significantly increased the demand for long-term care, presenting various challenges in elderly healthcare, especially in medication management. Ensuring timely and accurate medication delivery is essential to preventing medication errors, which could otherwise lead to severe health complications. This paper proposes an innovative solution to automate and enhance medication verification in nursing homes through object detection. The proposed system leverages the YOLO-GF framework, which integrates efficient computational techniques with high detection accuracy to enable real-time medicine package detection. A hybrid down sampling (HDS) module, combining max pooling, average pooling, and 2×2 stride convolution, optimizes feature map processing by reducing computational overhead while preserving key spatial information. Additionally, an enhanced multi-scale pyramid pooling (EMSPP) technique is introduced to improve multi-scale feature aggregation, enhancing the model’s ability to capture object features at various resolutions. Unlike existing detection systems that often suffer from high computational cost or poor generalization, our approach explicitly balances speed and accuracy through lightweight architectural design. Extensive experiments validate the effectiveness of the HDS and EMSPP modules in multi-scale feature learning. The YOLO-GF framework achieves a 100.00% mean average precision (mAP) on the medicine package dataset, with an inference speed of 130.16 frames per second, fully meeting the real-time monitoring requirements. Further evaluation of public datasets, including Dish20 (98.30% mAP) and Barcodes (97.34% mAP), demonstrates that YOLO-GF outperforms existing state-of-the-art models. These results highlight the superior generalization capabilities of the proposed framework. This work establishes a novel and efficient real-time solution for medicine package verification by integrating HDS and multi-scale enhancement into a unified YOLO-based framework, significantly improving accuracy and safety over manual processes in healthcare environments. The project can be found at http://www.ihpc.se.ritsumei.ac.jp/obidataset.html.
YOLO-GF overall framework
- [1] Y. Wang, S. Cang, and H. Yu, “A survey on wearable sensor modality centred human activity recognition in health care,” Expert Systems with Applications, Vol.137, pp. 167-190, 2019. https://doi.org/10.1016/j.eswa.2019.04.057
- [2] X. Yue, B. Lyu, H. Li, L. Meng, and K. Furumoto, “Real-time medicine packet recognition system in dispensing medicines for the elderly,” Measurement: Sensors, Vol.18, Article No.100072, 2021. https://doi.org/10.1016/j.measen.2021.100072
- [3] B. Lyu, Z. Wang, H. Li, A. Tanaka, K. Funumoto, and L. Meng, “Deep Leaning based Medicine Packaging Information Recognition for Medication Use in the Elderly,” Procedia Computer Science, Vol.187, pp. 194-199, 2021. https://doi.org/10.1016/j.procs.2021.04.108
- [4] Y. Zhou, “A YOLO-NL object detector for real-time detection,” Expert Systems with Applications, Vol.238, Part E, Article No.122256, 2024. https://doi.org/10.1016/j.eswa.2023.122256
- [5] X. Yue, H. Li, M. Shimizu, S. Kawamura, and L. Meng, “YOLO-GD: A Deep Learning-Based Object Detection Algorithm for Empty-Dish Recycling Robots,” Machines, Vol.10, Issue 5, Article No.294, 2022. https://doi.org/10.3390/machines10050294
- [6] Y. Ge, Z. Li, X. Yue, H. Li, Q. Li, and L. Meng, “IoT-based automatic deep learning model generation and the application on empty-dish recycling robots,” Internet of Things, Vol.25, Article No.101047, 2024. https://doi.org/10.1016/j.iot.2023.101047
- [7] X. Yue, H. Li, Y. Fujikawa, and L. Meng, “Dynamic Dataset Augmentation for Deep Learning-Based Oracle Bone Inscriptions Recognition,” ACM J. Comput. Cult. Herit., Vol.15, Issue 4, Article No.76, 2022. https://doi.org/10.1145/3532868
- [8] X. Yue, Z. Wang, R. Ishibashi, H. Kaneko, and L. Meng, “An unsupervised automatic organization method for Professor Shirakawa’s hand-notated documents of oracle bone inscriptions,” Int. J. on Document Analysis and Recognition (IJDAR), Vol.27, pp. 583-601, 2024. https://doi.org/10.1007/s10032-024-00463-0
- [9] N. Ullah, I. De Falco, and G. Sannino, “A Novel Deep Learning Approach for Colon and Lung Cancer Classification Using Histopathological Images,” 2023 IEEE 19th Int. Conf. on e-Science (e-Science), 2023. https://doi.org/10.1109/e-Science58273.2023.10254909
- [10] N. Ullah, M. Hassan, J. A. Khan et al., “Enhancing explainability in brain tumor detection: A novel DeepEBTDNet model with LIME on MRI images,” Int. J. of Imaging Systems and Technology, Vol.34, Issue 1, Article No.e23012, 2024. https://doi.org/10.1002/ima.23012
- [11] N. Ullah, J. A. Khan et al., “A Lightweight Deep Learning-Based Model for Tomato Leaf Disease Classification,” Computers, Materials & Continua, Vol.77, No.3, pp. 3969-3992, 2023. https://doi.org/10.32604/cmc.2023.041819
- [12] G. Jocher, A. Stoken, J. Borovec et al., “ultralytics/yolov5: v3.0,” Zenodo, 2020. https://doi.org/10.5281/zenodo.3983579
- [13] A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection,” arXiv:2004.10934, 2020. https://doi.org/10.48550/arXiv.2004.10934
- [14] Z. Ge, S. Liu, F. Wang, Z. Li, and J. Sun, “YOLOX: Exceeding YOLO Series in 2021,” arXiv:2107.08430, 2021. https://doi.org/10.48550/arXiv.2107.08430
- [15] R. Girshick, “Fast R-CNN,” 2015 IEEE Int. Conf. on Computer Vision (ICCV), pp. 1440-1448, 2015. https://doi.org/10.1109/ICCV.2015.169
- [16] S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” arXiv:1506.01497, 2015. https://doi.org/10.48550/arXiv.1506.01497
- [17] K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 2980-2988, 2017. https://doi.org/10.1109/ICCV.2017.322
- [18] W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg, “SSD: Single Shot MultiBox Detector,” 14th European Conf. on Computer Vision (ECCV 2016), pp. 21-37, 2016. https://doi.org/10.1007/978-3-319-46448-0_2
- [19] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” 2017 IEEE Int. Conf. on Computer Vision (ICCV), pp. 2999-3007, 2017. https://doi.org/10.1109/ICCV.2017.324
- [20] J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” arXiv:1804.02767, 2018. https://doi.org/10.48550/arXiv.1804.02767
- [21] C.-Y. Wang, A. Bochkovskiy, and H.-Y. M. Liao, “YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors,” 2023 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7464-7475, 2023. https://doi.org/10.1109/CVPR52729.2023.00721
- [22] X. Yue, H. Li, and L. Meng, “An Ultralightweight Object Detection Network for Empty-Dish Recycling Robots,” IEEE Trans. on Instrumentation and Measurement, Vol.72, 2023. https://doi.org/10.1109/TIM.2023.3241078
- [23] W. Zhou, C. Cai, C. Li, H. Xu, and H. Shi, “AD-YOLO: A Real-Time YOLO Network with Swin Transformer and Attention Mechanism for Airport Scene Detection,” IEEE Trans. on Instrumentation and Measurement, Vol.73, 2024. https://doi.org/10.1109/TIM.2024.3472805
- [24] Y. Weng, X. Xiang, and L. Ma, “SCR-YOLOv8: an enhanced algorithm for target detection in sonar images,” J. of Real-Time Image Processing, Vol.22, Article No.62, 2025. https://doi.org/10.1007/s11554-025-01637-7
- [25] F. Hao, Z. Zhang, D. Ma, and H. Kong, “GSBF-YOLO: A lightweight model for tomato ripeness detection in natural environments,” J. of Real-Time Image Processing, Vol.22, Article No.47, 2025. https://doi.org/10.1007/s11554-025-01624-y
- [26] J.-J. Liu, Q. Hou, Z.-A. Liu, and M.-M. Cheng, “PoolNet+: Exploring the Potential of Pooling for Salient Object Detection,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, Issue 1, pp. 887-904, 2023. https://doi.org/10.1109/TPAMI.2021.3140168
- [27] Z. Huo, J. Yang, M. Xi, D. Chen, and J. Wen, “SP-MSResNet: A Multiscale Residual Network with Strip Pooling Module for Intrusion Pattern Recognition,” IEEE Sensors J., Vol.24, Issue 19, pp. 30136-30146, 2024. https://doi.org/10.1109/JSEN.2024.3444917
- [28] Y. Yang, L. Jiao, X. Liu, L. Li, F. Liu, S. Yang, and X. Zhang, “Efficient LWPooling: Rethinking the Wavelet Pooling for Scene Parsing,” IEEE Trans. on Circuits and Systems for Video Technology, Vol.34, Issue 9, pp. 8481-8493, 2024. https://doi.org/10.1109/TCSVT.2024.3383072
- [29] Y.-H. Wu, Y. Liu, X. Zhan, and M.-M. Cheng, “P2T: Pyramid Pooling Transformer for Scene Understanding,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.45, Issue 11, pp. 12760-12771, 2023. https://doi.org/10.1109/TPAMI.2022.3202765
- [30] F. Zhu, X. Zhang, B. Zhang, Y. Xu, and L. Cui, “Medicine Package Recommendation via Dual-Level Interaction Aware Heterogeneous Graph,” IEEE J. of Biomedical and Health Informatics, Vol.28, Issue 4, pp. 2294-2303, 2024. https://doi.org/10.1109/JBHI.2024.3361552
- [31] N. Ullah, J. A. Khan, I. De Falco, and G. Sannino, “Bridging Clinical Gaps: Multi-Dataset Integration for Reliable Multi-Class Lung Disease Classification with DeepCRINet and Occlusion Sensitivity,” 2024 IEEE Symp. on Computers and Communications (ISCC), 2024. https://doi.org/10.1109/ISCC61673.2024.10733651
- [32] X. Zhou, C. Yao, H. Wen, Y. Wang, S. Zhou, W. He, and J. Liang, “EAST: An Efficient and Accurate Scene Text Detector,” 2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 2642-2651, 2017. https://doi.org/10.1109/CVPR.2017.283
- [33] Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee, “Character Region Awareness for Text Detection,” 2019 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 9357-9366, 2019. https://doi.org/10.1109/CVPR.2019.00959
- [34] M. Liao, Z. Wan, C. Yao, K. Chen, and X. Bai, “Real-Time Scene Text Detection with Differentiable Binarization,” Proc. of the AAAI Conf. on Artificial Intelligence, Vol.34, No.7, pp. 11474-11481, 2020. https://doi.org/10.1609/aaai.v34i07.6812
- [35] C. Peng, X. Li, and Y. Wang, “TD-YOLOA: An Efficient YOLO Network with Attention Mechanism for Tire Defect Detection,” IEEE Trans. on Instrumentation and Measurement, Vol.72, 2023. https://doi.org/10.1109/TIM.2023.3312753
- [36] J. Hu, L. Shen, and G. Sun, “Squeeze-and-Excitation Networks,” 2018 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 7132-7141, 2018. https://doi.org/10.1109/CVPR.2018.00745
- [37] Q. Wang, B. Wu, P. Zhu, P. Li, W. Zuo, and Q. Hu, “ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 11531-11539, 2020. https://doi.org/10.1109/CVPR42600.2020.01155
- [38] S. Woo, J. Park, J.-Y. Lee, and I. S. Kweon, “CBAM: Convolutional Block Attention Module,” 15th European Conf. (ECCV 2018), 2018. https://doi.org/10.1007/978-3-030-01234-2_1
- [39] Z. Li, Z. Xie, P. Duan, X. Kang, and S. Li, “Dual Spatial Attention Network for Underwater Object Detection with Sonar Imagery,” IEEE Sensors J., Vol.24, Issue 5, pp. 6998-7008, 2024. https://doi.org/10.1109/JSEN.2023.3336899
- [40] X. Lyu, L. Tian, and S. Teng, “ECDet: Efficient oriented object detection on the aerial image with cross-layer attention,” J. of Real-Time Image Processing, Vol.22, Article No.42, 2025. https://doi.org/10.1007/s11554-024-01617-3
- [41] Y. Zhang, Z. Tu, Y. Zheng, T. Zhang, C. Wu, and N. Wang, “Parallel Attention for Multitask Road Object Detection in Autonomous Driving,” IEEE Sensors J., Vol.24, Issue 21, pp. 35975-35985, 2024. https://doi.org/10.1109/JSEN.2024.3454773
- [42] N. Ullah, J. A. Khan, I. De Falco, and G. Sannino, “Explainable Artificial Intelligence: Importance, Use Domains, Stages, Output Shapes, and Challenges,” ACM Computing Surveys, Vol.57, Issue 4, Article No.94, 2024. https://doi.org/10.1145/3705724
- [43] K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, “GhostNet: More Features From Cheap Operations,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 1577-1586, 2020. https://doi.org/10.1109/CVPR42600.2020.00165
- [44] C.-Y. Wang, H.-Y. M. Liao, Y.-H. Wu, P.-Y. Chen, J.-W. Hsieh, and I.-H. Yeh, “CSPNet: A New Backbone that can Enhance Learning Capability of CNN,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 1571-1580, 2020. https://doi.org/10.1109/CVPRW50498.2020.00203
- [45] X. Yue and L. Meng, “YOLO-SM: A Lightweight Single-Class Multi-Deformation Object Detection Network,” IEEE Trans. on Emerging Topics in Computational Intelligence, Vol.8, Issue 3, pp. 2467-2480, 2024. https://doi.org/10.1109/TETCI.2024.3367821
- [46] Q. Chen, Y. Wang, T. Yang, X. Zhang, J. Cheng, and J. Sun, “You Only Look One-level Feature,” 2021 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 13034-13043, 2021. https://doi.org/10.1109/CVPR46437.2021.01284
- [47] P. Hurtik, V. Molek, J. Hula, M. Vajgl, P. Vlasanek, and T. Nejezchleba, “Poly-YOLO: Higher speed, more precise detection and instance segmentation for YOLOv3,” Neural Computing and Applications, Vol.34, No.10, pp. 8275-8290, 2022. https://doi.org/10.1007/s00521-021-05978-9
- [48] A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images,” arXiv preprint, arXiv:1601.07140, 2016. https://doi.org/10.48550/arXiv.1601.07140
- [49] C.-K. Ch’ng, C. S. Chan, and C.-L. Liu, “Total-Text: Toward orientation robustness in scene text detection,” Int. J. on Document Analysis and Recognition (IJDAR), Vol.23, No.1, pp. 31-52, 2020. https://doi.org/10.1007/s10032-019-00334-z
- [50] X. Yue and L. Meng, “YOLO-MSA: A Multiscale Stereoscopic Attention Network for Empty-Dish Recycling Robots,” IEEE Trans. on Instrumentation and Measurement, Vol.72, 2023. https://doi.org/10.1109/TIM.2023.3315355
- [51] X. Yan, L. Jia, H. Cao, Y. Yu, T. Wang, F. Zhang, and Q. Guan, “Multitargets Joint Training Lightweight Model for Object Detection of Substation,” IEEE Trans. on Neural Networks and Learning Systems, Vol.35, Issue 2, pp. 2413-2424, 2022. https://doi.org/10.1109/TNNLS.2022.3190139
- [52] M. Tan, R. Pang, and Q. V. Le, “EfficientDet: Scalable and Efficient Object Detection,” 2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR), pp. 10778-10787, 2020. https://doi.org/10.1109/CVPR42600.2020.01079
- [53] Z. Tian, C. Shen, H. Chen, and T. He, “FCOS: A Simple and Strong Anchor-Free Object Detector,” IEEE Trans. on Pattern Analysis and Machine Intelligence, Vol.44, Issue 4, pp. 1922-1933, 2022. https://doi.org/10.1109/TPAMI.2020.3032166
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.