single-jc.php

JACIII Vol.30 No.4 pp. 966-976
(2026)

Research Paper:

Optimizing Speech Recognition and Feedback in Spoken English Learning Using Long Short-Term Memory Networks

Yu Shang and Shang Xue

Zhengzhou Railway Vocational and Technical College
No.9 Qiancheng Road, Zhengdong NewDistrict, Zhengzhou, Henan 451460, China

Corresponding author

Received:
July 23, 2025
Accepted:
January 25, 2026
Published:
July 20, 2026
Keywords:
oral English learning, long short-term memory network, speech recognition, feedback mechanism, machine learning
Abstract

In the era of globalization, learning English through oral means is becoming increasingly important. However, inaccurate speech recognition and imperfect feedback mechanisms have significantly hindered the improvement of the oral ability of learners. To solve this problem, this study proposes an English oral learning speech recognition and feedback system based on a multilayer improved long short-term memory network (MLSTM). The study uses the Texas Instruments and Massachusetts Institute of Technology (TIMIT) speech database and employs a hidden Markov model (HMM) and standard LSTM systems as controls to conduct a comprehensive test of the proposed model. The experimental results showed that the MLSTM model achieved high speech recognition accuracy, with an overall accuracy of 86.3% that was significantly higher than the 65.2% of HMM and 80.1% of LSTM. In terms of feedback information targeting and effectiveness, the MLSTM model scored 4.35 and 4.5, respectively, that were suggestively better than those of the comparison models. This showed that the MLSTM model could accurately recognize speech and provide learners with highly personalized and effective feedback. The research results enrich the theory of computer-assisted language learning, provide practical and effective tools for oral English learning, and promote oral English learning toward greater intelligence and efficiency.

Cite this article as:
Y. Shang and S. Xue, “Optimizing Speech Recognition and Feedback in Spoken English Learning Using Long Short-Term Memory Networks,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 966-976, 2026.
Data files:
References
  1. [1] E. Shafaei-Bajestan, M. Moradipour-Tari, P. Uhrig, and R. H. Baayen, “LDL-AURIS: A computational model, grounded in error-driven learning, for the comprehension of single spoken words,” Language Cognition and Neuroscience, Vol.38, No.4, pp. 509-536, 2023. https://doi.org/10.1080/23273798.2021.1954207
  2. [2] K. Junttila, A-R. Smolander, R. Karhila, M. Kurimo, and S. Ylinen, “Non-game like training benefits spoken foreign-language processing in children with dyslexia,” Frontiers in Human Neuroscience, Vol.17, Article No.1122886, 2023. https://doi.org/10.3389/fnhum.2023.1122886
  3. [3] J. Yang, A. Wagner, Y. Zhang, and L. Xu, “Recognition of vocoded speech in English by Mandarin-speaking English-learners,” Speech Communication, Vol.136, pp. 63-75, 2022. https://doi.org/10.1016/j.specom.2021.11.008
  4. [4] G. Karakasidis, M. Kurimo, P. Bell, and T. Grósz, “Comparison and analysis of new curriculum criteria for end-to-end ASR,” Speech Communication, Vol.163, Article No.103113, 2024. https://doi.org/10.1016/j.specom.2024.103113
  5. [5] J. Xu and T. Li, “Application of multimodal NLP instruction combined with speech recognition in oral English practice,” Mobile Information Systems, Article No.222696, 2022. https://doi.org/10.1155/2022/2262696
  6. [6] H. Zhang, “Research on spoken English analysis model based on transfer learning and machine learning algorithms,” J. of Intelligent & Fuzzy Systems, Vol.38, No.6, pp. 7377-7387, 2020. https://doi.org/10.3233/jifs-179811
  7. [7] C. Y. Tzeng, M. L. Russell, and L. C. Nygaard, “Attention modulates perceptual learning of non-native-accented speech,” Attention Perception & Psychophysics, Vol.86, pp. 339-353, 2024. https://doi.org/10.3758/s13414-023-02790-6
  8. [8] Y. Chen and B. Martinuzzi, “Machine learning for predictive analytics in the improvement of English speech feature recognition,” Mobile Information Systems, Article No.3541667, 2022. https://doi.org/10.1155/2022/3541667
  9. [9] M. Habbash, S. Mnasri, M. Alghamdi, M. Alrashidi, A. S. Tarawneh, A. Gumair et al., “Recognition of Arabic accents from English spoken speech using deep learning approach,” IEEE Access, Vol.12, pp. 37219-37230, 2024. https://doi.org/10.1109/access.2024.3374768
  10. [10] M. Zhu, “The Application of intelligent speech analysis technology in the spoken English language learning model,” Mobile Information Systems, Article No.3192892, 2022. https://doi.org/10.1155/2022/3192892
  11. [11] J. Oruh and S. Viriri, “Deep Learning-based classification of spoken English digits,” Computational Intelligence and Neuroscience, Article No.3364141, 2022. https://doi.org/10.1155/2022/3364141
  12. [12] R. E. Bieber, M. J. Makashay, B. Simpson, B. M. Sheffield, and D. S. Brungart, “Short-term retention of learning after rapid adaptation to native and non-native speech,” J. of the Acoustical Society of America, Vol.153, No.6, pp. 3362-3371, 2023. https://doi.org/10.1121/10.0019749
  13. [13] H. M. Zhu, “Construction of English spoken language system based on machine learning algorithm and natural language recognition,” J. of Intelligent & Fuzzy Systems, Vol.39, No.4, pp. 4891-4902, 2020. https://doi.org/10.3233/jifs-179975
  14. [14] X. Xu and K. X. Xiao, “Oral business English recognition method based on RankNet model and endpoint detection algorithm,” J. of Sensors, Article No.7426303, 2022. https://doi.org/10.1155/2022/7426303
  15. [15] P. Drozdova, R. van Hout, S. Mattys, and O. Scharenborg, “The effect of intermittent noise on lexically-guided perceptual learning in native and non-native listening,” Speech Communication, Vol.126, pp. 61-70, 2021. https://doi.org/10.1016/j.specom.2020.12.002
  16. [16] A. Srinivasan, D. Singh, C. Yarra, A. Illa, and P. K. Ghosh, “A robust speaking rate estimator using a CNN-BLSTM Network,” Circuits Systems and Signal Processing, Vol.40, No.12, pp. 6098-6120, 2021. https://doi.org/10.1007/s00034-021-01754-1
  17. [17] L. Wang, “A machine learning assessment system for spoken English based on linear predictive coding,” Mobile Information Systems, 2022. https://doi.org/10.1155/2022/6131572
  18. [18] J. Oruh, S. Viriri, and A. Adegun, “Long short-term memory recurrent neural network for automatic speech recognition,” IEEE Access, Vol.10, pp. 30069-30079, 2022. https://doi.org/10.1109/access.2022.3159339
  19. [19] C. J. Wang, “Multimedia network English reading teaching model based on speech recognition confidence learning algorithm,” Mathematical Problems in Engineering, Article No.5641528, 2021. https://doi.org/10.1155/2021/5641528
  20. [20] J. Yang, J. Barrett, Z. G. Yin, and L. Xu, “Recognition of foreign-accented vocoded speech by native English listeners,” Acta Acustica, Vol.7, Arrticle No.43, 2023. https://doi.org/10.1051/aacus/2023038
  21. [21] L. Y. Wang, J. Yang, Y. S. Wang, Y. Qi, S. Wang, and J. Li, “Integrating Large Language Models (LLMs) and deep representations of emotional features for the recognition and evaluation of emotions in spoken English,” Applied Sciences-Basel, Vol.14, No.9, Article No.3543, 2024. https://doi.org/10.3390/app14093543
  22. [22] T. A. Ayall, C. J. Zhou, H. W. Liu, G. M. Brhanemeskel, S. T. Abate, and M. Adjeisah, “Amharic spoken digits recognition using convolutional neural network,” J. of Big Data, Vol.11, Article No.64, 2024. https://doi.org/10.1186/s40537-024-00910-z

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026