Research Paper:
Optimizing Speech Recognition and Feedback in Spoken English Learning Using Long Short-Term Memory Networks
Yu Shang and Shang Xue
Zhengzhou Railway Vocational and Technical College
No.9 Qiancheng Road, Zhengdong NewDistrict, Zhengzhou, Henan 451460, China
Corresponding author
In the era of globalization, learning English through oral means is becoming increasingly important. However, inaccurate speech recognition and imperfect feedback mechanisms have significantly hindered the improvement of the oral ability of learners. To solve this problem, this study proposes an English oral learning speech recognition and feedback system based on a multilayer improved long short-term memory network (MLSTM). The study uses the Texas Instruments and Massachusetts Institute of Technology (TIMIT) speech database and employs a hidden Markov model (HMM) and standard LSTM systems as controls to conduct a comprehensive test of the proposed model. The experimental results showed that the MLSTM model achieved high speech recognition accuracy, with an overall accuracy of 86.3% that was significantly higher than the 65.2% of HMM and 80.1% of LSTM. In terms of feedback information targeting and effectiveness, the MLSTM model scored 4.35 and 4.5, respectively, that were suggestively better than those of the comparison models. This showed that the MLSTM model could accurately recognize speech and provide learners with highly personalized and effective feedback. The research results enrich the theory of computer-assisted language learning, provide practical and effective tools for oral English learning, and promote oral English learning toward greater intelligence and efficiency.
- [1] E. Shafaei-Bajestan, M. Moradipour-Tari, P. Uhrig, and R. H. Baayen, “LDL-AURIS: A computational model, grounded in error-driven learning, for the comprehension of single spoken words,” Language Cognition and Neuroscience, Vol.38, No.4, pp. 509-536, 2023. https://doi.org/10.1080/23273798.2021.1954207
- [2] K. Junttila, A-R. Smolander, R. Karhila, M. Kurimo, and S. Ylinen, “Non-game like training benefits spoken foreign-language processing in children with dyslexia,” Frontiers in Human Neuroscience, Vol.17, Article No.1122886, 2023. https://doi.org/10.3389/fnhum.2023.1122886
- [3] J. Yang, A. Wagner, Y. Zhang, and L. Xu, “Recognition of vocoded speech in English by Mandarin-speaking English-learners,” Speech Communication, Vol.136, pp. 63-75, 2022. https://doi.org/10.1016/j.specom.2021.11.008
- [4] G. Karakasidis, M. Kurimo, P. Bell, and T. Grósz, “Comparison and analysis of new curriculum criteria for end-to-end ASR,” Speech Communication, Vol.163, Article No.103113, 2024. https://doi.org/10.1016/j.specom.2024.103113
- [5] J. Xu and T. Li, “Application of multimodal NLP instruction combined with speech recognition in oral English practice,” Mobile Information Systems, Article No.222696, 2022. https://doi.org/10.1155/2022/2262696
- [6] H. Zhang, “Research on spoken English analysis model based on transfer learning and machine learning algorithms,” J. of Intelligent & Fuzzy Systems, Vol.38, No.6, pp. 7377-7387, 2020. https://doi.org/10.3233/jifs-179811
- [7] C. Y. Tzeng, M. L. Russell, and L. C. Nygaard, “Attention modulates perceptual learning of non-native-accented speech,” Attention Perception & Psychophysics, Vol.86, pp. 339-353, 2024. https://doi.org/10.3758/s13414-023-02790-6
- [8] Y. Chen and B. Martinuzzi, “Machine learning for predictive analytics in the improvement of English speech feature recognition,” Mobile Information Systems, Article No.3541667, 2022. https://doi.org/10.1155/2022/3541667
- [9] M. Habbash, S. Mnasri, M. Alghamdi, M. Alrashidi, A. S. Tarawneh, A. Gumair et al., “Recognition of Arabic accents from English spoken speech using deep learning approach,” IEEE Access, Vol.12, pp. 37219-37230, 2024. https://doi.org/10.1109/access.2024.3374768
- [10] M. Zhu, “The Application of intelligent speech analysis technology in the spoken English language learning model,” Mobile Information Systems, Article No.3192892, 2022. https://doi.org/10.1155/2022/3192892
- [11] J. Oruh and S. Viriri, “Deep Learning-based classification of spoken English digits,” Computational Intelligence and Neuroscience, Article No.3364141, 2022. https://doi.org/10.1155/2022/3364141
- [12] R. E. Bieber, M. J. Makashay, B. Simpson, B. M. Sheffield, and D. S. Brungart, “Short-term retention of learning after rapid adaptation to native and non-native speech,” J. of the Acoustical Society of America, Vol.153, No.6, pp. 3362-3371, 2023. https://doi.org/10.1121/10.0019749
- [13] H. M. Zhu, “Construction of English spoken language system based on machine learning algorithm and natural language recognition,” J. of Intelligent & Fuzzy Systems, Vol.39, No.4, pp. 4891-4902, 2020. https://doi.org/10.3233/jifs-179975
- [14] X. Xu and K. X. Xiao, “Oral business English recognition method based on RankNet model and endpoint detection algorithm,” J. of Sensors, Article No.7426303, 2022. https://doi.org/10.1155/2022/7426303
- [15] P. Drozdova, R. van Hout, S. Mattys, and O. Scharenborg, “The effect of intermittent noise on lexically-guided perceptual learning in native and non-native listening,” Speech Communication, Vol.126, pp. 61-70, 2021. https://doi.org/10.1016/j.specom.2020.12.002
- [16] A. Srinivasan, D. Singh, C. Yarra, A. Illa, and P. K. Ghosh, “A robust speaking rate estimator using a CNN-BLSTM Network,” Circuits Systems and Signal Processing, Vol.40, No.12, pp. 6098-6120, 2021. https://doi.org/10.1007/s00034-021-01754-1
- [17] L. Wang, “A machine learning assessment system for spoken English based on linear predictive coding,” Mobile Information Systems, 2022. https://doi.org/10.1155/2022/6131572
- [18] J. Oruh, S. Viriri, and A. Adegun, “Long short-term memory recurrent neural network for automatic speech recognition,” IEEE Access, Vol.10, pp. 30069-30079, 2022. https://doi.org/10.1109/access.2022.3159339
- [19] C. J. Wang, “Multimedia network English reading teaching model based on speech recognition confidence learning algorithm,” Mathematical Problems in Engineering, Article No.5641528, 2021. https://doi.org/10.1155/2021/5641528
- [20] J. Yang, J. Barrett, Z. G. Yin, and L. Xu, “Recognition of foreign-accented vocoded speech by native English listeners,” Acta Acustica, Vol.7, Arrticle No.43, 2023. https://doi.org/10.1051/aacus/2023038
- [21] L. Y. Wang, J. Yang, Y. S. Wang, Y. Qi, S. Wang, and J. Li, “Integrating Large Language Models (LLMs) and deep representations of emotional features for the recognition and evaluation of emotions in spoken English,” Applied Sciences-Basel, Vol.14, No.9, Article No.3543, 2024. https://doi.org/10.3390/app14093543
- [22] T. A. Ayall, C. J. Zhou, H. W. Liu, G. M. Brhanemeskel, S. T. Abate, and M. Adjeisah, “Amharic spoken digits recognition using convolutional neural network,” J. of Big Data, Vol.11, Article No.64, 2024. https://doi.org/10.1186/s40537-024-00910-z
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.