single-jc.php

JACIII Vol.30 No.4 pp. 989-1003
(2026)

Research Paper:

Automatic Music Generation Model Integrating Gradient Penalty Strategy and Wasserstein Distance

Yuhong He ORCID Icon

School of Music and Dance, Aba Teachers University
Shuimo Town, Wenchuan, Aba, Sichuan 623002, China

Received:
August 28, 2025
Accepted:
January 27, 2026
Published:
July 20, 2026
Keywords:
automatic generation of music, gradient penalty, Wasserstein distance, MuseGAN, VGGish
Abstract

This study addresses the problems of imbalanced voice sequence, insufficient training stability, and symbol-audio modality mismatch in multi-track music generation. To this end, a gradient penalty-constrained multi-track adversarial training framework and a Wasserstein distance-driven cross-modal distribution alignment mechanism are studied and designed to optimize voice coordination accuracy and auditory perception authenticity. The core innovation lies in building a dual-module collaborative optimization architecture, pioneering gradient penalty constraints to eliminate multi-track training oscillations, and proposing Wasserstein feature space mapping mechanism to eradicate cross-modal perception mismatch, establishing a theoretical paradigm and technical path for joint optimization of voice and sound effects. Experimental verification: the voice conflict rate reaches 0.60%, the gradient norm variance is 1.13×10-3, and the training stability is controlled. The cross-modal distribution distance is 0.28, achieving precise alignment. The feature alignment error of 0.17 exceeds the technical limit, and the style fidelity is 91.7% (Baroque 94% / Jazz Blues 95%) to restore artistic expression. The dynamic expressive power is 4.66 points, approaching human creativity, and the auditory similarity is 0.89, establishing perceptual authenticity. The conflict rate of voice parts in the ablation test decreases by 52%, and cross-modal mismatch compression is reduced by 49%. Parameter sensitivity analysis shows that when the hidden space dimension is 64, the structural entropy is 0.810 and the perceptual similarity is 0.87, reaching the global optimum. This model significantly improves the coordination of multi-track structures and cross-modal perception quality, providing core technical support solutions for industrial-grade artificial intelligence (AI) music creation platforms such as film and television music composition and digital composition.

Cite this article as:
Y. He, “Automatic Music Generation Model Integrating Gradient Penalty Strategy and Wasserstein Distance,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 989-1003, 2026.
Data files:
References
  1. [1] P. P. Groumpos, “A critical historic overview of artificial intelligence: Issues, challenges, opportunities, and threats,” Artif. Intell. Appl., Vol.1, No.4, pp. 181-197, 2023. https://doi.org/10.47852/bonviewAIA3202689
  2. [2] S. Ji, X. Yang, and J. Luo, “A survey on deep learning for symbolic music generation: Representations, algorithms, evaluations, and challenges,” ACM Comput. Surv., Vol.56, No.1, Article No.7, 2023. https://doi.org/10.1145/3597493
  3. [3] L. Wang et al., “A review of intelligent music generation systems,” Neural Comput. Appl., Vol.36, No.12, pp. 6381-6401, 2024. https://doi.org/10.1007/s00521-024-09418-2
  4. [4] F. Ding and Y. Cui, “MuseFlow: Music accompaniment generation based on flow,” Appl. Intell., Vol.53, No.20, pp. 23029-23038, 2023. https://doi.org/10.1007/s10489-023-04664-8
  5. [5] S. Serrano, M. L. Scarpa, and O. Serghini, “VGGish for music/speech classification in radio broadcasting,” Proc. 38th ECMS Int. Conf. Model. Simul.Eur. Counc. Model. Simul., pp. 550-557, 2024. https://doi.org/10.7148/2024-0550
  6. [6] W. Liu, “Literature survey of multi-track music generation model based on generative confrontation network in intelligent composition,” J. Supercomput.., Vol.79, No.6, pp. 6560-6582, 2023. https://doi.org/10.1007/s11227-022-04914-5
  7. [7] Z. Qiu, H. Wang, C. Liao, Z. Lu, and Y. Kuang, “Sound recognition of harmful bird species related to power grid faults based on VGGish transfer learning,” J. Electr. Eng. Technol., Vol.18, No.3, pp. 2447-2456, 2023. https://doi.org/10.1007/s42835-022-01284-z
  8. [8] P. Luo, Z. Yin, D. Yuan, F. Gao, and J. Liu, “A novel generative adversarial networks via music theory knowledge for early fault intelligent diagnosis of motor bearings,” IEEE Trans. Ind. Electron., Vol.71, No.8, pp. 9777-9788, 2024. https://doi.org/10.1109/TIE.2023.3321984
  9. [9] A. Zhang et al., “EEG data augmentation for emotion recognition with a multiple generator conditional Wasserstein GAN,” Complex Intell. Syst., Vol.8, No.4, pp. 3059-3071, 2022. https://doi.org/10.1007/s40747-021-00336-7
  10. [10] S.-L. Wu, C. Donahue, S. Watanabe, and N. J. Bryan, “Music ControlNet: Multiple time-varying controls for music generation,” IEEE/ACM Trans. Audio Speech Lang. Process., Vol.32, pp. 2692-2703, 2024. https://doi.org/10.1109/TASLP.2024.3399026
  11. [11] S. Tian, C. Zhang, W. Yuan, W. Tan, and W. Zhu, “XMusic: Towards a generalized and controllable symbolic music generation framework,” IEEE Trans. Multimed., Vol.27, pp. 6857-6871, 2025. https://doi.org/10.1109/TMM.2025.3590912
  12. [12] F. Wang, “Application of artificial intelligence-based music generation technology in popular music production,” J. Comb. Math. Comb. Comput., Vol.127a, pp. 655-671, 2025. https://doi.org/10.61091/jcmcc127a-038
  13. [13] Y. Tie et al., “Hybrid learning module-based transformer for multitrack music generation with music theory,” IEEE Trans. Comput. Soc. Syst., Vol.12, No.2, pp. 862-872, 2025. https://doi.org/10.1109/TCSS.2024.3486604
  14. [14] Y.-J. Shih, S.-L. Wu, F. Zalkow, M. Müller, and Y.-H. Yang, “Theme transformer: Symbolic music generation with theme-conditioned transformer,” IEEE Trans. Multimed., Vol.25, pp. 3495-3508, 2023. https://doi.org/10.1109/TMM.2022.3161851
  15. [15] C. Jin et al., “A transformer generative adversarial network for multi-track music generation,” CAAI Trans. Intell. Technol., Vol.7, No.3, pp. 369-380, 2022. https://doi.org/10.1049/cit2.12065
  16. [16] O. Medialov, Y. Babiak, and T. Basyuk, “Evaluation and comparison of text-to-audio generation models for media applications,” Her. Khmelnytskyi Natl. Univ. Tech. Sci., Vol.351, No.3.1, pp. 28-34, 2025. https://doi.org/10.31891/2307-5732-2025-351-3
  17. [17] W. Huang, Y. Xue, Z. Xu, G. Peng, and Y. Wu, “Polyphonic music generation generative adversarial network with Markov decision process,” Multimed. Tools Appl., Vol.81, No.21, pp. 29865-29885, 2022. https://doi.org/10.1007/s11042-022-12925-w
  18. [18] D. Zhang, M. Ma, and L. Xia, “A comprehensive review on GANs for time-series signals,” Neural Comput. Appl., Vol.34, No.5, pp. 3551-3571, 2022. https://doi.org/10.1007/s00521-022-06888-0
  19. [19] C. Zhang, Y. Ren, K. Zhang, and S. Yan, “SDMuse: Stochastic differential music editing and generation via hybrid representation,” IEEE Trans. Multimed., Vol.26, pp. 1681-1689, 2024. https://doi.org/10.1109/TMM.2023.3284996
  20. [20] H. Wang, Y. Zou, H. Cheng, and L. Ye, “DiffuseRoll: Multi-track multi-attribute music generation based on diffusion model,” Multimed. Syst., Vol.30, No.1, Article No.19, 2024. https://doi.org/10.1007/s00530-023-01220-9
  21. [21] J. Min, Z. Gao, L. Wang, and A. Zhang, “Application research of short-time Fourier transform in music generation based on the parallel WaveGan system,” IEEE Trans. Ind. Inform., Vol.20, No.9, pp. 10770-10778, 2024. https://doi.org/10.1109/TII.2024.3397344
  22. [22] S. Zhao, Q. Li, T. He, and J. Wen, “A step-by-step gradient penalty with similarity calculation for text summary generation,” Neural Process. Lett., Vol.55, No.4, pp. 4111-4126, 2023. https://doi.org/10.1007/s11063-022-11031-0
  23. [23] C. R. Lekshmi and R. Rajeev, “Multiple predominant instruments recognition in polyphonic music using spectro/modgd-gram fusion,” Circuits Syst. Signal Process., Vol.42, No.6, pp. 3464-3484, 2023. https://doi.org/10.1007/s00034-022-02278-y
  24. [24] F. Liu, D.-L. Chen, R.-Z. Zhou, S. Yang, and F. Xu, “Self-supervised music motion synchronization learning for music-driven conducting motion generation,” J. Comput. Sci. Technol., Vol.37, No.3, pp. 539-558, 2022. https://doi.org/10.1007/s11390-022-2030-z
  25. [25] Y.-N. Hung et al., “A large TV dataset for speech and music activity detection,” EURASIP J. Audio Speech Music Process., Vol.2022, Article No.21, 2022. https://doi.org/10.1186/s13636-022-00253-8
  26. [26] T. Kanwal, R. Mahum, A. M. AlSalman, M. Sharaf, and H. Hassan, “Fake speech detection using VGGish with attention block,” EURASIP J. Audio Speech Music Process., Vol.2024, Article No.35, 2024. https://doi.org/10.1186/s13636-024-00348-4
  27. [27] F. Özcan and A. Alkan, “Explainable audio CNNs applied to neural decoding: Sound category identification from inferior colliculus,” Signal Image Video Process., Vol.18, No.2, pp. 1193-1204, 2024. https://doi.org/10.1007/s11760-023-02825-3
  28. [28] J. Cai, Y. Zhang, S. Wang, J. Fan, and W. Guo, “Wasserstein embedding learning for deep clustering: A generative approach,” IEEE Trans. Multimed., Vol.26, pp. 7567-7580, 2024. https://doi.org/10.1109/TMM.2024.3369862
  29. [29] Z. Su, G. Zhang, Z. Shi, D. Hu, and W. Zhang, “Message-driven generative music steganography using MIDI-GAN,” IEEE Trans. Dependable Secure Comput., Vol.21, No.6, pp. 5196-5207, 2024. https://doi.org/10.1109/TDSC.2024.3372139
  30. [30] J. Yang, J. Jiang, and Y. Guo, “MHA-WoML: Multi-head attention and Wasserstein-OT for few-shot learning,” Int. J. Multimed. Inf. Retr., Vol.11, No.4, pp. 681-694, 2022. https://doi.org/10.1007/s13735-022-00254-5
  31. [31] A. Saeed, M. F. Hayat, T. Habib, D. A. Ghaffar, and M. A. Qureshi, “A novel multi-speakers Urdu singing voices synthesizer using Wasserstein generative adversarial network,” Speech Commun., Vol.137, pp. 103-113, 2022. https://doi.org/10.1016/j.specom.2021.12.005
  32. [32] D. Edwards, S. Dixon, E. Benetos, A. Maezawa, and Y. Kusaka, “A data-driven analysis of robust automatic piano transcription,” IEEE Signal Process. Lett., Vol.31, pp. 681-685, 2024. https://doi.org/10.1109/LSP.2024.3363646
  33. [33] A. Dash and K. Agres, “AI-based affective music generation systems: A review of methods and challenges,” ACM Comput. Surv., Vol.56, No.11, Article No.287, 2024. https://doi.org/10.1145/3672554
  34. [34] Z. Yin, F. Reuben, S. Stepney, and T. Collins, “Deep learning’s shallow gains: A comparative evaluation of algorithms for automatic music generation,” Mach. Learn., Vol.112, No.5, pp. 1785-1822, 2023. https://doi.org/10.1007/s10994-023-06309-w
  35. [35] S. Ji and X. Yang, “EmoMusicTV: Emotion-conditioned symbolic music generation with hierarchical transformer VAE,” IEEE Trans. Multimed., Vol.26, pp. 1076-1088, 2024. https://doi.org/10.1109/TMM.2023.3276177

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026