Research Paper:
LLM Prompt Optimization: From Zero-Shot to Automatic Role-Playing Generation Using Ontology
Asyafa Ditra Al Hauna*
, Siti Khomsah**
, Andi Prademon Yunus*,
, and Masanori Fukui***

*Informatics Engineering Study Program, Telkom University
Jl. DI Panjaitan No.128, Karangreja, Purwokerto Kidul, Kecamatan, Purwokerto Selatan, Kabupaten Banyumas, Jawa Tengah 53147, Indonesia
Corresponding author
**Data Science Study Program, Telkom University
Jl. DI Panjaitan No.128, Karangreja, Purwokerto Kidul, Kecamatan, Purwokerto Selatan, Kabupaten Banyumas, Jawa Tengah 53147, Indonesia
***Department of Information Engineering, Mie University
1577 Kurimamachiya-cho, Tsu, Mie 514-8507, Japan
The performance of an LLM application is highly dependent on how the user commands the application through prompts. Previous researchers invented role-play prompting techniques to overcome these problems. Including roles relevant to the task can improve the accuracy of the responses generated by LLMs. Nevertheless, the assignment of roles in this study remained a human-driven process. This study addresses the research gap in previous studies by developing deep learning models to effectively predict relevant roles based on tasks. In addition, the impact of the ontology data as an enrichment of the training data on the model performance in two scenarios (with and without ontology) was examined. In the with-ontology setup, enrichment data are extracted from an occupational ontology that includes roles, skills, and abilities along with their corresponding definitions. Data that were restricted to four predefined occupations were incorporated into the benchmark datasets. The benchmark datasets consisted of questions used to assess the LLM performance aligned with predetermined occupations. In without-ontology settings, the models were trained using only the benchmark datasets. An experiment on three variants of recurrent-based and a graph-based model in both scenarios showed that these models exhibited stable convergence and performed best on ontology-enriched data. Notably, the recurrent neural network (RNN) exhibited significant performance gains, achieving improvements of 34% and 1.86% in the F1 score and receiver operating characteristic-area under the curve (ROC–AUC), respectively. Despite its great performance—92% F1 score and 97% ROC–AUC, the gated recurrent unit is computationally the most expensive model compared to long short-term memory, RNN, and graph convolutional network. Further research is required to explore alternative model architectures and expand the scope of their predictable roles.
Automatic role-play prompt generation with ontology
- [1] A. Bhargava, C. Witkowski, S.-Z. Looi, and M. Thomson, “What’s the Magic Word? A Control Theory of LLM Prompting,” arXiv preprint, arXiv:2310.04444, 2023. https://doi.org/10.48550/arXiv.2310.04444
- [2] Y. Zhou, A. I. Muresanu, Z. Han, K. Paster, S. Pitis, H. Chan, and J. Ba, “Large Language Models Are Human-Level Prompt Engineers,” 11th Int. Conf. on Learning Representations, 2022. https://doi.org/10.48550/arXiv.2211.01910
- [3] H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al., “LLaMA: Open and Efficient Foundation Language Models,” arXiv preprint, arXiv:2302.13971, 2023.
- [4] S. K. Singh, S. Kumar, and P. S. Mehra, “Chat GPT & Google Bard AI: A Review,” 2023 Int. Conf. on IoT, Communication and Automation Technology (ICICAT), 2023. https://doi.org/10.1109/ICICAT57735.2023.10263706
- [5] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al., “GPT-4 Technical Report,” arXiv preprint, arXiv:2303.08774, 2023. https://doi.org/10.48550/arXiv.2303.08774
- [6] S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann, “BloombergGPT: A Large Language Model for Finance,” arXiv preprint, arXiv:2303.17564, 2023. https://doi.org/10.48550/arXiv.2303.17564
- [7] K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis et al., “Toward expert-level medical question answering with large language models,” Nature Medicine, Vol.31, No.3, pp. 943-950, 2025. https://doi.org/10.1038/s41591-024-03423-7
- [8] X. Wang, C. Li, Z. Wang, F. Bai, H. Luo, J. Zhang, N. Jojic, E. P. Xing, and Z. Hu, “PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization,” arXiv preprint, arXiv:2310.16427, 2023. https://doi.org/10.48550/arXiv.2310.16427
- [9] A. Kong, S. Zhao, H. Chen, Q. Li, Y. Qin, R. Sun, X. Zhou, E. Wang, and X. Dong, “Better Zero-Shot Reasoning with Role-Play Prompting,” Proc. of the 2024 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol.1: Long Papers), pp. 4099-4113, 2024. https://doi.org/10.18653/v1/2024.naacl-long.228
- [10] J. Beverley, S. Smith, M. A. Diller, W. D. Duncan, J. Zheng, J. W. Judkins, W. R. Hogan, R. McGill, D. Dooley, and Y. He, “The Occupational Ontology (OccO): Building a Bridge between Global Occupational Standards,” OSS2023: Ontologies for Services and Society, 9th Joint Ontology Workshops (JOWO 2023), 2023. https://doi.org/10.2139/ssrn.5345722
- [11] L. Zhang, B. Li, K. K. Thekumparampil, S. Oh, and N. He, “DPZero: Private Fine-Tuning of Language Models without Backpropagation,” arXiv preprint, arXiv:2310.09639, 2023. https://doi.org/10.48550/arXiv.2310.09639
- [12] J. Cheng, X. Liu, K. Zheng, P. Ke, H. Wang, Y. Dong, J. Tang, and M. Huang, “Black-Box Prompt Optimization: Aligning Large Language Models without Model Training,” Proc. of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol.1: Long Papers), pp. 3201-3219, 2024. https://doi.org/10.18653/v1/2024.acl-long.176
- [13] R. Wang, F. Mi, Y. Chen, B. Xue, H. Wang, Q. Zhu, K.-F. Wong, and R. Xu, “Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models,” Findings of the Association for Computational Linguistics: NAACL 2024, pp. 2243-2255, 2024. https://doi.org/10.18653/v1/2024.findings-naacl.145
- [14] Z. Liu, C. Gan, J. Wang, Y. Zhang, Z. Bo, M. Sun, H. Chen, and W. Zhang, “OntoTune: Ontology-Driven Self-training for Aligning Large Language Models,” Proc. of the ACM on Web Conf. 2025, pp. 119-133, 2025. https://doi.org/10.1145/3696410.3714816
- [15] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All You Need,” Advances in Neural Information Processing Systems, Vol.30, 2017.
- [16] F. D. Keles, P. M. Wijewardena, and C. Hegde, “On The Computational Complexity of Self-Attention,” Proc. of the 34th Int. Conf. on Algorithmic Learning Theory, Vol.201, pp. 597-619, 2023.
- [17] B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski et al., “RWKV: Reinventing RNNs for the Transformer Era,” Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 14048-14077, 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.936
- [18] B. Peng, D. Goldstein, Q. Anthony, A. Albalak, E. Alcaide, S. Biderman, E. Cheah, X. Du, T. Ferdinan, H. Hou et al., “Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence,” arXiv preprint, arXiv:2404.05892, 2024. https://doi.org/10.48550/arXiv.2404.05892
- [19] T. N. Kipf and M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks,” arXiv preprint, arXiv:1609.02907, 2016. https://doi.org/10.48550/arXiv.1609.02907
- [20] A. Hikmah, S. Adi, and M. Sulistiyono, “The Best Parameter Tuning on RNN Layers for Indonesian Text Classification,” 2020 3rd Int. Seminar on Research of Information Technology and Intelligent Systems (ISRITI), pp. 94-99, 2020. https://doi.org/10.1109/ISRITI51436.2020.9315425
- [21] X. Chen, K. Hirota, Y. Dai, and Z. Jia, “Estimation of SOC Based on LSTM-RNN and Design of Intelligent Equalization Charging System,” J. Adv. Comput. Intell. Intell. Inform., Vol.24, No.7, pp. 855-863, 2020. https://doi.org/10.20965/jaciii.2020.p0855
- [22] S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation, Vol.9, Issue 8, pp. 1735-1780, 1997. https://doi.org/10.1162/neco.1997.9.8.1735
- [23] H. Asrawi, A. Sunyoto, and B. Setiaji, “Implementation of Bidirectional Gated Recurrent Units for Text Classification,” 2023 6th Int. Conf. on Information and Communications Technology (ICOIACT), pp. 464-469, 2023. https://doi.org/10.1109/ICOIACT59844.2023.10455822
- [24] W. Gu, S. Zheng, R. Wang, and C. Dong, “Forecasting Realized Volatility Based on Sentiment Index and GRU Model,” J. Adv. Comput. Intell. Intell. Inform., Vol.24, No.3, pp. 299-306, 2020. https://doi.org/10.20965/jaciii.2020.p0299
- [25] S. Roy and D. Roth, “Solving General Arithmetic Word Problems,” Proc. of the 2015 Conf. on Empirical Methods in Natural Language Processing, pp. 1743-1752, 2015. https://doi.org/10.18653/v1/D15-1202
- [26] M. Goswami, V. Sanil, A. Choudhry, A. Srinivasan, C. Udompanyawit, and A. Dubrawski, “AQuA: A Benchmarking Tool for Label Quality Assessment,” Advances in Neural Information Processing Systems, Vol.36, pp. 79792-79807, 2023. https://doi.org/10.52202/075280-3494
- [27] M. J. Hosseini, H. Hajishirzi, O. Etzioni, and N. Kushman, “Learning to Solve Arithmetic Word Problems with Verb Categorization,” Proc. of the 2014 Conf. on Empirical Methods in Natural Language Processing (EMNLP), pp. 523-533, 2014. https://doi.org/10.3115/v1/D14-1058
- [28] A. Srivastava, A. Rastogi, A. Rao, A. A. Md Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso et al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” Trans. on Machine Learning Research, 2023.
- [29] M. Geva, D. Khashabi, E. Segal, T. Khot, D. Roth, and J. Berant, “Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies,” Trans. of the Association for Computational Linguistics, Vol.9, pp. 346-361, 2021. https://doi.org/10.1162/tacl_a_00370
- [30] A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,” Proc. of the 2019 Conf. of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Vol.1 (Long and Short Papers), pp. 4149-4158, 2019. https://doi.org/10.18653/v1/N19-1421
- [31] D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits, “What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Medical Exams,” Applied Sciences, Vol.11, Issue 14, Article No.6421, 2021. https://doi.org/10.3390/app11146421
- [32] J. Pennington, R. Socher, and C. Manning, “GloVe: Global Vectors for Word Representation,” Proc. of the 2014 Conf. on Empirical Methods in Natural Language Processing (EMNLP), pp. 1532-1543, 2014. https://doi.org/10.3115/v1/D14-1162
- [33] F. M. Shiri, T. Perumal, N. Mustapha, and R. Mohamed, “A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU,” arXiv preprint, arXiv:2305.17473, 2023. https://doi.org/10.48550/arXiv.2305.17473
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.