Research Paper:
Category-Centric Initialization for Knowledge Graph Embedding
Runyu Ni
, Hiroki Shibata
, and Yasufumi Takama

School of Systems Design, Tokyo Metropolitan University
6-6 Ashigaoka, Hino, Tokyo 191-0065, Japan
This paper proposes a category-centric initialization that introduces prior knowledge for knowledge graph embedding (KGE) at a low cost. KGE is a technology that maps symbols to embeddings to utilize large-scale knowledge graphs, and it has been widely applied because of its simplicity and efficiency. However, the initialization challenge with this technology has long been overlooked. This critical issue has implications for the training cost, the stability of training, and even the final performance of the model. KGE predominantly utilizes random initialization, which overlooks the wealth of prior knowledge embedded within knowledge graphs. To counteract this, pre-training initialization has been introduced as a way to utilize the prior knowledge. While this strategy can lead to enhanced model performance and quicker convergence rates, it increases computational demands and restricts application breadth. To address these challenges, we propose a novel initialization called category-centric initialization (CCI). CCI is designed to be universally applicable across any scenario involving the training of KGE models from scratch. It utilizes the weighted sum of the category embedding and the random embedding as the initial embedding of entities. By integrating explicit category information into the random initialization, CCI effectively utilizes prior knowledge while avoiding excessive computational cost. The results of experiments demonstrate that the proposed method can effectively reduce the training cost of advanced KGE models without degrading the final performance. Additionally, the results of experiments without category information show that our method can be applied in scenarios where explicit categories are not given to entities.
Category-centric KGE initialization
- [1] Q. Wang, Z. Mao, B. Wang, and L. Guo, “Knowledge graph embedding: A survey of approaches and applications,” IEEE Trans. Knowl. Data Eng., Vol.29, No.12, pp. 2724-2743, 2017. https://doi.org/10.1109/TKDE.2017.2754499
- [2] A. Rossi, D. Barbosa, D. Firmani, A. Matinata, and P. Merialdo, “Knowledge graph embedding for link prediction: A comparative analysis,” ACM Trans. Knowl. Discov. Data, Vol.15, No.2, Article No.14, 2021. https://doi.org/10.1145/3424672
- [3] X. Ge, Y. C. Wang, B. Wang, and C.-C. J. Kuo, “Knowledge graph embedding: An overview,” APSIPA Trans. Signal Inf. Process., Vol.13, No.1, 2024. https://doi.org/10.1561/116.00000065
- [4] A. Bordes, N. Usunier, A. Garcia-Durán, J. Weston, and O. Yakhnenko, “Translating embeddings for modeling multi-relational data,” Proc. 27th Int. Conf. Neural Inf. Process. Syst. (NIPS’13), Vol.2, pp. 2787-2795, 2013.
- [5] Y. Cao, X. Wang, X. He, Z. Hu, and T.-S. Chua, “Unifying knowledge graph learning and recommendation: Towards a better understanding of user preferences,” Proc. World Wide Web Conf. (WWW’19), pp. 151-161, 2019. https://doi.org/10.1145/3308558.3313705
- [6] X. Wang, X. He, Y. Cao, M. Liu, and T.-S. Chua, “KGAT: Knowledge graph attention network for recommendation,” Proc. 25th ACM SIGKDD Int. Conf. Knowl. Discov. Data (KDD’19), pp. 950-958, 2019. https://doi.org/10.1145/3292500.3330989
- [7] Q. Ai, V. Azizi, X. Chen, and Y. Zhang, “Learning heterogeneous knowledge base embeddings for explainable recommendation,” Algorithms, Vol.11, No.9, Article No.137, 2018. https://doi.org/10.3390/a11090137
- [8] Q. Lu, W. Du, W. Xu, and J. Ma, “KSGAN: Knowledge-aware subgraph attention network for scholarly community recommendation,” Inf. Syst., Vol.119, Article No.102282, 2023. https://doi.org/10.1016/j.is.2023.102282
- [9] S. K. Mohamed, A. Nounu, and V. Nováček, “Biological applications of knowledge graph embedding models,” Brief. Bioinform., Vol.22, No.2, pp. 1679-1693, 2021. https://doi.org/10.1093/bib/bbaa012
- [10] M. Zitnik, M. Agrawal, and J. Leskovec, “Modeling polypharmacy side effects with graph convolutional networks,” Bioinformatics, Vol.34, No.13, pp. i457-i466, 2018. https://doi.org/10.1093/bioinformatics/bty294
- [11] C. Xiong, R. Power, and J. Callan, “Explicit semantic ranking for academic search via knowledge graph embedding,” Proc. 26th Int. Conf. World Wide Web (WWW’17), pp. 1271-1279, 2017. https://doi.org/10.1145/3038912.3052558
- [12] M. N. Alam and M. M. Ali, “Loan default risk prediction using knowledge graph,” 14th Int. Conf. Knowl. Smart Technol. (KST), pp. 34-39, 2022. https://doi.org/10.1109/KST53302.2022.9729073
- [13] Y. Liu, Q. Zeng, H. Yang, and A. Carrio, “Stock price movement prediction from financial news with deep learning and knowledge graph embedding,” Proc. 15th Int. Workshop Knowl. Manag. Acquis. Intell. Syst., pp. 102-113, 2018. https://doi.org/10.1007/978-3-319-97289-3_8
- [14] J. Deng et al., “A unified model for video understanding and knowledge embedding with heterogeneous knowledge graph dataset,” Proc. 2023 ACM Int. Conf. Multimed. Retr., pp. 95-104, 2023. https://doi.org/10.1145/3591106.3592258
- [15] H. Ding et al., “Genre classification empowered by knowledge-embedded music representation,” IEEE/ACM Trans. Audio Speech Lang. Process., Vol.32, pp. 2764-2776, 2024. https://doi.org/10.1109/TASLP.2024.3402115
- [16] B. Yang, W.-T. Yih, X. He, J. Gao, and L. Deng, “Embedding entities and relations for learning and inference in knowledge bases,” 3rd Int. Conf. Learn. Represent., 2015.
- [17] T. Trouillon, J. Welbl, S. Riedel, É. Gaussier, and G. Bouchard, “Complex embeddings for simple link prediction,” Proc. 33rd Int. Conf. Mach. Learn., pp. 2071-2080, 2016.
- [18] Z. Wang, J. Zhang, J. Feng, and Z. Chen, “Knowledge graph embedding by translating on hyperplanes,” Proc. AAAI Conf. Artif. Intell., Vol.28, No.1, pp. 1112-1119, 2014. https://doi.org/10.1609/aaai.v28i1.8870
- [19] Z. Sun, Z.-H. Deng, J.-Y. Nie, and J. Tang, “RotatE: Knowledge graph embedding by relational rotation in complex space,” 7th Int. Conf. Learn. Represent., 2019.
- [20] D. Krompaß, S. Baier, and V. Tresp, “Type-constrained representation learning in knowledge graphs,” Proc. 14th Int. Semant. Web Conf., pp. 640-655, 2015. https://doi.org/10.1007/978-3-319-25007-6_37
- [21] G. Niu, B. Li, Y. Zhang, and S. Pu, “CAKE: A scalable commonsense-aware framework for multi-view knowledge graph completion,” Proc. 60th Annu. Meet. Assoc. Comput. Linguist. (Vol.1: Long Pap.), pp. 2867-2877, 2022. https://doi.org/10.18653/v1/2022.acl-long.205
- [22] W. Hu et al., “OGB-LSC: A large-scale challenge for machine learning on graphs,” arXiv:2103.09430, 2021. https://doi.org/10.48550/arXiv.2103.09430
- [23] P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov, “Enriching word vectors with subword information,” Trans. Assoc. Comput. Linguist., Vol.5, pp. 135-146, 2017. https://doi.org/10.1162/tacl_a_00051
- [24] M. Galkin, E. Denis, J. Wu, and W. L. Hamilton, “NodePiece: Compositional and parameter-efficient representations of large knowledge graphs,” 10th Int. Conf. Learn. Represent., 2022.
- [25] H. Wang, Y. Wang, D. Lian, and J. Gao, “A lightweight knowledge graph embedding framework for efficient inference and storage,” Proc. 30th ACM Int. Conf. Inf. Knowl. Manag., pp. 1909-1918, 2021. https://doi.org/10.1145/3459637.3482224
- [26] M. Chen et al., “Entity-agnostic representation learning for parameter-efficient knowledge graph embedding,” Proc. AAAI Conf. Artif. Intell., Vol.37, No.4, pp. 4182-4190, 2023. https://doi.org/10.1609/aaai.v37i4.25535
- [27] M. V. Narkhede, P. P. Bartakke, and M. S. Sutaone, “A review on weight initialization strategies for neural networks,” Artif. Intell. Rev., Vol.55, No.1, pp. 291-322, 2022. https://doi.org/10.1007/s10462-021-10033-z
- [28] D. Ruffinelli, S. Broscheit, and R. Gemulla, “You CAN teach an old dog new tricks! On training knowledge graph embeddings,” 8th Int. Conf. Learn. Represent., 2020.
- [29] M. Ali et al., “Bringing light into the dark: A large-scale evaluation of knowledge graph embedding models under a unified framework,” IEEE Trans. Pattern Anal. Mach. Intell., Vol.44, No.12, pp. 8825-8845, 2022. https://doi.org/10.1109/TPAMI.2021.3124805
- [30] X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” Proc. of the 30th Int. Conf. on Artif. Intell. Stat., pp. 249-256, 2010.
- [31] A. M. Saxe, J. L. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” 2nd Int. Conf. Learn. Represent., 2014.
- [32] C. Chung and J. J. Whang, “Knowledge graph embedding via metagraph learning,” Proc. 44th Int. ACM SIGIR Conf. Res. Dev. Inf. Retr., pp. 2212-2216, 2021. https://doi.org/10.1145/3404835.3463072
- [33] V. Kocijan and T. Lukasiewicz, “Knowledge base completion meets transfer learning,” Proc. 2021 Conf. Empir. Methods Nat. Lang. Process., pp. 6521-6533, 2021. https://doi.org/10.18653/v1/2021.emnlp-main.524
- [34] L. Yao, C. Mao, and Y. Luo, “KG-BERT: BERT for knowledge graph completion,” arXiv:1909.03193, 2019. https://doi.org/10.48550/arXiv.1909.03193
- [35] H. Chen, B. Perozzi, Y. Hu, and S. Skiena, “HARP: Hierarchical representation learning for networks,” Proc. AAAI Conf. Artif. Intell., Vol.32, No.1, pp. 2127-2134, 2018. https://doi.org/10.1609/aaai.v32i1.11849
- [36] W. Lin, F. He, F. Zhang, X. Cheng, and H. Cai, “Initialization for network embedding: A graph partition approach,” Proc. 13th Int. Conf. Web Search Data Min., pp. 367-374, 2020. https://doi.org/10.1145/3336191.3371781
- [37] C. Moon, P. Jones, and N. F. Samatova, “Learning entity type embeddings for knowledge graph completion,” Proc. 2017 ACM Conf. Inf. Knowl. Manag., pp. 2215-2218, 2017. https://doi.org/10.1145/3132847.3133095
- [38] B. Y. Weisfeiler and A. A. Leman, “The reduction of a graph to canonical form and the algebra which appears therein,” G. Ryabov (Trans.), Nauchno-Technicheskaya Informatsia, Vol.2, No.9, pp. 12-16, 1968.
- [39] M. Ali et al., “PyKEEN 1.0: A Python library for training and evaluating knowledge graph embeddings,” J. Mach. Learn. Res., Vol.22, No.82, pp. 1-6, 2021.
- [40] T. Lacroix, N. Usunier, and G. Obozinski, “Canonical tensor decomposition for knowledge base completion,” Proc. 35th Int. Conf. Mach. Learn., pp. 2863-2872, 2018.
- [41] S. Friedland and L.-H. Lim, “Nuclear norm of higher-order tensors,” Mathematics of Computation, Vol.87, No.311, pp. 1255-1281, 2018.
- [42] R. Socher, D. Chen, C. D. Manning, and A. Y. Ng, “Reasoning with neural tensor networks for knowledge base completion,” Proc. 27th Int. Conf. Neural Inf. Process. Syst., pp. 926-934, 2013.
- [43] Y. Zhao, A. Zhang, R. Xie, K. Liu, and X. Wang, “Connecting embeddings for knowledge graph entity typing,” Proc. 58th Annu. Meet. Assoc. Comput. Linguist., pp. 6419-6428, 2020. https://doi.org/10.18653/v1/2020.acl-main.572
- [44] W. Pan, W. Wei, and X.-L. Mao, “Context-aware entity typing in knowledge graphs,” Find. Assoc. Comput. Linguist. EMNLP 2021, pp. 2240-2250, 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.193
- [45] X. Ge, Y.-C. Wang, B. Wang, and C. C. J. Kuo, “CORE: A knowledge graph entity type prediction method via complex space regression and embedding,” Pattern Recognit. Lett., Vol.157, pp. 97-103, 2022. https://doi.org/10.1016/j.patrec.2022.03.024
- [46] X. Ge, Y. C. Wang, B. Wang, and C.-C. J. Kuo, “TypeEA: Type-associated embedding for knowledge graph entity alignment,” APSIPA Trans. Signal Inf. Process., Vol.12, No.1, 2023. https://doi.org/10.1561/116.00000139
- [47] Y.-C. Wang, X. Ge, B. Wang, and C.-C. J. Kuo, “AsyncET: Asynchronous representation learning for knowledge graph entity typing,” Proc. 30th ACM SIGKDD Conf. Knowl. Discov. Data Min., pp. 3267-3276, 2024. https://doi.org/10.1145/3637528.3671832
- [48] J. Weston, A. Bordes, O. Yakhnenko, and N. Usunier, “Connecting language and knowledge bases with embedding models for relation extraction,” Proc. 2013 Conf. Empir. Methods Nat. Lang. Process., pp. 1366-1371, 2013.
- [49] X. Li, A. Henriksson, M. Duneld, J. Nouri, and Y. Wu, “Evaluating embeddings from pre-trained language models and knowledge graphs for educational content recommendation,” Future Internet, Vol.16, No.1, Article No.12, 2024. https://doi.org/10.3390/fi16010012
- [50] S. K. Mohamed, A. Nounu, and V. Nováček, “Drug target discovery using knowledge graph embeddings,” Proc. 34th ACM/SIGAPP Symp. Appl. Comput., pp. 11-18, 2019. https://doi.org/10.1145/3297280.3297282
- [51] S. Bonner et al., “Understanding the performance of knowledge graph embeddings in drug discovery,” Artif. Intell. Life Sci., Vol.2, Article No.100036, 2022. https://doi.org/10.1016/j.ailsci.2022.100036
- [52] L. Luo, “Combining knowledge graph and artificial intelligence to conduct financial report quality detection research,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.4, pp. 787-795, 2025. https://doi.org/10.20965/jaciii.2025.p0787
- [53] Z. Wang et al., “Federated recommendation with explicitly encoding item bias,” Proc. AAAI Conf. Artif. Intell., Vol.39, No.12, pp. 12792-12800, 2025. https://doi.org/10.1609/aaai.v39i12.33395
- [54] J. Zhu et al., “BGCL: Bi-subgraph network based on graph contrastive learning for cold-start QoS prediction,” Knowl.-Based Syst., Vol.263, Article No.110296, 2023. https://doi.org/10.1016/j.knosys.2023.110296
- [55] S. Xie et al., “PBScaler: A bottleneck-aware autoscaling framework for microservice-based applications,” IEEE Trans. on Serv. Comput., Vol.17, No.2, pp. 604-616, 2024. https://doi.org/10.1109/TSC.2024.3376202
- [56] L. Liang et al., “KAG: Boosting LLMs in professional domains via knowledge augmented generation,” Companion Proc. ACM Web Conf. 2025, pp. 334-343, 2025. https://doi.org/10.1145/3701716.3715240
- [57] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation,” D. E. Rumelhart, J. L. McClelland, and the PDP Research Group (Eds.), “Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol.1, Foundations,” Chapter 8, pp. 318-362, The MIT Press, 1986.
- [58] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proc. 2019 Conf. N. Am. Chapter Assoc. Comput. Linguist. Hum. Lang. Technol. (NAACL-HLT), Vol.1, pp. 4171-4186, 2019. https://doi.org/10.18653/v1/N19-1423
- [59] G. Karypis and V. Kumar, “A fast and high quality multilevel scheme for partitioning irregular graphs,” SIAM J. Sci. Comput., Vol.20, No.1, pp. 359-392, 1998. https://doi.org/10.1137/S1064827595287997
- [60] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature, Vol.323, No.6088, pp. 533-536, 1986. https://doi.org/10.1038/323533a0
- [61] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” 3rd Int. Conf. Learn. Represent., 2015.
- [62] J. Hao, M. Chen, W. Yu, Y. Sun, and W. Wang, “Universal representation learning of knowledge bases by jointly embedding instances and ontological concepts,” Proc. 25th ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 1709-1719, 2019. https://doi.org/10.1145/3292500.3330838
- [63] W. Xiong, T. Hoang, and W. Y. Wang, “DeepPath: A reinforcement learning method for knowledge graph reasoning,” Proc. 2017 Conf. Empir. Methods Nat. Lang. Process., pp. 564-573, 2017. https://doi.org/10.18653/v1/D17-1060
- [64] A. Carlson et al., “Toward an architecture for never-ending language learning,” Proc. AAAI Conf. Artif. Intell., Vol.24, No.1, pp. 1306-1313, 2010. https://doi.org/10.1609/aaai.v24i1.7519
- [65] K. Toutanova and D. Chen, “Observed versus latent features for knowledge base and text inference,” Proc. 3rd Workshop Contin. Vector Space Models Their Compos., pp. 57-66, 2015. https://doi.org/10.18653/v1/W15-4007
- [66] N. Morgan and H. Bourlard, “Generalization and parameter estimation in feedforward nets: Some experiments,” Proc. 3rd Int. Conf. Neural Inf. Process. Syst., pp. 630-637, 1989.
- [67] Y. Bai et al., “Understanding and improving early stopping for learning with noisy labels,” Proc. 35th Int. Conf. Neural Inf. Process. Syst., pp. 24392-24403, 2021.
- [68] L. Prechelt, “Early stopping—But when?” G. Montavon, G. B. Orr, and K.-R. Müller, “Neural Networks: Tricks of the Trade (2nd edition),” pp. 53-67, Springer, 2012. https://doi.org/10.1007/978-3-642-35289-8_5
- [69] L. Prechelt, “Automatic early stopping using cross validation: Quantifying the criteria,” Neural Netw., Vol.11, No.4, pp. 761-767, 1998. https://doi.org/10.1016/S0893-6080(98)00010-0
- [70] T. Caliński and J. Harabasz, “A dendrite method for cluster analysis,” Commun. Stat., Vol.3, No.1, pp. 1-27, 1974. https://doi.org/10.1080/03610927408827101
- [71] A. Grover and J. Leskovec, “node2vec: Scalable feature learning for networks,” Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., pp. 855-864, 2016. https://doi.org/10.1145/2939672.2939754
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.