single-jc.php

JACIII Vol.30 No.4 pp. 1093-1101
(2026)

Research Paper:

Mixture of Experts Approach for Domain Adaptation in Finance

Masahiro Suzuki* ORCID Icon, Hiroki Sakaji** ORCID Icon, and Kiyoshi Izumi* ORCID Icon

*The University of Tokyo
7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan

**Hokkaido University
Kita 14, Nishi 9, Kita-ku, Sapporo, Hokkaido 060-0814, Japan

Received:
June 29, 2025
Accepted:
February 11, 2026
Published:
July 20, 2026
Keywords:
Mixture of Experts, domain adaptation, financial text mining, pre-trained language model
Abstract

Pre-trained language models (PLMs) have demonstrated high performance across various tasks and domains. Among these PLMs, Mixture of Experts (MoE) models also exhibit high performance with fewer active parameters. In domain adaptation, generally, continual pre-training is performed with existing models using domain-specific corpora. However, few efforts have been made to transform and train these models into models with MoE architectures. We propose a method to construct domain-adapted MoE models from general pre-trained models that do not initially have MoE architectures. By independently training multiple experts using domain corpora and integrating them into an MoE architecture, we constructed a domain-adapted MoE model. We performed this MoE transformation in the financial domain and verified its effectiveness in financial tasks. The evaluation results indicate that our domain-adapted MoE models perform better than those without MoE architectures. Our domain-adapted MoE models achieved consistent improvements over domain-adaptive pretraining, task-adaptive pretraining, and domain- and task-adaptive pretraining, with an average improvement of approximately 0.08 in F1 across six financial benchmark tasks for encoder-based models.

Two-stage MoE domain adaptation

Two-stage MoE domain adaptation

Cite this article as:
M. Suzuki, H. Sakaji, and K. Izumi, “Mixture of Experts Approach for Domain Adaptation in Finance,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 1093-1101, 2026.
Data files:
References
  1. [1] OpenAI, “GPT-4 technical report,” arXiv:2303.08774, 2024. https://doi.org/10.48550/arXiv.2303.08774
  2. [2] H. Touvron et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288, 2023. https://doi.org/10.48550/arXiv.2307.09288
  3. [3] X. Li et al., “Are ChatGPT and GPT-4 general-purpose solvers for financial text analytics? A study on several typical tasks,” Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing: Industry Track, pp. 408-422, 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.39
  4. [4] S. Gururangan et al., “Don’t stop pretraining: Adapt language models to domains and tasks,” Proc. of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8342-8360, 2020. https://doi.org/10.18653/v1/2020.acl-main.740
  5. [5] K. Singhal et al., “Large language models encode clinical knowledge,” Nature, Vol.620, No.7972, pp. 172-180, 2023. https://doi.org/10.1038/s41586-023-06291-2
  6. [6] Y. Xie, K. Aggarwal, and A. Ahmad, “Efficient continual pre-training for building domain specific large language models,” Findings of the Association for Computational Linguistics: ACL 2024, pp. 10184-10201, 2024. https://doi.org/10.18653/v1/2024.findings-acl.606
  7. [7] D. Araci, “FinBERT: Financial sentiment analysis with pre-trained language models,” arXiv:1908.10063, 2019. https://doi.org/10.48550/arXiv.1908.10063
  8. [8] R. Shah et al., “When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain,” Proc. of the 2022 Conf. on Empirical Methods in Natural Language Processing, pp. 2322-2335, 2022. https://doi.org/10.18653/v1/2022.emnlp-main.148
  9. [9] S. Wu et al., “BloombergGPT: A large language model for finance,” arXiv:2303.17564, 2023. https://doi.org/10.48550/arXiv.2303.17564
  10. [10] N. Pahari and K. Shimada, “Layer configurations of BERT for multitask learning and data augmentation,” J. Adv. Comput. Intell. Intell. Inform., Vol.28, No.1, pp. 29-40, 2024. https://doi.org/10.20965/jaciii.2024.p0029
  11. [11] S. O. Khairunnisa, Z. Chen, and M. Komachi, “Improving domain-specific NER in the Indonesian language through domain transfer and data augmentation,” J. Adv. Comput. Intell. Intell. Inform., Vol.28, No.6, pp. 1299-1312, 2024. https://doi.org/10.20965/jaciii.2024.p1299
  12. [12] N. Shazeer et al., “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” 5th Int. Conf. on Learning Representations, 2017. https://openreview.net/forum?id=B1ckMDqlg [Accessed June 27, 2026]
  13. [13] A. Q. Jiang et al., “Mixtral of experts,” arXiv:2401.04088, 2024. https://doi.org/10.48550/arXiv.2401.04088
  14. [14] W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” J. of Machine Learning Research, Vol.23, Article No.120, 2022.
  15. [15] M. Suzuki, H. Sakaji, M. Hirano, and K. Izumi, “FinDeBERTaV2: Word-segmentation-free pre-trained language model for finance,” Trans. of the Japanese Society for Artificial Intelligence, Vol.39, No.4, pp. FIN23-G_1-14, 2024 (in Japanese). https://doi.org/10.1527/tjsai.39-4_FIN23-G
  16. [16] D. Lepikhin et al., “GShard: Scaling giant models with conditional computation and automatic sharding,” 9th Int. Conf. on Learning Representations, 2021. https://openreview.net/forum?id=qrwe7XHTmYb [Accessed June 27, 2026]
  17. [17] Y. Zhou et al., “Mixture-of-experts with expert choice routing,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 7103-7114, 2022.
  18. [18] J. Lee-Thorp and J. Ainslie, “Sparse mixers: Combining MoE and mixing to build a more efficient BERT,” Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 58-75, 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.5
  19. [19] N. Du et al., “GLaM: Efficient scaling of language models with mixture-of-experts,” Proc. of the 39th Int. Conf. on Machine Learning, pp. 5547-5569, 2022.
  20. [20] S. Diao, T. Xu, R. Xu, J. Wang, and T. Zhang, “Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models’ memories,” Proc. of the 61st Annual Meeting of the Association for Computational Linguistics (Vol.1: Long Papers), pp. 5113-5129, 2023. https://doi.org/10.18653/v1/2023.acl-long.280
  21. [21] M. Suzuki, H. Sakaji, M. Hirano, and K. Izumi, “Constructing and analyzing domain-specific language model for financial text mining,” Information Processing & Management, Vol.60, No.2, Article No.103194, 2023. https://doi.org/https://doi.org/10.1016/j.ipm.2022.103194
  22. [22] Z. Liu, D. Huang, K. Huang, Z. Li, and J. Zhao, “FinBERT: A pre-trained financial language representation model for financial text mining,” Proc. of the 29th Int. Joint Conf. on Artificial Intelligence, pp. 4513-4519, 2020. https://doi.org/10.24963/ijcai.2020/622
  23. [23] J. Hoffmann et al., “Training compute-optimal large language models,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 30016-30030, 2022.
  24. [24] BigScience Workshop (T. L. Scao et al.), “BLOOM: A 176B-parameter open-access multilingual language model,” arXiv:2211.05100, 2023. https://doi.org/10.48550/arXiv.2211.05100
  25. [25] S. Biderman et al., “Pythia: A suite for analyzing large language models across training and scaling,” Proc. of the 40th Int. Conf. on Machine Learning, pp. 2397-2430, 2023.
  26. [26] J. Wei et al., “Finetuned language models are zero-shot learners,” 10th Int. Conf. on Learning Representations, 2022. https://openreview.net/forum?id=gEZrGCozdqR [Accessed June 27, 2026]
  27. [27] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv:1707.06347, 2017. https://doi.org/10.48550/arXiv.1707.06347
  28. [28] Y. Bai et al., “Training a helpful and harmless assistant with reinforcement learning from human feedback,” arXiv:2204.05862, 2022. https://doi.org/10.48550/arXiv.2204.05862
  29. [29] L. Ouyang et al., “Training language models to follow instructions with human feedback,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 27730-27744, 2022.
  30. [30] Q. Xie et al., “PIXIU: A comprehensive benchmark, instruction dataset and large language model for finance,” 37th Conf. on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. https://openreview.net/forum?id=vTrRq6vCQH [Accessed June 27, 2026]
  31. [31] H. Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv:2302.13971, 2023. https://doi.org/10.48550/arXiv.2302.13971
  32. [32] P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing,” The 11th Int. Conf. on Learning Representations, 2023. https://openreview.net/forum?id=sE7-XhLxHA [Accessed June 27, 2026]
  33. [33] C. Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. of Machine Learning Research, Vol.21, Article No.140, 2020.
  34. [34] AI@Meta, “Llama 3 model card,” 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md [Accessed June 27, 2026]
  35. [35] T. Wolf et al., “Transformers: State-of-the-art natural language processing,” Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38-45, 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6
  36. [36] J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters,” Proc. of the 26th ACM SIGKDD Int. Conf. on Knowledge Discovery & Data Mining, pp. 3505-3506, 2020. https://doi.org/10.1145/3394486.3406703
  37. [37] P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala, “Good debt or bad debt: Detecting semantic orientations in economic texts,” J. of the Association for Information Science and Technology, Vol.65, No.4, pp. 782-796, 2014. https://doi.org/10.1002/asi.23062
  38. [38] M. Maia et al., “WWW’18 Open Challenge: Financial opinion mining and question answering,” Companion Proc. of the Web Conf. 2018, pp. 1941-1942, 2018. https://doi.org/10.1145/3184558.3192301
  39. [39] A. Sinha and T. Khandait, “Impact of news on the commodity market: Dataset and results,” Advances in Information and Communication: Proc. of the 2021 Future of Information and Communication Conf., Vol.2, pp. 589-601, 2021. https://doi.org/10.1007/978-3-030-73103-8_41
  40. [40] A. Shah, S. Paturi, and S. Chava, “Trillion dollar words: A new financial dataset, task & market analysis,” Proc. of the 61st Annual Meeting of the Association for Computational Linguistics (Vol.1: Long Papers), pp. 6664-6679, 2023. https://doi.org/10.18653/v1/2023.acl-long.368
  41. [41] J. C. S. Alvarado, K. Verspoor, and T. Baldwin, “Domain adaption of named entity recognition to support credit risk assessment,” Proc. of the Australasian Language Technology Association Workshop 2015, pp. 84-90, 2015.
  42. [42] A. Shah, A. Gullapalli, R. Vithani, M. Galarnyk, and S. Chava, “FiNER: Financial named entity recognition dataset and weak-supervision model,” arXiv:2302.11157v1, 2023. https://doi.org/10.48550/arXiv.2302.11157
  43. [43] Q. Xie et al., “FinBen: An holistic financial benchmark for large language models,” arXiv:2402.12659, 2024. https://doi.org/10.48550/arXiv.2402.12659

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026