Research Paper:
Mixture of Experts Approach for Domain Adaptation in Finance
Masahiro Suzuki*
, Hiroki Sakaji**
, and Kiyoshi Izumi*

*The University of Tokyo
7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan
**Hokkaido University
Kita 14, Nishi 9, Kita-ku, Sapporo, Hokkaido 060-0814, Japan
Pre-trained language models (PLMs) have demonstrated high performance across various tasks and domains. Among these PLMs, Mixture of Experts (MoE) models also exhibit high performance with fewer active parameters. In domain adaptation, generally, continual pre-training is performed with existing models using domain-specific corpora. However, few efforts have been made to transform and train these models into models with MoE architectures. We propose a method to construct domain-adapted MoE models from general pre-trained models that do not initially have MoE architectures. By independently training multiple experts using domain corpora and integrating them into an MoE architecture, we constructed a domain-adapted MoE model. We performed this MoE transformation in the financial domain and verified its effectiveness in financial tasks. The evaluation results indicate that our domain-adapted MoE models perform better than those without MoE architectures. Our domain-adapted MoE models achieved consistent improvements over domain-adaptive pretraining, task-adaptive pretraining, and domain- and task-adaptive pretraining, with an average improvement of approximately 0.08 in F1 across six financial benchmark tasks for encoder-based models.
Two-stage MoE domain adaptation
- [1] OpenAI, “GPT-4 technical report,” arXiv:2303.08774, 2024. https://doi.org/10.48550/arXiv.2303.08774
- [2] H. Touvron et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv:2307.09288, 2023. https://doi.org/10.48550/arXiv.2307.09288
- [3] X. Li et al., “Are ChatGPT and GPT-4 general-purpose solvers for financial text analytics? A study on several typical tasks,” Proc. of the 2023 Conf. on Empirical Methods in Natural Language Processing: Industry Track, pp. 408-422, 2023. https://doi.org/10.18653/v1/2023.emnlp-industry.39
- [4] S. Gururangan et al., “Don’t stop pretraining: Adapt language models to domains and tasks,” Proc. of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 8342-8360, 2020. https://doi.org/10.18653/v1/2020.acl-main.740
- [5] K. Singhal et al., “Large language models encode clinical knowledge,” Nature, Vol.620, No.7972, pp. 172-180, 2023. https://doi.org/10.1038/s41586-023-06291-2
- [6] Y. Xie, K. Aggarwal, and A. Ahmad, “Efficient continual pre-training for building domain specific large language models,” Findings of the Association for Computational Linguistics: ACL 2024, pp. 10184-10201, 2024. https://doi.org/10.18653/v1/2024.findings-acl.606
- [7] D. Araci, “FinBERT: Financial sentiment analysis with pre-trained language models,” arXiv:1908.10063, 2019. https://doi.org/10.48550/arXiv.1908.10063
- [8] R. Shah et al., “When FLUE meets FLANG: Benchmarks and large pretrained language model for financial domain,” Proc. of the 2022 Conf. on Empirical Methods in Natural Language Processing, pp. 2322-2335, 2022. https://doi.org/10.18653/v1/2022.emnlp-main.148
- [9] S. Wu et al., “BloombergGPT: A large language model for finance,” arXiv:2303.17564, 2023. https://doi.org/10.48550/arXiv.2303.17564
- [10] N. Pahari and K. Shimada, “Layer configurations of BERT for multitask learning and data augmentation,” J. Adv. Comput. Intell. Intell. Inform., Vol.28, No.1, pp. 29-40, 2024. https://doi.org/10.20965/jaciii.2024.p0029
- [11] S. O. Khairunnisa, Z. Chen, and M. Komachi, “Improving domain-specific NER in the Indonesian language through domain transfer and data augmentation,” J. Adv. Comput. Intell. Intell. Inform., Vol.28, No.6, pp. 1299-1312, 2024. https://doi.org/10.20965/jaciii.2024.p1299
- [12] N. Shazeer et al., “Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,” 5th Int. Conf. on Learning Representations, 2017. https://openreview.net/forum?id=B1ckMDqlg [Accessed June 27, 2026]
- [13] A. Q. Jiang et al., “Mixtral of experts,” arXiv:2401.04088, 2024. https://doi.org/10.48550/arXiv.2401.04088
- [14] W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” J. of Machine Learning Research, Vol.23, Article No.120, 2022.
- [15] M. Suzuki, H. Sakaji, M. Hirano, and K. Izumi, “FinDeBERTaV2: Word-segmentation-free pre-trained language model for finance,” Trans. of the Japanese Society for Artificial Intelligence, Vol.39, No.4, pp. FIN23-G_1-14, 2024 (in Japanese). https://doi.org/10.1527/tjsai.39-4_FIN23-G
- [16] D. Lepikhin et al., “GShard: Scaling giant models with conditional computation and automatic sharding,” 9th Int. Conf. on Learning Representations, 2021. https://openreview.net/forum?id=qrwe7XHTmYb [Accessed June 27, 2026]
- [17] Y. Zhou et al., “Mixture-of-experts with expert choice routing,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 7103-7114, 2022.
- [18] J. Lee-Thorp and J. Ainslie, “Sparse mixers: Combining MoE and mixing to build a more efficient BERT,” Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 58-75, 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.5
- [19] N. Du et al., “GLaM: Efficient scaling of language models with mixture-of-experts,” Proc. of the 39th Int. Conf. on Machine Learning, pp. 5547-5569, 2022.
- [20] S. Diao, T. Xu, R. Xu, J. Wang, and T. Zhang, “Mixture-of-domain-adapters: Decoupling and injecting domain knowledge to pre-trained language models’ memories,” Proc. of the 61st Annual Meeting of the Association for Computational Linguistics (Vol.1: Long Papers), pp. 5113-5129, 2023. https://doi.org/10.18653/v1/2023.acl-long.280
- [21] M. Suzuki, H. Sakaji, M. Hirano, and K. Izumi, “Constructing and analyzing domain-specific language model for financial text mining,” Information Processing & Management, Vol.60, No.2, Article No.103194, 2023. https://doi.org/https://doi.org/10.1016/j.ipm.2022.103194
- [22] Z. Liu, D. Huang, K. Huang, Z. Li, and J. Zhao, “FinBERT: A pre-trained financial language representation model for financial text mining,” Proc. of the 29th Int. Joint Conf. on Artificial Intelligence, pp. 4513-4519, 2020. https://doi.org/10.24963/ijcai.2020/622
- [23] J. Hoffmann et al., “Training compute-optimal large language models,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 30016-30030, 2022.
- [24] BigScience Workshop (T. L. Scao et al.), “BLOOM: A 176B-parameter open-access multilingual language model,” arXiv:2211.05100, 2023. https://doi.org/10.48550/arXiv.2211.05100
- [25] S. Biderman et al., “Pythia: A suite for analyzing large language models across training and scaling,” Proc. of the 40th Int. Conf. on Machine Learning, pp. 2397-2430, 2023.
- [26] J. Wei et al., “Finetuned language models are zero-shot learners,” 10th Int. Conf. on Learning Representations, 2022. https://openreview.net/forum?id=gEZrGCozdqR [Accessed June 27, 2026]
- [27] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv:1707.06347, 2017. https://doi.org/10.48550/arXiv.1707.06347
- [28] Y. Bai et al., “Training a helpful and harmless assistant with reinforcement learning from human feedback,” arXiv:2204.05862, 2022. https://doi.org/10.48550/arXiv.2204.05862
- [29] L. Ouyang et al., “Training language models to follow instructions with human feedback,” Proc. of the 36th Int. Conf. on Neural Information Processing Systems, pp. 27730-27744, 2022.
- [30] Q. Xie et al., “PIXIU: A comprehensive benchmark, instruction dataset and large language model for finance,” 37th Conf. on Neural Information Processing Systems Datasets and Benchmarks Track, 2023. https://openreview.net/forum?id=vTrRq6vCQH [Accessed June 27, 2026]
- [31] H. Touvron et al., “LLaMA: Open and efficient foundation language models,” arXiv:2302.13971, 2023. https://doi.org/10.48550/arXiv.2302.13971
- [32] P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing,” The 11th Int. Conf. on Learning Representations, 2023. https://openreview.net/forum?id=sE7-XhLxHA [Accessed June 27, 2026]
- [33] C. Raffel et al., “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. of Machine Learning Research, Vol.21, Article No.140, 2020.
- [34] AI@Meta, “Llama 3 model card,” 2024. https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md [Accessed June 27, 2026]
- [35] T. Wolf et al., “Transformers: State-of-the-art natural language processing,” Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38-45, 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6
- [36] J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters,” Proc. of the 26th ACM SIGKDD Int. Conf. on Knowledge Discovery & Data Mining, pp. 3505-3506, 2020. https://doi.org/10.1145/3394486.3406703
- [37] P. Malo, A. Sinha, P. Korhonen, J. Wallenius, and P. Takala, “Good debt or bad debt: Detecting semantic orientations in economic texts,” J. of the Association for Information Science and Technology, Vol.65, No.4, pp. 782-796, 2014. https://doi.org/10.1002/asi.23062
- [38] M. Maia et al., “WWW’18 Open Challenge: Financial opinion mining and question answering,” Companion Proc. of the Web Conf. 2018, pp. 1941-1942, 2018. https://doi.org/10.1145/3184558.3192301
- [39] A. Sinha and T. Khandait, “Impact of news on the commodity market: Dataset and results,” Advances in Information and Communication: Proc. of the 2021 Future of Information and Communication Conf., Vol.2, pp. 589-601, 2021. https://doi.org/10.1007/978-3-030-73103-8_41
- [40] A. Shah, S. Paturi, and S. Chava, “Trillion dollar words: A new financial dataset, task & market analysis,” Proc. of the 61st Annual Meeting of the Association for Computational Linguistics (Vol.1: Long Papers), pp. 6664-6679, 2023. https://doi.org/10.18653/v1/2023.acl-long.368
- [41] J. C. S. Alvarado, K. Verspoor, and T. Baldwin, “Domain adaption of named entity recognition to support credit risk assessment,” Proc. of the Australasian Language Technology Association Workshop 2015, pp. 84-90, 2015.
- [42] A. Shah, A. Gullapalli, R. Vithani, M. Galarnyk, and S. Chava, “FiNER: Financial named entity recognition dataset and weak-supervision model,” arXiv:2302.11157v1, 2023. https://doi.org/10.48550/arXiv.2302.11157
- [43] Q. Xie et al., “FinBen: An holistic financial benchmark for large language models,” arXiv:2402.12659, 2024. https://doi.org/10.48550/arXiv.2402.12659
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.