single-jc.php

JACIII Vol.30 No.4 pp. 1243-1257
(2026)

Research Paper:

fMRI Encoding and Decoding with LLMs: Input Embeddings vs. Hidden State Representations

Muxuan Liu ORCID Icon and Ichiro Kobayashi ORCID Icon

Ochanomizu University
2-1-1 Ohtsuka, Bunkyo-ku, Tokyo 112-8610, Japan

Corresponding author

Received:
September 11, 2025
Accepted:
March 12, 2026
Published:
July 20, 2026
Keywords:
fMRI, language models, encoding–decoding
Abstract

Systematically comparing how linguistic representations relate to brain activity has become an important topic in the field of computational neuroscience. Prior studies have mainly relied on contextual hidden states combined with linear regression, leaving open questions about the role of static input embeddings and the benefits of nonlinear mappings. In this study, we compare input embeddings and hidden states from multiple language model families (BERT, GPT-2, and LLaMA) within both encoding frameworks, which map text features to brain responses, and decoding frameworks, which reconstruct linguistic features from brain activity. We benchmarked voxel-wise ridge regression against bidirectional long short-term memory (BiLSTMs) models, using repeat-split cross-validation and explainable variance normalization on functional magnetic resonance imaging (fMRI) data from three subjects. Our analyses demonstrate that input embeddings, despite being context-invariant, remain competitive and, in some cases, outperform hidden states, while BiLSTMs provide modest but region-specific improvements over ridge regression. Fine-grained voxel-level results further revealed distinct cortical distributions of stable versus context-dependent features. Together, these findings clarify the trade-off between predictive performance and interpretability and highlight that input embeddings offer a strong and interpretable baseline for representational alignment between language models and brain activity.

fMRI encoding flatmap

fMRI encoding flatmap

Cite this article as:
M. Liu and I. Kobayashi, “fMRI Encoding and Decoding with LLMs: Input Embeddings vs. Hidden State Representations,” J. Adv. Comput. Intell. Intell. Inform., Vol.30 No.4, pp. 1243-1257, 2026.
Data files:
References
  1. [1] T. Naselaris, K. N. Kay, S. Nishimoto, and J. L. Gallant, “Encoding and decoding in fMRI,” NeuroImage, Vol.56, No.2, pp. 400-410, 2011. https://doi.org/10.1016/j.neuroimage.2010.07.073
  2. [2] A. G. Huth, S. Nishimoto, A. T. Vu, and J. L. Gallant, “A continuous semantic space describes the representation of thousands of object and action categories across the human brain,” Neuron, Vol.76, No.6, pp. 1210-1224, 2012. https://doi.org/10.1016/j.neuron.2012.10.014
  3. [3] A. Mittal, P. Aggarwal, L. Pessoa, and A. Gupta, “Robust Brain State Decoding Using Bidirectional Long Short Term Memory Networks in functional MRI,” bioRxiv, 2021. https://doi.org/10.1101/2021.06.18.449069
  4. [4] K. Qiao, J. Chen, L. Wang, C. Zhang, L. Zeng, L. Tong, and B. Yan, “Category Decoding of Visual Stimuli From Human Brain Activity Using a Bidirectional Recurrent Neural Network to Simulate Bidirectional Information Flows in Human Visual Cortices,” Frontiers in Neuroscience, Vol.13, Article No.692, 2019. https://doi.org/10.3389/fnins.2019.00692
  5. [5] M. Kucukosmanoglu, J. O. Garcia, J. Brooks, and K. Bansal, “Influence of cognitive networks and task performance on fMRI-based state classification using DNN models,” Scientific Reports, Vol.15, No.1, Article No.23689, 2025. https://doi.org/10.1038/s41598-025-05690-x
  6. [6] Y. Sun, D. Chahine, Q. Wen, T. Liu, X. Li, Y. Yuan, F. Calamante, and J. Lv, “Voxel-Level Brain States Prediction Using Swin Transformer,” IEEE J. of Biomedical and Health Informatics, Vol.29, No.12, pp. 8719-8726, 2025. https://doi.org/10.1109/JBHI.2025.3613793
  7. [7] Y. Sun, M. Cabezas, J. Lee, C. Wang, W. Zhang, F. Calamante, and J. Lv, “Predicting Human Brain States with Transformer,” A. Schroder et al. (Eds.), “Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 Workshops,” pp. 136-146, Springer Nature Switzerland, 2025. https://doi.org/10.1007/978-3-031-84525-3_12
  8. [8] M. Toneva and L. Wehbe, “Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain),” Advances in Neural Information Processing Systems, Vol.32, 2019.
  9. [9] M. Schrimpf, I. A. Blank, G. Tuckute, C. Kauf, E. A. Hosseini, N. Kanwisher, J. B. Tenenbaum, and E. Fedorenko, “The neural architecture of language: Integrative modeling converges on predictive processing,” Proc. of the National Academy of Sciences, Vol.118, No.45, Article No.e2105646118, 2021. https://doi.org/10.1073/pnas.2105646118
  10. [10] C. Caucheteux and J.-R. King, “Brains and algorithms partially converge in natural language processing,” Communications Biology, Vol.5, No.1, Article No.134, 2022. https://doi.org/10.1038/s42003-022-03036-1
  11. [11] M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” arXiv preprint, arXiv:1802.05365, 2018. https://doi.org/10.48550/arXiv.1802.05365
  12. [12] M. Liu and I. Kobayashi, “Do Feature Representations from Different Language Models Affect Accuracy of Brain Encoding Models’ Predictions?,” 2024 IEEE Int. Conf. on Systems, Man, and Cybernetics (SMC), pp. 2766-2771, 2024. https://doi.org/10.1109/SMC54092.2024.10831584
  13. [13] R. Antonello, A. Vaidya, and A. G. Huth, “Scaling laws for language encoding models in fMRI,” Advances in Neural Information Processing Systems, Vol.36, pp. 21895-21907, 2023.
  14. [14] J. Tang, M. Du, V. A. Vo, V. Lal, and A. G. Huth, “Brain encoding models based on multimodal transformers can transfer across language and vision,” Proc. of the 37th Int. Conf. on Neural Information Processing Systems (NIPS ’23), 2023.
  15. [15] S. R. Oota, Z. Chen, M. Gupta, R. S. Bapi, G. Jobard, F. Alexandre, and X. Hinaut, “Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey),” arXiv preprint, arXiv:2307.10246, 2024. https://doi.org/10.48550/arXiv.2307.10246
  16. [16] T. D. la Tour, M. Eickenberg, A. O. Nunez-Elizalde, and J. L. Gallant, “Feature-space selection with banded ridge regression,” NeuroImage, Vol.264, Article No.119728, 2022. https://doi.org/https://doi.org/10.1016/j.neuroimage.2022.119728
  17. [17] N. Affolter, B. Egressy, D. Pascual, and R. Wattenhofer, “Brain2Word: Decoding Brain Activity for Language Generation,” arXiv preprint, arXiv:2009.04765, 2020. https://doi.org/10.48550/arXiv.2009.04765
  18. [18] J. Tang et al., “Semantic reconstruction of continuous language from non-invasive brain recordings,” Nature Neuroscience, Vol.26, pp. 858-866, 2023. https://doi.org/10.1038/s41593-023-01304-9
  19. [19] P. Liu, G. Dong, D. Guo, K. Li, F. Li, X. Yang, M. Wang, and X. Ying, “A Survey on fMRI-based Brain Decoding for Reconstructing Multimodal Stimuli,”  arXiv preprint, arXiv:2503.15978, 2025. https://doi.org/10.48550/arXiv.2503.15978
  20. [20] A. Goldstein, A. Grinstein-Dabush, M. Schain, H. Wang, Z. Hong, B. Aubrey, S. A. Nastase, Z. Zada, E. Ham, A. Feder et al., “Alignment of brain embeddings and artificial contextual embeddings in natural language points to common geometric patterns,” Nature Communications, Vol.15, No.1, Article No.2768, 2024. https://doi.org/10.1038/s41467-024-46631-y
  21. [21] L. Zhao, Z. Wu, H. Dai, Z. Liu, X. Hu, T. Zhang, D. Zhu, and T. Liu, “A generic framework for embedding human brain function with temporally correlated autoencoder,” Medical Image Analysis, Vol.89, Article No.102892, 2023. https://doi.org/https://doi.org/10.1016/j.media.2023.102892
  22. [22] H. Li and Y. Fan, “Interpretable, highly accurate brain decoding of subtly distinct brain states from functional MRI using intrinsic functional networks and long short-term memory recurrent neural networks,” NeuroImage, Vol.202, Article No.116059, 2019. https://doi.org/10.1016/j.neuroimage.2019.116059
  23. [23] S. Ma, L. Wang, L. Hou, S. Hou, and B. Yan, “An fMRI visual neural encoding method with multimodal large language model,” Knowledge-Based Systems, Vol.326, Article No.114049, 2025. https://doi.org/10.1016/j.knosys.2025.114049
  24. [24] A. LeBel, L. Wagner, S. Jain, A. Adhikari-Desai, B. Gupta, A. Morgenthal, J. Tang, L. Xu, and A. G. Huth, “A natural language fMRI dataset for voxelwise encoding models,” Scientific Data, Vol.10, No.1, Article No.555, 2023. https://doi.org/10.1038/s41597-023-02437-z
  25. [25] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” arXiv preprint, arXiv:2212.04356, 2022. https://doi.org/10.48550/arXiv.2212.04356
  26. [26] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A Lite BERT for Self-supervised Learning of Language Representations,” arXiv preprint, arXiv:1909.11942, 2019. https://doi.org/10.48550/arXiv.1909.11942
  27. [27] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint, arXiv:1810.04805, 2018. https://doi.org/10.48550/arXiv.1810.04805
  28. [28] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint, arXiv:1907.11692, 2019. https://doi.org/10.48550/arXiv.1907.11692
  29. [29] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog, 2019.
  30. [30] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models,” arXiv preprint, arXiv:2307.09288, 2023. https://doi.org/10.48550/arXiv.2307.09288
  31. [31] A. Grattafiori et al., “The Llama 3 Herd of Models,” 2024. LLaMA 3.1-8B model: https://huggingface.co/meta-llama/Llama-3.1-8B [Accessed March 20, 2025]
  32. [32] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush, “Transformers: State-of-the-Art Natural Language Processing,” Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38-45, 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6
  33. [33] M. Sahani and J. F. Linden, “How Linear are Auditory Cortical Responses?,” Advances in Neural Information Processing Systems, Vol.15, 2002.
  34. [34] A. Hsu, A. Borst, and F. E. Theunissen, “Quantifying variability in neural responses and its application for the validation of model predictions,” Network: Computation in Neural Systems, Vol.15, No.2, pp. 91-109, 2004. https://doi.org/10.1088/0954-898X/15/2/002
  35. [35] O. Schoppe, N. S. Harper, B. D. B. Willmore, A. J. King, and J. W. H. Schnupp, “Measuring the performance of neural models,” Frontiers in Computational Neuroscience, Vol.10, Article No.10, 2016. https://doi.org/10.3389/fncom.2016.00010
  36. [36] F. Reichel, “On Bessel’s Correction: Unbiased Variance, Degrees of Freedom, and the Sum of Pairwise Differences,” Qeios, 2025. https://doi.org/10.32388/3GJGNA
  37. [37] W. J. Reichmann, “The Use and Abuse of Statistics,” Oxford University Press, 1961.
  38. [38] G. Upton and I. Cook, “A Dictionary of Statistics, 2nd ed.,” Oxford University Press, 2008.
  39. [39] A. G. Huth and G. Lab, “Voxelwise encoding model (VEM) tutorials,” 2021. https://gallantlab.org/voxelwise_tutorials/ [Accessed September 10, 2025]
  40. [40] J. Brownlee, “Deep Learning for Time Series Forecasting: Predict the Future with MLPs, CNNs and LSTMs in Python,” Machine Learning Mastery, 2018.
  41. [41] Data Science StackExchange, “Sliding window leads to overfitting in LSTM?,” 2018. https://datascience.stackexchange.com/questions/27628/ [Accessed September 10, 2025]
  42. [42] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” Proc. of the Int. Conf. on Learning Representations (ICLR), 2019.
  43. [43] I. Loshchilov and F. Hutter, “SGDR: Stochastic Gradient Descent with Warm Restarts,” Proc. of the Int. Conf. on Learning Representations (ICLR), 2017.
  44. [44] L. Prechelt, “Early Stopping – But When,?” G. B. Orr and K.-R. Müller (Eds.), “Neural Networks: Tricks of the Trade," Lecture Notes in Computer Science, Vol.1524, pp. 55-69, Springer, 1998. https://doi.org/10.1007/3-540-49430-8_3

*This site is desgined based on HTML5 and CSS3 for modern browsers, e.g. Chrome, Firefox, Safari, Edge, Opera.

Last updated on Jul. 19, 2026