Research Paper:
Deep Learning for the Production of Official Statistics: Density Ratio Estimation Using Biased Transaction Data for Japanese Labor Statistics
Yuya Takada*,**
, Yuri Murayama*
, and Kiyoshi Izumi*

*Department of Systems Innovation, School of Engineering, The University of Tokyo
7-3-1 Hongo, Bunkyo-ku, Tokyo 113-8656, Japan
**Re Data Science Co., Ltd.
Kashiwa-no-ha Open Innovation Lab, 178-4 Wakashiba, Kashiwa, Chiba 277-0871, Japan
National statistical offices are exploring alternative data sources for official statistics. Such data, including point-of-sale (POS) records and mobile phone GPS logs, are originally collected for operational rather than statistical purposes. In the era of big data, the private sector accumulates vast volumes of transaction data, and leveraging such data for official statistics has become an emerging priority. However, these efforts face significant challenges, primarily due to severe selection bias stemming from such data. Prior research has shown that, even if non-representative, transaction data can produce timely statistics via density ratio estimation methods from machine learning. As a proof of concept, that study demonstrated that preliminary estimates could be generated using biased data from a Japanese private employment agency, enabling the early release of a labor market indicator otherwise delayed by up to a year. Building on this, the present study incorporates deep learning into density ratio estimation to improve accuracy. While deep learning, when applied to density ratio estimation, is often considered prone to overfitting, this study demonstrates that it can improve estimation accuracy without overfitting. Moreover, although deep learning is typically regarded as requiring extensive hyperparameter tuning, we show that it can be implemented without a significant tuning burden, supporting its practical use in the production of official statistics.
Errors and time lags in official statistics
- [1] Y. Takada and K. Izumi, “Implementation of biased big data to the Japanese official labor statistics using supervised learning under covariate shift,” 2022 IEEE Int. Conf. on Big Data (Big Data), pp. 2062-2071, 2022. https://doi.org/10.1109/BigData55660.2022.10020563
- [2] Y. Takada and K. Izumi, “Machine learning for the production of official statistics: Density ratio estimation using biased transaction data for Japanese labor statistics,” arXiv preprint, arXiv:2510.24153, 2025. https://doi.org/10.48550/arXiv.2510.24153
- [3] M. Sugiyama, T. Suzuki, and T. Kanamori, “Density-ratio matching under the Bregman divergence: A unified framework of density-ratio estimation,” Annals of the Institute of Statistical Mathematics, Vol.64, pp. 1009-1044, 2012. https://doi.org/10.1007/s10463-011-0343-8
- [4] M. Kato and T. Teshima, “Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio Estimation,” Proc. of the 38th Int. Conf. on Machine Learning, Vol.PMLR139, pp. 5320-5333, 2021. https://proceedings.mlr.press/v139/kato21a.html
- [5] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Advances in Neural Information Processing Systems, Vol.25, pp. 1097-1105, 2012. https://doi.org/10.1145/3065386
- [6] Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural probabilistic language model,” J. of Machine Learning Research, Vol.3, pp. 1137-1155, 2003.
- [7] United Nations Statistics Division, “Task Teams: Mobile Phone Data,” https://unstats.un.org/bigdata/task-teams/mobile-phone/index.cshtml [Accessed June 21, 2026]
- [8] U. N. S. Division, “Handbook on the use of mobile phone data for official statistics,” 2019.
- [9] J. Kroon, “Mobile Positioning as a Possible Data Source for International Travel Service Statistics,” UNECE Conf. of European Statisticians, 2012.
- [10] R. C. Feenstra and M. D. Shapiro, “High-frequency substitution and the measurement of price indexes,” NBER Chapters, Scanner Data and Price Indexes, pp. 123-146, 2003. https://doi.org/10.7208/chicago/9780226239668.003.0007
- [11] J. D. Haan and H. A. V. der Grient, “Eliminating chain drift in price indexes based on scanner data,” J. of Econometrics, Vol.161, No.1, pp. 36-46, 2011. https://doi.org/10.1016/j.jeconom.2010.09.004
- [12] L. Ivancic, W. E. Diewert, and K. J. Fox, “Scanner data, time aggregation and the construction of price indexes,” J. of Econometrics, Vol.161, No.1, pp. 24-35, 2011. https://doi.org/10.1016/j.jeconom.2010.09.003
- [13] K. Watanabe and T. Watanabe, “Estimating daily inflation using scanner data: A progress report,” CARF Working Paper, Article No.CARF-F-342, 2014.
- [14] A. Cavallo, “Online and official price indexes: Measuring Argentina’s inflation,” J. of Monetary Economics, Vol.60, No.2, pp. 152-165, 2013. https://doi.org/10.1016/j.jmoneco.2012.10.002
- [15] A. Cavallo and R. Rigobon, “The billion prices project: Using online prices for measurement and research,” J. of Economic Perspectives, Vol.30, No.2, pp. 151-178, 2016. https://doi.org/10.1257/jep.30.2.151
- [16] A. Cavallo, W. E. Diewert, R. C. Feenstra, R. Inklaar, and M. P. Timmer, “Using online prices for measuring real consumption across countries,” AEA Papers and Proc., Vol.108, pp. 483-487, 2018. https://doi.org/10.1257/pandp.20181037
- [17] N. Woloszko, “Tracking activity in real time with Google Trends,” OECD Economics Department Working Papers, Article No.1634, 2020. https://doi.org/10.1787/6b9c7518-en
- [18] C. Schiavoni, F. Palm, S. Smeekes, and J. V. D. Brakel, “A dynamic factor model approach to incorporate big data in state space models for official statistics,” J. of the Royal Statistical Society Series A: Statistics in Society, Vol.184, No.1, pp. 324-353, 2019. https://doi.org/10.1111/rssa.12626
- [19] M. Sugiyama, T. Suzuki, and T. Kanamori, “Density Ratio Estimation in Machine Learning,” Cambridge University Press, 2012. https://doi.org/10.1017/CBO9781139035613
- [20] H. Shimodaira, “Improving predictive inference under covariate shift by weighting the log-likelihood function,” J. of Statistical Planning and Inference, Vol.90, No.2, pp. 227-244, 2000. https://doi.org/10.1016/S0378-3758(00)00115-4
- [21] D. B. Rubin, “Inference and missing data,” Biometrika, Vol.63, No.3, pp. 581-592, 1976. https://doi.org/10.1093/biomet/63.3.581
- [22] Y. Yang, A. K. Kuchibhotla, and E. Tchetgen Tchetgen, “Doubly robust calibration of prediction sets under covariate shift,” J. of the Royal Statistical Society Series B: Statistical Methodology, Vol.86, No.4, pp. 943-965, 2024. https://doi.org/10.1093/jrsssb/qkae009
- [23] J. W. Graham, “Missing data analysis: Making it work in the real world,” Annual Review of Psychology, Vol.60, pp. 549-576, 2009. https://doi.org/10.1146/annurev.psych.58.110405.085530
- [24] K. Morikawa and J. K. Kim, “Semiparametric optimal estimation with nonignorable nonresponse data,” The Annals of Statistics, Vol.49, No.5, pp. 2991-3014, 2021. https://doi.org/10.1214/21-AOS2070
- [25] T. Kanamori, S. Hido, and M. Sugiyama, “A least-squares approach to direct importance estimation,” The J. of Machine Learning Research, Vol.10, pp. 1391-1445, 2009.
- [26] A. Gretton, A. Smola, J. Huang, M. Schmittfull, K. Borgwardt, and B. Schölkopf, “Covariate Shift by Kernel Mean Matching,” J. Quiñonero-Candela, M. Sugiyama, A. Schwaighofer, and N. D. Lawrence (Eds.), “Dataset Shift in Machine Learning,” pp. 131-160, MIT Press, 2013. https://doi.org/10.7551/mitpress/9780262170055.003.0008
- [27] J. Huang, A. Gretton, K. Borgwardt, B. Schölkopf, and A. Smola, “Correcting sample selection bias by unlabeled data,” Proc. of the 19th Int. Conf. on Neural Information Processing Systems, pp. 601-608, 2006.
- [28] J. Qin, “Inferences for case-control and semiparametric two-sample density ratio models,” Biometrika, Vol.85, No.3, pp. 619-630, 1998. https://doi.org/10.1093/biomet/85.3.619
- [29] K. F. Cheng and C. K. Chu, “Semiparametric density estimation under a two-sample density ratio model,” Bernoulli, Vol.10, No.4, pp. 583-604, 2004. https://doi.org/10.3150/bj/1093265631
- [30] S. Bickel, M. Brückner, and T. Scheffer, “Discriminative learning for differing training and test distributions,” Proc. of the 24th Int. Conf. on Machine Learning, pp. 81-88, 2007. https://doi.org/https://doi.org/10.1145/1273496.1273507
- [31] M. Sugiyama, T. Suzuki, T. Kanamori, J. Sese, and I. Takeuchi, “Direct importance estimation for covariate shift adaptation,” Ann. Inst. Stat. Math., Vol.60, pp. 699-746, 2008. https://doi.org/10.1007/s10463-008-0197-x
- [32] Y. Tsuboi, H. Kashima, S. Hido, S. Bickel, and M. Sugiyama, “Direct density ratio estimation for large-scale covariate shift adaptation,” J. Inf. Process., Vol.17, pp. 138-155, 2009. https://doi.org/10.2197/ipsjjip.17.138
- [33] M. Yamada and M. Sugiyama, “Direct importance estimation with Gaussian mixture models,” IEICE Trans. Inf. Syst., Vol.E92-D, No.10, pp. 2159-2162, 2009. https://doi.org/10.1587/transinf.E92.D.2159
- [34] X. Nguyen, M. J. Wainwright, and M. I. Jordan, “Estimating divergence functionals and the likelihood ratio by convex risk minimization,” IEEE Trans. Inf. Theory, Vol.56, No.11, pp. 5847-5861, 2010. https://doi.org/10.1109/TIT.2010.2068870
- [35] M. Yamada, M. Sugiyama, G. Wichern, and J. Simm, “Direct importance estimation with a mixture of probabilistic principal component analyzers,” IEICE Trans. Inf. Syst., Vol.E93-D, No.10, pp. 2846-2849, 2010. https://doi.org/10.1587/transinf.E93.D.2846
- [36] R. Kiryo, G. Niu, M. C. du Plessis, and M. Sugiyama, “Positive-Unlabeled learning with non-negative risk estimator,” Advances in Neural Information Processing Systems, Vol.30, pp. 1674-1684, 2017.
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.