Research Paper:
Size Classification for Software Maintenance Projects Through a Simplified Minimalist Machine Learning Algorithm
Cuauhtémoc López-Martín*
, Cornelio Yáñez-Márquez**,
, and Ali Bou Nassif***

*Department of Information Systems, Universidad de Guadalajara
Periférico Norte No.799 Núcleo Universitario, C. Prol. Belenes, Zapopan, Jalisco 45100, México
**Centro de Investigación en Computación, Instituto Politécnico Nacional
Av. Juan de Dios Bátiz S/N, Nueva Industrial Vallejo, Gustavo A. Madero, Ciudad de México 07738, México
Corresponding author
***Department of Computer Engineering, University of Sharjah
P.O. Box 27272, United Arab Emirates
In the software engineering field, the size of projects is used as explanatory variable for predicting the effort, duration, defects, costs, or risks of a project. Thus, the software project size has been considered the most influential factor. A systematic mapping study on the use of categorical data in software prediction concludes that the use of categorical data as explanatory variable in prediction models is an important issue because in the first phases of the software development process, the information is expressed in a categorical manner rather than numerical; however, at date, software size has mostly been used in its quantitative form rather than in its categorical representation. Because the software enhancement maintenance has the highest impact on business, our purpose is to classify software enhancement projects from their size by applying a new model termed simplified minimalist machine learning (S-MML) algorithm, which belongs to the minimalist machine learning (MML) paradigm. S-MML reduces five attributes commonly used for sizing a software project to only one. The performance of the S-MML is compared to those obtained from three classifiers. Seven public datasets of projects obtained from an international public repository of software projects were used to train and test the classifiers. Results showed that the S-MML had a better f-measure than the other three classifiers for all of the datasets at 95% confidence. We can conclude that the S-MML can be applied to classify the size of software enhancement projects.
- [1] D. Minh, H. X. Wang, Y. F. Li, and T. N. Nguyen “Explainable artificial intelligence: A comprehensive review,” Artificial Intelligence Review, Vol.55, pp. 3503-3568, 2022. https://doi.org/10.1007/s10462-021-10088-y
- [2] P. P. Angelov, E. A. Soares, R. Jiang, N. I. Arnold, and P. M. Atkinson, “Explainable artificial intelligence: An analytical review,” Data Mining and Knowledge Discovery, Vol.11, Issue 5, Article No.e1424, 2021. https://doi.org/10.1002/widm.1424
- [3] Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature, Vol.521, pp. 436-444, 2015. https://doi.org/10.1038/nature14539
- [4] A. Adadi and M. Berrada, “Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI),” IEEE ACCESS, Vol.6, pp. 52138-52160, 2018. https://doi.org/10.1109/ACCESS.2018.2870052
- [5] H. Hagras, “Toward Human-Understandable, Explainable AI,” Computer, Vol.51, Issue 9, pp. 28-36, 2018. https://doi.org/10.1109/MC.2018.3620965
- [6] D. Gunning and D. W. Aha, “DARPA’s explainable artificial intelligence program,” AI Magazine, Vol.40, Issue 2, pp. 3-84, 2019. https://doi.org/10.1609/aimag.v40i2.2850
- [7] C. Yáñez-Márquez, “Toward the Bleaching of the Black Boxes: Minimalist Machine Learning,” IT Professional, Vol.22, Issue 4, pp. 51-56, 2020. https://doi.org/10.1109/MITP.2020.2994188
- [8] Y. Villuendas-Rey, C. F. Rey-Benguría, A. Ferreira-Santiago, O. Camacho-Nieto, and C. Yáñez-Márquez, “The Naïve Associative Classifier (NAC): A novel, simple, transparent, and accurate classification model evaluated on financial data,” Neurocomputing, Vol.265, pp. 105-115, 2017. https://doi.org/10.1016/j.neucom.2017.03.085
- [9] C. Cortes and V. Vapnik, “Support-vector networks,” Machine Learning, Vol.20, No.3, pp. 273-297, 1995. https://doi.org/10.1007/BF00994018
- [10] P. Tripathy and K. Naik, “Software Evolution and Maintenance: A Practitioner’s Approach,” Wiley, 2015. https://doi.org/10.1002/9781118964637
- [11] ISBSG, “Guidelines for use of the ISBSG data Release 2018,” Int. Software Benchmarking Standards Group, 2018.
- [12] P. Bourque and R. Fairley, “Guide to the Software Engineering Body of Knowledge: SWEBoK V3.0,” IEEE Computer Society, 2014.
- [13] A. B. Nassif, M. Azzeh, L. F. Capretz, and D. Ho, “Neural network models for software development effort estimation: A comparative study,” Neural Computing and Applications, Vol.27, pp. 2369-2381, 2016. https://doi.org/10.1007/s00521-015-2127-1
- [14] P. Pospieszny, B. Czarnacka-Chrobot, and A. Kobylinski, “An effective approach for software project effort and duration estimation with machine learning algorithms,” J. of Systems and Software, Vol.137, pp. 184-196, 2018. https://doi.org/10.1016/j.jss.2017.11.066
- [15] P. Ardimento, L. Aversano, M. L. Bernardi, M. Cimitile, and M. Iammarino, “Just-in-time software defect prediction using deep temporal convolutional networks,” Neural Computing and Applications, Vol.34, pp. 3981-4001, 2022. https://doi.org/10.1007/s00521-021-06659-3
- [16] L. Lavazza and S. Morasca, “Empirical evaluation and proposals for bands-based COSMIC early estimation methods,” Information and Software Technology, Vol.109, pp. 108-125, 2019. https://doi.org/10.1016/j.infsof.2019.02.002
- [17] M. Ochodek, “Functional size approximation based on use-case names,” Information and Software Technology, Vol.80, pp. 73-88, 2016. https://doi.org/10.1016/j.infsof.2016.08.007
- [18] A. Ali and C. Gravino, “A systematic literature review of software effort prediction using machine learning methods,” J. of Software: Evolution and Process, Vol.31, Issue 10, Article No.e2211, 2019. https://doi.org/10.1002/smr.2211
- [19] A. Abran, J. M. Desharnais, S. Oligny, D. St-Pierre, and C. Symons, “The COSMIC Functional Size Measurement Method Version 3.0.1, Measurement Manual (The COSMIC Implementation Guide for ISO/IEC 19761: 2003),” The COSMIC Group, 2009.
- [20] A. Abran, “Software Project Estimation: The Fundamentals for Providing High Quality Information to Decision Makers,” IEEE Computer Society, 2015. https://doi.org/10.1002/9781118959312
- [21] ISBSG, “Field descriptions ISBSG D&E repository,” Int. Software Benchmarking Standards Group, 2018.
- [22] F. G. Wilkie, I. R. McChesney, P. Morrow, C. Tuxworth, and N. G. Lester, “The value of software sizing,” Information and Software Technology, Vol.53, Issue 11, pp. 1236-1249, 2011. https://doi.org/10.1016/j.infsof.2011.05.008
- [23] M. Alyahya, R. Ahmad, and S. P. Lee, “Impact of CMMI-based process maturity levels on effort, productivity and diseconomy of scale,” The Int. Arab J. of Information Technology, Vol.9, No.4, pp. 352-360, 2012.
- [24] J. T. Hancock and T. M. Khoshgoftaar, “Survey on categorical data for neural networks,” J. of Big Data, Vol.7, Article No.28, 2020. https://doi.org/10.1186/s40537-020-00305-w
- [25] L. A. Zadeh, “From computing with numbers to computing with words – From manipulation of measurements to manipulation of perceptions,” IEEE Trans. on Circuits and Systems I: Fundamental Theory and Applications, Vol.46, Issue 1, pp. 105-119, 1999. https://doi.org/10.1109/81.739259
- [26] F.-A. Amazal, A. Idri, and A. Abran, “An Analogy-Based Approach to Estimation of Software Development Effort Using Categorical Data,” 2014 Joint Conf. of the Int. Workshop on Software Measurement and the Int. Conf. on Software Process and Product Measurement, 2014. https://doi.org/10.1109/IWSM.Mensura.2014.31
- [27] S. S. Stevens, “On the Theory of Scales of Measurement,” Science, Vol.103, Issue 2684, pp. 677-680, 1946. https://doi.org/10.1126/science.103.2684.677
- [28] R. A. Fisher, “Statistical Methods for Research Workers (13th ed.),” Oliver and Boyd, 1938.
- [29] C. Hayashi, “On the quantification of qualitative data from the mathematico-statistical point of view,” Annals of the Institute of Statistical Mathematics, Vol.2, pp. 35-47, 1950. https://doi.org/10.1007/BF02919500
- [30] C. Hayashi, “On the prediction of phenomena from qualitative data and the quantification of qualitative data from the mathematico-statistical point of view,” Annals of the Institute of Statistical Mathematics, Vol.3, pp. 69-98, 1951. https://doi.org/10.1007/BF02949778
- [31] F. A. Amazal and A. Idri, “Handling of Categorical Data in Software Development Effort Estimation: A Systematic Mapping Study,” Proc. of the 2019 Federated Conf. on Computer Science and Information Systems (FedCSIS), Vol.18, pp. 763-770, 2019. https://doi.org/10.15439/2019F222
- [32] D. C. Montgomery, E. A. Peck, and G. G. Vining, “Introduction to Linear Regression Analysis (5th ed.),” Wiley, 2012.
- [33] D. Garmus and D. Herron, “Measuring the software process, A practical guide to functional measurements,” Prentice-Hall, 1996.
- [34] B. Kitchenham and E. Mendes, “Why comparative effort prediction studies may be invalid,” Proc. of the 5th Int. Conf. on Predictor Models in Software Engineering, 2009. https://doi.org/10.1145/1540438.1540444
- [35] P. Filzmoser, K. Hron, and M. Templ, “Discriminant Analysis,” Applied Compositional Data Analysis, Springer Series in Statistics (SSS), pp. 163-179, 2018. https://doi.org/10.1007/978-3-319-96422-5_9
- [36] J. Wen, S. Li, Z. Lin, Y. Hu, and C. Huang, “Systematic literature review of machine learning based software development effort estimation models,” Information and Software Technology, Vol.54, Issue 1, pp. 41-59, 2012. https://doi.org/10.1016/j.infsof.2011.09.002
- [37] S. Haykin, “Neural networks and learning machines (3rd ed.),” Pearson, 2012.
- [38] L. H. Son, N. Pritam, M. Khari, R. Kumar, P. T. M. Phuong, and P. H. Thong, “Empirical Study of Software Defect Prediction: A Systematic Mapping,” Symmetry, Vol.11, Issue 2, Article No.212, 2019. https://doi.org/10.3390/sym11020212
- [39] R. Özakınc and A. Tarhan, “Early software defect prediction: A systematic map and review,” J. of Systems and Software, Vol.144, pp. 216-239, 2018. https://doi.org/10.1016/j.jss.2018.06.025
- [40] C. Catal and B. Diri, “Investigating the effect of dataset size, metrics sets, and feature selection techniques on software fault prediction problem,” Information Sciences, Vol.179, Issue 8, pp. 1040-1058, 2009. https://doi.org/10.1016/j.ins.2008.12.001
- [41] K. El Emam, S. Benlarbi, N. Goel, W. Melo, H. Lounis, and S. N. Rai, “The optimal class size for object-oriented software,” IEEE Trans. on Software Engineering, Vol.28, Issue 5, pp. 494-509, 2002. https://doi.org/10.1109/TSE.2002.1000452
- [42] L. Fink and Y. Lichtenstein, “Why project size matters for contract choice in software development outsourcing,” ACM SIGMIS Database: The DATABASE for Advances in Information Systems, Vol.45, Issue 3, pp. 54-71, 2014. https://doi.org/10.1145/2659254.2659258
- [43] B. Kitchenham and E. Mendes, “Software productivity measurement using multiple size measures,” IEEE Trans. on Software Engineering, Vol.30, Issue 12, pp. 1023-1035, 2004. https://doi.org/10.1109/TSE.2004.104
- [44] H. B. K. Tan, Y. Zhao, and H. Zhang, “Conceptual data model-based software size estimation for information systems,” ACM Trans. on Software Engineering and Methodology (TOSEM), Vol.19, Issue 2, Article No.4, 2009. https://doi.org/10.1145/1571629.1571630
- [45] K. Lind and R. Heldal, “A Practical Approach to Size Estimation of Embedded Software Components,” IEEE Trans. on Software Engineering, Vol.38, Issue 5, pp. 993-1007, 2012. https://doi.org/10.1109/TSE.2011.86
- [46] S. G. MacDonell, “Software source code sizing using fuzzy logic modeling,” Information and Software Technology, Vol.45, Issue 7, pp. 389-404, 2003. https://doi.org/10.1016/S0950-5849(03)00011-9
- [47] M. Ochodek, “Functional size approximation based on use-case names,” Information and Software Technology, Vol.80, pp. 73-88, 2016. https://doi.org/10.1016/j.infsof.2016.08.007
- [48] P. C. Pendharkar, “An exploratory study of object-oriented software component size determinants and the application of regression tree forecasting models,” Information & Management, Vol.42, Issue 1, pp. 61-73, 2004. https://doi.org/10.1016/j.im.2003.12.004
- [49] J. Verner and G. Tate, “A software size model,” IEEE Trans. on Software Engineering, Vol.18, Issue 4, pp. 265-278, 1992. https://doi.org/10.1109/32.129216
- [50] J. Aguilar, M. Sánchez, C. Fernández-y-Fernández, E. Rocha, D. Martínez, and J. Figueroa, “The Size of Software Projects Developed by Mexican Companies,” 2014 Int. Conf. Software Engineering, Research and Practice, 2014. https://doi.org/10.48550/arXiv.1408.1068
- [51] J.-L. Solorio-Ramírez, M. Saldana-Perez, M. D. Lytras, M.-A. Moreno-Ibarra, and C. Yáñez-Márquez, “Brain Hemorrhage Classification in CT Scan Images Using Minimalist Machine Learning,” Diagnostics, Vol.11, Issue 8, Article No.1449, 2021. https://doi.org/10.3390/diagnostics11081449
- [52] V. N. Vapnik, “Statistical learning theory,” Wiley-Interscience, 1998.
- [53] B. E. Boser, I. M. Guyon, and V. N. Vapnik, “A training algorithm for optimal margin classifiers,” Proc. of the 5th Annual Workshop on Computational Learning Theory, pp. 144-152, 1992. https://doi.org/10.1145/130385.130401
- [54] B. Schölkopf, A. J. Smola, R. C. Williamson, and P. L. Bartlett, “New Support Vector Algorithms,” Neural Computation, Vol.12, Issue 5, pp. 1207-1245, 2000. https://doi.org/10.1162/089976600300015565
- [55] D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature, Vol.323, pp. 533-536, 1986. https://doi.org/10.1038/323533a0
- [56] A. Luque, A. Carrasco, A. Martín, and A. de las Heras, “The impact of class imbalance in classification performance metrics based on the binary confusion matrix,” Pattern Recognition, Vol.91, pp. 216-231, 2019. https://doi.org/10.1016/j.patcog.2019.02.023
- [57] J. Ortigosa-Hernández, I. Inza, and J. A. Lozano, “Measuring the class-imbalance extent of multi-class problems,” Pattern Recognition Letters, Vol.98, pp. 32-38, 2017. https://doi.org/10.1016/j.patrec.2017.08.002
- [58] A. Fernández, S. García, M. J. del Jesús, and F. Herrera, “A study of the behaviour of linguistic fuzzy rule based classification systems in the framework of imbalanced data-sets,” Fuzzy Sets and Systems, Vol.159, Issue 18, pp. 2378-2398, 2008. https://doi.org/10.1016/j.fss.2007.12.023
- [59] B. W. Boehm, “Software Engineering Economics,” Prentice Hall, 1981.
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.