A Hybrid Machine Learning and Six Sigma Framework for Software Defect Prediction
SDGs: Primary SDG: SDG 9 - Industry, Innovation, and Infrastructure. | Secondary SDGs: SDG 8 - Decent Work and Economic Growth; SDG 12 - Responsible Consumption and Production
Keywords:
Software Defect Prediction, Six Sigma, Machine Learning, Statistical Quality Control, Process Capability, DMAIC, Random ForestAbstract
Software defect prediction (SDP) is a proactive approach that assists in identifying the modules which are prone to fault early enough so that the remediation and the use of the testing resources can be used in the most optimum manner. Machine Learning (ML) models have demonstrated significant predictions in the area, but in general, are applied without the use of well-established process improvement models; therefore, their role in quality control on a long-term basis is quite insignificant. The gap in this research is bridged by proposing the new hybrid framework which would integrate the predictive power of the Machine Learning models with the statistical method of the Six Sigma i.e. Define, Measure, Analyze, Improve, Control (DMAIC) approach. Here this research has empirically compared four classifiers, viz, Logistic Regression, Decision Tree, Random Forest, and Support Vector machine, on NASA MDP KC1 data. Random Forest model became the most excellent with an accuracy of 91.25 and Receiver Operating Characteristic-Area Under the Curve (ROC-AUC) of 0.964. The research work has found that cooperation between the ML based prediction and the statistical evaluation of Six Sigma provides a powerful data-driven paradigm to not only identify defects but also measure and maintain software quality improvement. The research work concludes with contributions to SDGs.