Date of Submission
Fall 2025
Supervisor
Dr. Faisal Iradat, Assistant Professor, Department of Computer Science, School of Mathematics and Computer Science (SMCS)
Co-Supervisor
Dr. Belgin Metin (Bogazici University), Dr. Nazım Taşkın (Bogazici University)
Committee Member 1
Dr. Tariq Mahmood, Examiner – I, Institute of Business Administration (IBA), Karachi, Institute of Business Administration (IBA), Karachi
Committee Member 2
Dr. Mian Muhammad Waseem Iqbal, Examiner-II, NUST
Degree
Master of Science in Data Science
Department
Department of Computer Science
Faculty/ School
School of Mathematics and Computer Science (SMCS)
Keywords
Explainable Artificial Intelligence, API Vulnerability Detection, Machine Learning, SHapley Additive exPlanations, Feature Importance Analysis, Model Interpretability, Security, Shift-Left Security
Abstract
The digital transformation sweeping the globe relies heavily on Application Programming Interfaces (API) for information exchange which results in an upsurge in API use which often remain unmanaged and introduces substantial security risks. The OWASP API Security Top Ten list highlights the critical vulnerabilities surrounding unmanaged APIs which plays a key role in guiding existing research and industry efforts. The existing research around API Security is more focused on shift right security paradigm leaving ample room for shift left security i.e. security early in software development life cycle (SDLC). Catching security issues early in SDLC requires vast data on best practices, which then can be made to learned through machine learning models. Though acquiring the data on best practices is a struggling objective, hampering the focus shift left security. This study leverages the OWASP API Top Ten Vulnerabilities to construct a focused, expert-validated dataset of 60 labeled Python authentication snippets, serving as a benchmark for Shift-Left vulnerability detection. We employed this dataset to train various Machine Learning models using targeted Natural Language Processing (NLP) feature extraction tools, which have been studied at length though for human language. Our results indicate that ensemble methods, particularly the Extra Trees Classifier, significantly outperform single classifiers, achieving a test accuracy of 83.3% and demonstrating effective learning even from limited high-quality data. To enhance interpretability, we integrated eXplainable Artificial Intelligence (XAI) using SHapley Additive exPlanations (SHAP). This analysis revealed critical feature importance distributions, identifying that security-critical tokens are often overlooked in standard development patterns, thus validating the model’s decision-making process for developers.
Document Type
Restricted Access
Submission Type
Thesis
Recommended Citation
Jehanzeb, S. (2025). Code Validation Through Machine Learning: A Shift-Left Focus Exploratory Study (Unpublished Unpublished graduate thesis). Retrieved from https://ir.iba.edu.pk/etd-ms-ds/18
