Student Name

Date of Submission

Fall 2025

Supervisor

Dr. Faisal Iradat, Assistant Professor, Department of Computer Science, School of Mathematics and Computer Science (SMCS)

Co-Supervisor

Dr. Belgin Metin (Bogazici University), Dr. Nazım Taşkın (Bogazici University)

Committee Member 1

Dr. Tariq Mahmood, Examiner – I, Institute of Business Administration (IBA), Karachi, Institute of Business Administration (IBA), Karachi

Committee Member 2

Dr. Mian Muhammad Waseem Iqbal, Examiner-II, NUST

Degree

Master of Science in Data Science

Department

Department of Computer Science

Faculty/ School

School of Mathematics and Computer Science (SMCS)

Keywords

Explainable Artificial Intelligence, API Vulnerability Detection, Machine Learning, SHapley Additive exPlanations, Feature Importance Analysis, Model Interpretability, Security, Shift-Left Security

Abstract

The digital transformation sweeping the globe relies heavily on Application Programming Interfaces (API) for information exchange which results in an upsurge in API use which often remain unmanaged and introduces substantial security risks. The OWASP API Security Top Ten list highlights the critical vulnerabilities surrounding unmanaged APIs which plays a key role in guiding existing research and industry efforts. The existing research around API Security is more focused on shift right security paradigm leaving ample room for shift left security i.e. security early in software development life cycle (SDLC). Catching security issues early in SDLC requires vast data on best practices, which then can be made to learned through machine learning models. Though acquiring the data on best practices is a struggling objective, hampering the focus shift left security. This study leverages the OWASP API Top Ten Vulnerabilities to construct a focused, expert-validated dataset of 60 labeled Python authentication snippets, serving as a benchmark for Shift-Left vulnerability detection. We employed this dataset to train various Machine Learning models using targeted Natural Language Processing (NLP) feature extraction tools, which have been studied at length though for human language. Our results indicate that ensemble methods, particularly the Extra Trees Classifier, significantly outperform single classifiers, achieving a test accuracy of 83.3% and demonstrating effective learning even from limited high-quality data. To enhance interpretability, we integrated eXplainable Artificial Intelligence (XAI) using SHapley Additive exPlanations (SHAP). This analysis revealed critical feature importance distributions, identifying that security-critical tokens are often overlooked in standard development patterns, thus validating the model’s decision-making process for developers.

Document Type

Restricted Access

Submission Type

Thesis

Share

COinS