Deep Learning-based Software Vulnerability Severity Prediction Using Natural Language Processing

Paul Teye

Department of ICT Education, University of Education, Winneba, Ghana.

Stephen Opoku Oppong

Department of ICT Education, University of Education, Winneba, Ghana.

Dickson Keddy Wornyo

Information Technology Studies, University of Professional Studies, Accra, Ghana.

Daniel Danso Essel

Department of ICT Education, University of Education, Winneba, Ghana.

Benjamin Ghansah *

Department of ICT Education, University of Education, Winneba, Ghana.

Muhammed Siraj

Department of Computer Science, Academic City University, Accra, Ghana.

*Author to whom correspondence should be addressed.


Abstract

Software vulnerability assessment is essential for prioritising remediation, yet the manual assignment of Common Vulnerability Scoring System metrics is time-consuming, subjective, and prone to error. This study develops and evaluates a design science artefact for predicting vulnerability severity directly from natural-language descriptions. More than 15,000 confirmed vulnerability records were collected from ten open-source projects represented in the CVE-NVD repositories. The descriptions were cleaned through punctuation removal, tokenisation, stopword removal, and Porter stemming. Textual features were represented using n-grams, term frequency-inverse document frequency, and Word2Vec, while Information Gain was applied for feature selection. A deep neural network with three hidden layers was developed for binary classification of vulnerabilities as Severe or Non-Severe. Its performance was compared with Random Forest, Support Vector Machine, Decision Tree, k-nearest neighbours, and Naïve Bayes classifiers using stratified 10-fold cross-validation. The deep neural network achieved average accuracy, precision, and recall values of 71.2%, 68.5%, and 72.0%, respectively, and generally outperformed the conventional classifiers across the reported metrics. However, performance varied among projects, with AUC-ROC values ranging from 0.385 for Gentoo to 0.720 for Windows 7, and the Decision Tree exceeded the deep neural network on the Gentoo dataset. These findings indicate that natural-language descriptions can support automated severity prediction, while also showing that model performance depends on project-specific data characteristics and data availability.

Keywords: Software vulnerability prediction, deep learning, common vulnerability scoring system (cvss), natural language processing


How to Cite

Teye, Paul, Stephen Opoku Oppong, Dickson Keddy Wornyo, Daniel Danso Essel, Benjamin Ghansah, and Muhammed Siraj. 2026. “Deep Learning-Based Software Vulnerability Severity Prediction Using Natural Language Processing”. Asian Journal of Research in Computer Science 19 (8):124-38. https://doi.org/10.9734/ajrcos/2026/v19i8896.

Downloads

Download data is not yet available.