Identification Of An Email Author Using Machine Learning And Natural Language Processing
En cours de chargement...
Date
Nom de la revue
ISSN de la revue
Titre du volume
Éditeur
Résumé
Emails are often used in cybercrime today, so it is important to verify
the identity of the email author. This paper proposes different Machine
Learning models: Naive Bayes (NB), Logistic Regression (LR) and Sup-
port Vector Machine (SVM), to solve the problem of anonymous email
author attribution. The main task is to find the author of an anonymous
email among the many suspected targets, to verify if an email was in fact
written by the sender.
In this project, the models were trained and tested using the same
email dataset where we analyze writing style, vocabulary usage, and tex-
tual patterns, plus date and time information using supervised machine
learning in addition to natural language processing techniques, with the
intention of combining multiple factors to identify the author.
The tests proved that the accuracy varies from one model to another
indicating that some are better than others. In email authorship verifica-
tion experiments, usually the average accuracy reaches 89.9%. while our
model’s accuracy rate to a well distributed dataset is 89,06% for SVM ,
81,77% for Naive Bays and 88.02% for Logistic Regression. And with a
not so well distributed dataset the accuracy rate is 91,40% for SVM model,
76,74% for Naive Bays and for Logistic Regression 92.29%. Proving that
a good identification system relies on two aspects, the model chosen for
the task and the dataset veracity.
