Identification Of An Email Author Using Machine Learning And Natural Language Processing

En cours de chargement...
Vignette d'image

Date

Nom de la revue

ISSN de la revue

Titre du volume

Éditeur

Résumé

Emails are often used in cybercrime today, so it is important to verify the identity of the email author. This paper proposes different Machine Learning models: Naive Bayes (NB), Logistic Regression (LR) and Sup- port Vector Machine (SVM), to solve the problem of anonymous email author attribution. The main task is to find the author of an anonymous email among the many suspected targets, to verify if an email was in fact written by the sender. In this project, the models were trained and tested using the same email dataset where we analyze writing style, vocabulary usage, and tex- tual patterns, plus date and time information using supervised machine learning in addition to natural language processing techniques, with the intention of combining multiple factors to identify the author. The tests proved that the accuracy varies from one model to another indicating that some are better than others. In email authorship verifica- tion experiments, usually the average accuracy reaches 89.9%. while our model’s accuracy rate to a well distributed dataset is 89,06% for SVM , 81,77% for Naive Bays and 88.02% for Logistic Regression. And with a not so well distributed dataset the accuracy rate is 91,40% for SVM model, 76,74% for Naive Bays and for Logistic Regression 92.29%. Proving that a good identification system relies on two aspects, the model chosen for the task and the dataset veracity.

Description

Citation

Collections