Web attacks are a significant threat to online security. Traditional WAFs (Rule-based) may struggle to detect certain attacks due to sophisticated obfuscation techniques used by attackers.
Example: Consider a simple XSS payload:
<script>alert('XSS Attack!');</script>
Now, observe its obfuscated version, which might evade rule-based WAFs:
%253Cscript%253Ealert%2528%2527XSS%2520Attack%2521%2527%2529%253B%253C%252Fscript%253E
However, Natural Language Processing (NLP) techniques have shown promise in resisting such obfuscation techniques. In this project, we leverage the power of NLP to enhance web security.
- Utilizes simple NLP techniques: Bag of Words and TF-IDF.
- Implements multiple machine learning algorithms including Multinomial Naive Bayes, Logistic Regression, Support Vector Machines (SVM), Random Forests, etc.
- Includes code for utilizing fastText, a library for efficient learning of word representations and sentence classification.
The dataset utilized is CSIC 2010. It comprises of sqli, xss, command injection and path traversal malicious payloads as well as normal http requests.
Link: https://www.kaggle.com/datasets/ispangler/csic-2010-web-application-attacks