Original title:
Rozšíření nástroje DomainRadar pro detekci škodlivých doménových jmen na základě obsahu webové stránky
Translated title:
DomainRadar Extension for Malicious Domain Name Detection using Web Page Contents
Authors:
Mazhirinov, Alisher ; Setinský, Jiří (referee) ; Hranický, Radek (advisor) Document type: Bachelor's theses
Year:
2025
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Tato bakalářská práce pojednává o metodách detekce phishingových, škodlivých a benigních domén na základě analýzy textového obsahu jejich webových stránek. Hlavní důraz je kladen na využití TF-IDF (Term Frequency - Inverse Document Frequency), metody, která umožňuje určit význam slov v dokumentech na základě jejich frekvence v textu a inverzní frekvence v celém datovém korpusu. Studie ukazuje, že obsah webových stránek obsahuje užitečné textové prvky, které lze použít k automatické klasifikaci domén. Na základě těchto vlastností byly vyvinuty a natrénovány modely klasifikátorů, které dosáhly přesností téměř 90% oba. Použití TF-IDF v kombinaci s metodami strojového učení umožňuje efektivně identifikovat phishing a škodlivé zdroje a také je odlišit od bezpečných domén. Výsledky potvrzují vysoký přínos analýzy textu při řešení problémů kybernetické bezpečnosti a lze je využít k vytvoření automatizovaných systémů pro monitorování a ochranu uživatelů na internetu.
This thesis discusses methods for detecting phishing, malicious, and benign domains based on the analysis of the text content of their webpages. The main focus is on the use of TF-IDF (Term Frequency - Inverse Document Frequency), a method that allows determining the significance of words in documents based on their frequency in the text and inverse frequency in the entire data corpus. The study shows that the content of web pages contains useful text features that can be used to automatically classify domains. Based on these features, two classifier models were developed, trained and achieved accuracies of almost 90% for both The use of TF-IDF in combination with machine learning methods allows you to effectively identify phishing and malicious resources, as well as distinguish them from benign domains. The results confirm the high benefit of text analysis in solving cybersecurity problems and can be used to create automated systems for monitoring and protecting users on the internet.
Keywords:
analýza obsahu webu; analýza webových stránek; automatický klasifikátor; bezpečné weby; datová sada.; detekce malwaru; detekce phishingu; extrakce znaků; klasifikace dokumentů; klasifikace domén; klasifikace textu; strojové učení; TF-IDF; automatic classifier; benign websites; data set.; document classification; domain classification; feature extraction; machine learning; malware detection; phishing detection; text classification; TF-IDF; web content analysis; web page analysis
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/254346