Original title:
Detekce závadných webů podle vizuální podobnosti
Translated title:
Detection of Malicious Websites based on Visual Similarity
Authors:
Dacík, Ondřej ; Žádník, Martin (referee) ; Hranický, Radek (advisor) Document type: Master’s theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[cze][eng]
Cílem této práce je návrh a implementace systému, který na základě vizuální podoby webové stránky odhadne, zda je závadná. Práce se zaměřuje na stránky phishingové, tedy takové, které napodobují legitimní služby ve snaze získat přihlašovací údaje uživatelů. Pro účely trénování a testování modelů byl v rámci práce vytvořen nástroj pro automatizovaný sběr dat, pomocí něhož byla sestavena datová sada obsahující 12 464 benigních a 10 558 phishingových stránek. Architektura systému je navržena v modelu klient-server. Klientskou část tvoří rozšíření do webového prohlížeče, které se dotazuje serveru na závadnost navštěvovaných webů. Analytická část serveru se skládá ze čtyř nezávislých modulů. Jádrem systému je grafová neuronová síť analyzující sémantické RDF grafy vykreslených webových stránek. Tento přístup doplňuje analýza celkového vzhledu stránky pomocí konvoluční neuronové sítě a ověřování identity služby na základě loga a faviconu. Dílčí predikce jsou následně agregovány meta-modelem, který provádí finální rozhodnutí. Evaluace systému prokázala vysokou účinnost zvoleného multi-modálního přístupu, který na nezávislé datové sadě dosáhl skóre F1 0,9326 a překonal tak existující nástroje.
The goal of this thesis is to design and implement a system that estimates whether a website is malicious based on its visual appearance. The thesis focuses on phishing sites, i.e., sites that mimic legitimate services in an attempt to obtain users’ login credentials. For the purposes of training and testing the models, a tool for automated data collection was developed as part of this work, which was used to compile a dataset containing 12,464 benign and 10,558 phishing pages. The system’s architecture is designed as a client-server model. The client-side consists of a web browser extension that queries the server about the maliciousness of visited websites. The analytical part of the server consists of four independent modules. The core of the system is a graph neural network that analyzes semantic RDF graphs of rendered web pages. This approach is complemented by an analysis of the page’s overall appearance using a convolutional neural network and verification of the service’s identity based on its logo and favicon. The individual predictions are then aggregated by a meta-model, which makes the final decision. Evaluation of the system demonstrated the high effectiveness of the chosen multimodal approach, which achieved an F1 score of 0.9326 on an independent dataset and outperformed existing tools.
Keywords:
cybersecurity; machine learning; phishing; visual analysis; web scraping; automatizovaný sběr dat; kybernetická bezpečnost; phishing; strojové učení; vizuální analýza
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260293