Original title:
Metody analýzy síťového provozu pro detekci malware
Authors:
Fukala, Jakub ; Burgetová, Ivana (referee) ; Ryšavý, Ondřej (advisor) Document type: Bachelor's theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[cze][eng]
Práce porovnává pět přístupů k detekci malware v TLS-šifrovaném síťovém provozu: signa- turní detekci, heuristickou detekci, Random Forest, Isolation Forest a banku autoencoderů. Všechny metody jsou vyhodnoceny na společné datové sadě 151 014 vzorků sestavené z 51 rodin malware analyzovaných v sandboxu Triage a z benigního provozu ze tří nezávislých zdrojů, se stejným rozdělením na trénovací a testovací část, hodnocením na úrovni vzorků a stejnými metrikami. Nejlepších výsledků dosahuje Random Forest, který při 1 % falešných poplachů zachytí 98,39 až 99,21 % vzorků malware, zatímco ostatní přístupy výrazně zao- stávají. Práce dále ukazuje, že signaturní detekce a Random Forest se vzájemně doplňují, zatímco heuristika a Isolation Forest selhávají u rodin, jejichž TLS profil se překrývá s be- nigním provozem. Výsledky jsou proto nutné číst s ohledem na rozdíl mezi sandboxovým a reálným prostředím, který omezuje jejich přímou přenositelnost mimo použitou datovou sadu.
This thesis compares five approaches to malware detection in TLS-encrypted network traf- fic: signature-based detection, heuristic detection, Random Forest, Isolation Forest, and a bank of autoencoders. All methods are evaluated on a common dataset of 151 014 sam- ples built from 51 malware families analyzed in the Triage sandbox and benign traffic from three independent sources, using the same train/test split, sample-level evaluation, and metrics. Random Forest achieves the strongest results, detecting 98.39 % (core) to 99.21 % (operational) of malware samples at a 1 % false positive rate, while the remaining appro- aches lag substantially behind. The thesis also shows that signature-based detection and Random Forest are complementary, whereas heuristic detection and Isolation Forest fail on families whose TLS profile overlaps with benign traffic. The results are interpreted with explicit regard to sandbox-induced domain bias, which limits direct generalization beyond the evaluated dataset.
Keywords:
autoencoder; detection cascade; encrypted traffic; heuris- tic detection; Isolation Forest; JA3; JA4; machine learning; Malware detection; Random Forest; signature-based detection; TLS; autoencoder; Detekce malware; heuristická detekce; Isolation Forest; JA3; JA4; kaskádové řazení; Random Forest; signaturní detekce; strojové učení; TLS; šifrovaný provoz
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258909