Original title:
Detekce a mitigace otravy dat v klasifikátorech strojového učení
Translated title:
Detection and Mitigation of Data Poisoning in Machine Learning Classifiers
Authors:
Špác, Petr ; Reš, Jakub (referee) ; Firc, Anton (advisor) Document type: Bachelor's theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Tato práce se zabývá zranitelností metody podpůrných vektorů (SVM) vůči otravě dat v kontextu klasifikace malwaru. Cílem práce bylo tyto útoky analyzovat a navrhnout proti nim vhodnou obranu. Práce se zaměřuje na čištění trénovacích dat od škodlivých vzorků a poskytuje návrh a implementaci filtrační metody L2-RONI. Navržená metoda je experimentálně vyhodnocena na datasetech Drebin-215 a LAMDA a proti třem útokům: random label flipping, furthest-first flipping a backdoor. Výsledky prokazují, že L2-RONI úspěšně kombinuje rychlost metody L2 s přesností algoritmu RONI, čímž zrychluje proces čištění a udržuje vysokou obranyschopnost proti furthest-first útoku i při vyšší míře otravy.
This work addresses the vulnerability of the Support Vector Machine (SVM) to data poisoning in the context of malware classification. The goal of the thesis was to analyze these attacks and propose a suitable defense. The work focuses on cleansing training data of malicious samples and provides the design and implementation of the L2-RONI filtering method. The proposed method is experimentally evaluated on the Drebin-215 and LAMDA datasets against three attacks: random label flipping, furthest-first flipping, and backdoor. The results demonstrate that L2-RONI successfully combines the speed of the L2 method with the accuracy of the RONI algorithm, thereby accelerating the cleansing process and maintaining high resilience against the furthest-first attack, even at a higher poisoning rate.
Keywords:
adversariální strojové učení; detekce; klasifikace; malware; mitigace; Otrava dat; počítačová bezpečnost; Python; sanitizace; SVM; zadní vrátka; záměna tříd; adversarial machine learning; backdoor; classification; computer security; Data poisoning; detection; label flipping; malware; mitigation; Python; sanitation; SVM
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258760