Original title:
Detekce vozidel z ptačí perspektivy na leteckých snímcích pomocí hlubokého učení
Translated title:
Detection of vehicles from a bird's eye view on aerial images using deep learning
Authors:
Marušinec, Matej ; Strýček, Šimon (referee) ; Dubovec, Pavol (advisor) Document type: Bachelor's theses
Year:
2026
Language:
slo Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[slo][eng]
Cieľom tejto práce je návrh a implementácia pokročilého systému na detekciu vozidiel z aerálnych snímok s využitím samo-riadených vizuálnych reprezentácií. Riešenie vychádza z inkrementálneho prístupu, ktorý zahŕňa počiatočný vývoj architektúry „from scratch“ pre analýzu limitov detekcie malých objektov a následnú integráciu príznakov modelu DINO (Self-Distillation with No Labels). Tento hybridný prístup spája priestorovú presnosť konvolučného backbone-u s hlbokým sémantickým kontextom Vision Transformers, čo zvyšuje robustnosť detekcie v komplexných scenároch. Na dosiahnutie vysokej stability a geometrickej presnosti lokalizácie je v práci implementovaná kompozitná stratová funkcia spájajúca L1 Loss, CIoU Loss a Distribution Focal Loss (DFL). Trénovanie a evaluácia prebiehali na unifikovaných datasetoch VisDrone, DOTA a UAVDT, pričom systém bol optimalizovaný s ohľadom na variabilitu mierky a hustotu objektov. Výsledkom práce je detekčný softvér schopný vykonávať lokalizáciu s F1 presnosťou takmer 62% a klasifikáciu vozidiel na výškových obrázkoch. Navrhnuté riešenie poskytuje vizuálnu spätnú väzbu vo forme bounding boxov a potvrdzuje efektivitu využitia samo-riadeného učenia pre praktické úlohy monitorovania dopravy a inteligentnej mestskej infraštruktúry.
The aim of this work is to design and implement an advanced system for detecting vehicles in aerial images using self-guided visual representations. The solution is based on an incremental approach that involves initial architecture development “from scratch” to analyze the detection limits of small objects, followed by the integration of features from the DINO (Self-Distillation with No Labels) model. This hybrid approach combines the spatial accuracy of a convolutional backbone with the deep semantic context of Vision Transformers, which enhances the robustness of detection in complex scenarios. To achieve high stability and geometric localization accuracy, the work implements a composite loss function combining L1 Loss, CIoU Loss, and Distribution Focal Loss (DFL). Training and evaluation were performed on the unified datasets VisDrone, DOTA, and UAVDT, with the system optimized to account for scale variability and object density. The result of this work is detection software capable of performing localization with an F1 score of nearly 62% and vehicle classification in aerial images. The proposed solution provides visual feedback in the form of bounding boxes and confirms the effectiveness of using self-supervised learning for practical tasks in traffic monitoring and smart urban infrastructure.
Keywords:
deep learning; DINOv2; drone images; neural networks; object detection; ResNet-50
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258771