Original title:
Návrh akcelerátoru konvolučních sítí
Authors:
Drlíčková, Alena ; Klhůfek, Jan (referee) ; Mrázek, Vojtěch (advisor) Document type: Master’s theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[cze][eng]
Tato diplomová práce se zabývá návrhem a implementací hardwarového akcelerátoru konvolučních neuronových sítí. Navržený akcelerátor je založen na systolickém poli a umožňuje provádět inferenci různých architektur CNN bez nutnosti modifikace hardwarového návrhu. Chování akcelerátoru je řízeno jednoduchou instrukční sadou, kterou generuje softwarový kompilátor na základě zadané architektury sítě. Kompilátor zároveň zajišťuje post-training kvantizaci vah a jejich převod do formátu vhodného pro celočíselnou aritmetiku akcelerátoru. Transformace konvoluční operace na maticové násobení je realizována přímo v hardwaru pomocí jednotky správy paměti provádějící im2col transformaci za chodu. Akcelerátor byl implementován ve VHDL a syntetizován pro čip XC7Z020. Správnost výpočtů byla ověřena simulací v nástroji Vivado XSim a porovnáním s referenční softwarovou implementací kvantizované sítě. Funkčnost byla demonstrována na architektuře LeNet-5. Syntéza potvrdila splnění časových požadavků při pracovní frekvenci 100 MHz s využitím přibližně 27 \% dostupných LUT cílového čipu.
This master's thesis presents the design and implementation of a hardware accelerator for convolutional neural networks. The proposed accelerator is based on a systolic array and supports inference of various CNN architectures without requiring any modification of the hardware design. The behavior of the accelerator is controlled by a simple instruction set generated by a software compiler according to the given network architecture. The compiler also performs post-training quantization of the network weights and converts them into a format suitable for the integer arithmetic used by the accelerator. The transformation of the convolution operation into matrix multiplication is performed directly in hardware by a memory management unit that carries out the im2col transformation on the fly. The accelerator was implemented in VHDL and synthesized for the XC7Z020 chip. The correctness of the computations was verified by simulation in Vivado XSim and by comparison with a reference software implementation of the quantized network. The functionality was demonstrated on the LeNet-5 architecture. Synthesis confirmed that the timing requirements are met at an operating frequency of 100 MHz, utilizing approximately 27\% of the available LUTs on the target chip.
Keywords:
Convolutional neural network; FPGA; hardware acceleration; im2col; inference; matrix multiplication; quantization; systolic array; VHDL; FPGA; hardwarová akcelerace; im2col; inference; Konvoluční neuronová síť; kvantizace; maticové násobení; systolické pole; VHDL
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260516