Original title:
Generativní neuronová síť pro tvorbu syntetických hudebních frází
Translated title:
Generative Neural Network for Synthetic Musical Phrase Creation
Authors:
Hofr, Šimon ; Přinosil, Jiří (referee) ; Říha, Kamil (advisor) Document type: Master’s theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta elektrotechniky a komunikačních technologií Abstract:
[cze][eng]
Tato práce se zabývá generováním syntetických hudebních frází pomocí generativních neuronových sítí (GAN) a především otázkou, jak lze kvalitu takto vytvořeného materiálu měřit. Cílem bylo navrhnout a implementovat model pro symbolickou reprezentaci hudby a doplnit jej opakovatelným postupem hodnocení výstupů. Navržená architektura je plně konvoluční Wassersteinův GAN s penalizací gradientu, pracující nad piano-roll mřížkou o rozměrech 64 x 64, který nevyužívá žádné doplňkové stabilizační techniky. Pro trénink byl zvolen Lakh Pianoroll Dataset, z něhož byly vybrány melodické nástrojové stopy, monofonizovány algoritmem skyline a nakrájeny na čtyřtaktové fráze. Hodnocení kvality kombinuje čtyři nezávislé pohledy, jimiž jsou vzdálenost rozdělení hudebních rysů, věrohodnost pod nezávislým referenčním modelem, klasifikátorový dvouvýběrový test a ověření, že model trénovací data nekopíruje. Výsledný model byl porovnán s rekurentní sítí typu LSTM a s algoritmickým generátorem.
This thesis deals with the generation of synthetic musical phrases using generative adversarial networks (GANs) and above all with the question of how the quality of such material can be measured. The aim was to design and implement a model for the symbolic representation of music and to complement it with a repeatable procedure for evaluating its outputs. The proposed architecture is a fully convolutional Wasserstein GAN with gradient penalty operating on a 64 x 64 piano-roll grid, which employs no supplementary stabilisation techniques. The Lakh Pianoroll Dataset was chosen for training, from which melodic instrument tracks were selected, monophonised by the skyline algorithm and cut into four-bar phrases. The quality assessment combines four independent views, namely the distance between distributions of musical features, the likelihood under an independent reference model, a classifier two-sample test and verification that the model does not copy the training data. The resulting model was compared with a recurrent LSTM network and with an algorithmic generator.
Keywords:
CNN; GAN; gradient penalty; Keras; Lakh; LSTM; MIDI; music generation; neural networks; Python; RNN; TensorFlow; WGAN-GP; CNN; GAN; generování hudby; Keras; Lakh; LSTM; MIDI; neuronové sítě; penalizace gradientu; Python; RNN; TensorFlow; WGAN-GP
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260652