Original title:
Predikce strukturálních charakteristik mozku z řeči pomocí akusticko-lingvistických biomarkerů
Translated title:
Prediction of brain structural characteristics from speech using acoustic-linguistic biomarkers
Authors:
Vengerová, Veronika ; Mekyska, Jiří (referee) ; Kováč, Daniel (advisor) Document type: Master’s theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta elektrotechniky a komunikačních technologií Abstract:
[eng][cze]
V priebehu posledných niekoľkých desaťročí sa zvýšil počet ľudí s diagnostikovanými neurodegeneratívnymi ochoreniami. Ak by sa digitálne biomarkery a algoritmy strojového učenia dali využiť na predpovedanie štrukturálnych charakteristík mozgu, výrazne by to znížilo záťaž pre pacientov. V tejto práci sa skúma možnosť využitia rečových biomarkerov na predpovedanie kortikálnej hrúbky v jednotlivých oblastiach mozgu. Z nahrávok spontánnej reči a ich prepisov sa extrahujú akustické a lingvistické príznaky. Na extrakciu akustických príznakov bola použitá knižnica Disvoice s dodatočnými úpravami. Pre lingvistické príznaky bolo použité označovanie slovných druhov (Part-of-Speech) z nástroja stanza. Pre neinterpretovateľné príznaky boli extrahované sémantické vektory z rôznych vrstiev modelov BERT a CzeGPT-2. Testovali sa tri typy modelov: modely trénované iba na interpretovateľných príznakoch, modely trénované iba na neinterpretovateľných príznakoch a modely trénované na interpretovateľných aj neinterpretovateľných príznakoch s použitím včasnej alebo neskorej fúzie. V tejto práci boli testované architektúry neurónových sietí, lineárnej regresie a XGBoost. Na vyhodnotenie najlepších modelov bola použitá stratifikovaná krížová validácia (Stratified K-Fold). Po získaní počiatočných výsledkov boli modely s najlepším výkonom následne testované pomocou krížovej validácie typu leave-one-out na piatich oblastiach mozgu. Spomedzi hodnotených oblastí dosiahli modely najlepšie výsledky pri predikcii insuly v ľavej aj pravej hemisfére. Najlepšie výsledky modelu včasnej fúzie sa podarilo dosiahnuť pomocou regresného modelu XGBoost, ktorý bol natrénovaný na vybraných príznakoch odvodených z analýzy hlavných komponentov (PCA) z 12. vrstvy modelu CzeGPT-2, interpretovateľných biomarkeroch a demografických príznakoch, vrátane veku, vzdelania a pohlavia. Tento model dosiahol priemernú percentuálnu chybu (MAPE) 0,0403, koeficientu determinácie (R2) 0,1105 a mieru chyby odhadu (EER) 0,1188. V prípade neskorej fúzie modely neprekonali modely trénované iba na interpretovateľných príznakoch. Najlepšie výsledky dosiahol model trénovaný s použitím výberu interpretovateľných príznakov, veku a vzdelania, s MAPE 0,0383, R2 0,1444 a EER 0,0991.
In the past few decades, the number of people diagnosed with neurodegenerative diseases has increased. If digital biomarkers and machine learning algorithms could be used to predict structural brain characteristics, it would significantly reduce the strain on patients. In this thesis, the feasibility of using speech biomarkers to predict the cortical thickness of brain regions is investigated. Acoustic and linguistic features are extracted from recordings of spontaneous speech and their transcriptions. Disvoice library with additional changes was used to extract acoustic features. For linguistic features, the Part-of-speech tagging from the stanza toolkit was used. For the non-interpretable features, the semantic vectors were extracted from different layers of the BERT and CzeGPT-2 models. Three types of models were tested: models trained on interpretable features only, models trained on non-interpretable features only, and models trained on both interpretable and non-interpretable features using early or late fusion. Neural networks, linear regression, and XGBoost were the architectures tested in this thesis. To evaluate the best models, Stratified K-Fold cross-validation was used. After the initial results, the best-performing models were then tested using Leave-one-out cross-validation on five brain regions. Among the evaluated regions, the models achieved the best results when predicting the insula in both the left and right hemispheres. The best results for the early fusion model were achieved using an XGBoost regressor trained on selected principal component analysis-derived features from the 12th layer of the CzeGPT-2 model, interpretable biomarkers, and demographic features including age, education, and gender. This model achieved a mean percentage error (MAPE) of 0.0403, R-squared (R2) of 0.1105, and estimation error rate (EER) of 0.1188. For late fusion, the models did not outperform those trained only on interpretable features. The best results were achieved by a model trained using a selection of interpretable features, age, and education, with MAPE of 0.0383, R2 of 0.1444, and EER of 0.0991.
Keywords:
fúzia; interpretovateľné parametre; kortikálna hrúbka; MCI; regresia; rečové biomarkery; sémantické vektory; cortical thickness; fusion; interpretable features; MCI; regression; semantic vectors; speech biomarkers
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/259567