Original title:
Řečové velké jazykové modely pro automatické rozpoznávání řeči více mluvčích
Translated title:
Speech-augmented large language models for multi-talker automatic speech recognition
Authors:
Vaculík, Martin ; Polok, Alexander (referee) ; Sedláček, Šimon (advisor) Document type: Bachelor's theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Moderní modely pro automatické zpracování řeči dosahují vysoké přesnosti v prostředí, kde mluví pouze jeden mluvčí, avšak vícemluvčí scénáře stále představují výzvu. Tato práce zkoumá, zda systémy postavené na velkých předtrénovaných řečových LLM dokážou tento problém řešit, a analyzuje různé metody rozšíření existujících modelů pro jednomluvčí scénáře. Nejprve trénujeme vlastní řečový LLM Mac-ASR s přibližně 2 miliardami parametrů, který dosahuje slušných výsledků na různých benchmarcích, avšak selhává při adaptaci na konverzační vícemluvčí data. Pro další experimenty s vícemluvčí řečí proto přecházíme na veřejně dostupný model Qwen3-ASR-0.6B. Na zvoleném dvoumluvčím datasetu porovnáváme několik strategií podmínění jazykového modelu, včetně speaker cache na úrovni promptu, target-speaker conditioningu pomocí cross-attention i promptového podmínění a dalších přístupů. Zjišťujeme, že v našem nastavení je pro model složitější přepisovat pouze cílového mluvčího než přepisovat všechny mluvčí současně, kde se identita rozlišuje pomocí speciálních speaker tokenů. Po rozšíření trénovacích dat pomocí volně dostupných datasetů AMI a LibriMix dosahuje konfigurace speaker cache hodnoty 22% cpWER na Fisheru, čímž se přibližuje systému DiCoW. Výsledky na AMI však zůstávají slabé a otevřenou otázkou nadále zůstává, zda lze tyto metody efektivně rozšířit na scénáře s více než dvěma mluvčími.
Modern speech recognition models achieve high accuracy in single-speaker settings. However, multi-speaker scenarios remain challenging. This thesis investigates whether systems built on top of large pretrained speech LLMs can address this problem and explores various methods for extending existing single-speaker models to multi-speaker settings. First, we train our own speech LLM, Mac-ASR, with roughly 2 billion parameters. While the model achieves solid results across several benchmarks, it fails to adapt to conversational multi-speaker data. For all subsequent multi-speaker experiments, we therefore transition to the publicly available Qwen3-ASR-0.6B model. On the Fisher corpus, we compare several conditioning strategies, including prompt-level speaker cache, target-speaker conditioning via cross-attention, prompt-level conditioning, and others. We find that, in our setting, it is more difficult for the model to transcribe only the target speaker than to transcribe all speakers jointly distinguishing them by special speaker tokens. After extending the training data with the AMI, and LibriMix corpora, the speaker cache configuration achieves 22% cpWER on Fisher, closing the performance gap with the DiCoW system. However, the results on AMI remain poor, and it is still an open question whether these methods can be effectively extended to scenarios involving more than two speakers.
Keywords:
automatické rozpoznávání řeči; hluboké učení; multimodální systémy; strojové učení; systémy bez diarizace; učení s učitelem; velké jazykové modely; vícemluvčí automatické rozpoznávání řeči; Whisper; zarovnané reprezentační prostory; řečové velké jazykové modely; aligned representation spaces; automatic speech recognition; deep learning; diarization-free systems; large language models; machine learning; multi-speaker automatic speech recognition; multimodal systems; speech large language models; supervised learning; whisper
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258876