Original title:
RAG systémy a jejich vyhodnocování
Translated title:
RAG Systems and Their Evaluation
Authors:
Martinček, Ľuboš ; Materna, Zdeněk (referee) ; Hradiš, Michal (advisor) Document type: Master’s theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Veľké jazykové modely sú náchylné na halucinácie kvôli svojej závislosti od statických trénovacích dát. Generovanie rozšírené o vyhľadávanie (RAG) tento problém zmierňuje ukotvením generovania v dynamicky vyhľadaných dôkazoch. Spoľahlivé hodnotiace benchmarky, najmä pre jazyky iné ako Angličtina, zostávajú vzácne. Táto práca navrhuje a implementuje benchmark na hodnotenie RAG systémov odvodený z OCR spracovaných českých historických dokumentov aplikácie semANT, spolu s dvoma rôznymi RAG systémami. Dataset s 536 vzorkami, pokrývajúci faktografické otázky, otázky vyžadujúce syntézu viacerých zdrojov a inferenčné otázky, bol vytvorený pomocou výberu seed chunkov algoritmom K-Means, obohacovania kontextu a frameworku RAGAS na generovanie testovacích sád, po ktorom nasledovala manuálna revízia. Viaceré konfigurácie RAG, vrátane naivného, inkrementálneho, agentného a adaptívneho multi-query variantu, sú porovnávané pomocou piatich metrík RAGAS: Context Recall, Context Relevance, Faithfulness, Answer Correctness a Answer Relevance. Experimenty ukazujú, že podobné skóre vyhľadávania nezaručuje podobnú kvalitu odpovede. Agentný systém dosahuje najvyššiu hodnotu metrík Answer Correctness a Answer Relevance.
Large Language Models are prone to hallucinations due to their reliance on static training data. Retrieval-Augmented Generation (RAG) mitigates this by grounding generation in dynamically retrieved evidence, yet robust evaluation benchmarks, especially for non-English settings, remain scarce. This thesis designs and implements a RAG evaluation benchmark derived from OCR-processed Czech historical documents from the semANT application, along with two different RAG systems. The 536-sample dataset, spanning factual, multi-source synthesis, and inference questions, was constructed using K-Means seed chunk selection, context enrichment, and the RAGAS testset generation framework, followed by manual review. Multiple RAG configurations, including naive, incremental, agentic, and adaptive multi-query variants, are compared using five RAGAS metrics: Context Recall, Context Relevance, Faithfulness, Answer Correctness, and Answer Relevance. The experiments demonstrate that similar retrieval scores do not guarantee similar answer quality. The agentic system achieves the highest Answer Correctness and Answer Relevance.
Keywords:
agentný RAG; benchmark; RAG; RAGAS; veľké jazykové modely; vyhodnotenie RAG; zodpovedanie otázok; agentic RAG; benchmark; large language models; question answering; RAG; RAG evaluation; RAGAS
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260410