Original title:
Inteligentní vyhledávání informací ve velkých kolekcích dokumentů
Translated title:
Intelligent information retrieval from large document collections
Authors:
Ondrušek, Tomáš ; Kohút, Jan (referee) ; Hradiš, Michal (advisor) Document type: Master’s theses
Year:
2026
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Táto diplomová práca sa zameriava na inteligentné vyhľadávanie informácií vo veľkých kolekciách dokumentov pomocou agentických architektúr vyhľadávania. Cieľom je navrhnúť a vyhodnotiť systémy inteligentného vyhľadávacieho systému podobného RAG systému bez generatívnej zložky so zameraním na efektívnosť a interpretovateľnosť. Riešenie integruje nástroje ako LangChain a LangGraph s vektorovými databázami, ako je Weaviate, a jazykovými modelmi z Ollama a OpenAI. Viacjazyčný hodnotiaci dataset so zameraním na český jazyk bol vytvorený na meranie kvality vyhľadávania, pričom výsledky poukazujú, ako štruktúra grafu a návrh systému ovplyvňujú presnosť a efektivitu a poskytujú usmernenia pre tvorbu odolných viacjazyčných vyhľadávacích systémov. Výsledky ukazujú, že zatiaľ čo hybridné vyhľadávanie a prehodnocovanie (reranking) výrazne zlepšujú výkon na viacjazyčných datasetoch, vysoko komplexné agentické architektúry a hĺbkové vyhľadávanie (deep search) automaticky nezaručujú lepšiu kvalitu vyhľadávania v porovnaní s jednoduchšími prístupmi, pričom často len zvyšujú latenciu bez primeraných ziskov. Efektivita komplexných systémov navyše silne závisí od schopnosti použitého jazykového modelu robiť spoľahlivé čiastkové rozhodnutia.
This thesis focuses on intelligent information retrieval from large document collections using agentic retrieval architectures. The goal is to design and evaluate an intelligent search system similar to retrieval-augmented systems without generative components, emphasizing effectiveness and interpretability. The solution integrates LangChain and LangGraph with vector databases such as Weaviate and language models from Ollama and OpenAI. A custom multilingual benchmark dataset, with a specific focus on the Czech language, among others, supports the evaluation of retrieval quality by showing how graph structure and system design influence accuracy and efficiency, and provides guidelines for building robust multilingual retrieval systems. The results demonstrate that while hybrid retrieval and cross-encoder reranking significantly improve performance on multilingual datasets, highly complex agentic pipelines and deep search architectures do not automatically guarantee better retrieval quality compared to simpler baselines, often increasing latency without proportionate gains. Additionally, the effectiveness of end-to-end agentic workflows is shown to depend heavily on the underlying language model's capability to make robust and reliable intermediate decisions.
Keywords:
agentické architektúry vyhľadávania; LangChain; LangGraph; sémantické vyhľadávanie; vektorové databázy; veľké jazykové modely; viacjazyčné hodnotenie; vyhľadávanie informácií; agentic retrieval architectures; information retrieval; LangChain; LangGraph; large language models; multilingual evaluation; semantic search; vector databases
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/260149