Original title:
Webová aplikace pro efektivní označování témat v kolekcích textu
Translated title:
Web Application for Efficient Topic Labeling in Text Collections
Authors:
Sucharda, Marek ; Bartl, Vojtěch (referee) ; Hradiš, Michal (advisor) Document type: Bachelor's theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[cze][eng]
Manuální kódování dokumentů je jedním z nejvíce časově náročných kroků kvalitativní i kvantitativní analýzy. Tato práce popisuje návrh a implementaci části webové aplikace s názvem semANT. Konkrétní implementovanou částí je modul zobrazující dokument a anotační editor s funkcemi podporovanými umělou inteligencí. V editoru lze kódovat, což je proces manuálního označování v textu a přiřazování štítků (kódů, tagů) těmto pasážím. Implementovány jsou dvě klíčové funkce využívající metod AI. První je automatický návrh anotací za pomoci LLM, kdy systém uživateli navrhne vhodné pasáže a přiřadí jim relevantní tag. Druhou funkcí je návrh nejvhodnějších tagů pro manuálně označenou část textu za pomoci zero-shot klasifikace modelem mDeBERTa, který vyhodnotí nejvíce relevantní tagy, a ty poté zobrazí uživatelské rozhraní. Logika těchto funkcí je oddělena do separátní Python knihovny Topicer. Uživatelské testování ukázalo zkrácení doby kódování při využití automatických návrhů. Testování shody manuálního kódování s klasifikátorem dosáhlo shody v rozmezí 56—85%. Výsledky naznačují, že integrace těchto funkcí do anotačního editoru má potenciál snížit časovou náročnost kódování v kvalitativní analýze.
Manual document annotation is one of the most time-consuming steps in both qualitative and quantitative analysis. This thesis describes the design and implementation of a component of a web application called semANT. Specifically, the implemented component consists of a document viewer and an annotation editor with AI-powered features. The editor allows for coding, which is the process of manually marking passages in the text and assigning labels (codes, tags) to them. Two key functions utilizing AI methods have been implemented. The first is automatic annotation suggestion using LLM, where the system suggests suitable passages to the user and assigns relevant tags to them. The second feature is the suggestion of the most suitable tags for a manually marked section of text using zero-shot classification with the mDeBERTa model, which evaluates the most relevant tags and then displays them in the user interface. The logic of these features is separated into a standalone Python library called Topicer. User testing showed a reduction in coding time when using automatic suggestions. Testing the agreement between manual coding and the classifier achieved an agreement rate of 56—85%. The results suggest that integrating these features into the annotation editor has the potential to reduce the time required for coding in qualitative analysis.
Keywords:
annotation editor; digitised documents; FastAPI; large language models; mDeBERTa; natural language processing; NLI; OCR; qualitative text analysis; text coding; Vue.js; Weaviate; zero-shot classification; anotační editor; digitalizované dokumenty; FastAPI; kvalitativní analýza textu; kódování textu; mDeBERTa; NLI; OCR; velké jazykové modely; Vue.js; Weaviate; zero-shot klasifikace; zpracování přirozeného jazyka
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258908