Original title:
Automatické zpracování obsahu dokumentu
Translated title:
Automated Processing of Table of Contents
Authors:
Shelest, Oleksii ; Materna, Zdeněk (referee) ; Kohút, Jan (advisor) Document type: Bachelor's theses
Year:
2026
Language:
cze Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[cze][eng]
Součástí digitalizačního procesu dokumentů je zpracování jejich struktury. Jednou z hlavních struktur, která umožňuje navigaci v dokumentu je obsah. Současné automatické zpracování obsahu se opírá především o pravidlový a hybridní přístupy, které však narážejí na problémy spojené s rozmanitostí formátování v jednotlivých knihách. Cílem této práce je navrhnout architekturu pro automatické zpracování obsahu digitalizovaných knih, která umožňuje extrakci dvěma odlišnými způsoby – metodou založenou na detekci objektů a OCR, a moderním přístupem využívajícím velké jazykové modely – a oba tyto přístupy mezi sebou porovnat. Navržená architektura doplňuje obě metody o mechanismus propojení extrahovaného obsahu se skutečnou strukturou knihy. Experimenty provedené na anotované datové sadě ukázaly, že každý z těchto dvou přístupů má své silné stránky. Metoda, založena na detekci objektů a OCR, zajišťuje vyšší geometrickou přesnost: průměrná hodnota F1 – 0.887. Zatímco metoda založená na jazykových modelech lépe zachycuje hierarchické struktury a lépe zpracovává text na stránce: průměrná hodnota TED – 0.1440, průměrná hodnota CER – 0.0968.
Processing the structure of documents is an important part of the digitization process. One of the main structures that enables navigation within a document is the table of contents. Current automated table-of-contents processing relies primarily on rule-based and hybrid approaches, which, however, face problems related to the diversity of formatting across individual books. The goal of this work is to design an architecture for the automatic processing of the table of contents in digitized books that enables extraction in two distinct ways – a method based on object detection and OCR, and a modern approach using large language models – and to compare these two approaches. The proposed architecture supplements both methods with a mechanism for linking the extracted table of contents to the actual structure of the book. Experiments performed on an annotated dataset showed that each of these two approaches has its strengths. The method based on object detection and OCR ensures higher geometric accuracy: average F1 score – 0.887. Meanwhile, the method based on language models better captures hierarchical structures and processes text on the page more effectively: average TED score – 0.1440, average CER score – 0.0968.
Keywords:
book digitization; Content processing; data extraction; hierarchical document structure; large language models; machine learning; OCR; YOLOv11; digitalizace knih; extrakce dat; hierarchická struktura dokumentu; OCR; strojové učení; velké jazykové modely; YOLOv11; Zpracování obsahu
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/258832