Original title:
Analýza struktury tabulek pomocí multimodálních transformerů
Translated title:
Table Structure Recognition Using Multimodal Transformers
Authors:
Vlach, Vojtěch ; Kišš, Martin (referee) ; Hradiš, Michal (advisor) Document type: Master’s theses
Year:
2025
Language:
eng Publisher:
Vysoké učení technické v Brně. Fakulta informačních technologií Abstract:
[eng][cze]
Tato práce se zabývá rozpoznáním struktury tabulek pro analýzu a rekonstrukci tabulek z vyfocených nebo skenovaných dokumentů. Práce představuje současné metody, rozšiřuje je a jejím výsledkem je systém pro rozpoznání tabulek. Konkrétně je použit systém rozpoznání písma (Optical Character Recognition - OCR) pro detekování a přepis jednotlivých slov. Struktura tabulky je tvořena pomocí matice sousedností reprezentující shlukové vztahy mezi slovy (shluky typu buňka, sloupec, řádek). Představená architektura je tvořena konvolučním předzpracováním, multimodálním transformerem, predikčními hlavami pro každý typ vztahu a algoritmem rekonstrukce tabulky. Architektura je funkční a je porovnatelná s referenční literaturou na datasetu PubTables-1M. Natrénované modely jsou také doladěné na novém datasetu HerritageTabNet s pozitivní změnou na obou datasetech.
This thesis introduces the topic of Table Structure Recognition (TSR), which is used to analyze and reconstruct scanned tables. Current methods are introduced and expanded upon to create a Table Structure Recognition system. First, the Optical Character Recognition (OCR) system detects and transcribes words. The table structure is created using adjacency matrices representing word relation classes (same cell, column clusters, row clusters). The proposed architecture consists of a CNN backbone, a multimodal decoder transformer, class-wise prediction heads, and post-processing table reconstruction algorithm. The architecture is proven to work and is comparable with refference literature on the PubTables-1M dataset. The trained models are also fine-tuned on a custom HerritageTabNet dataset with positive improvement on both datasets.
Keywords:
analýza struktury dokumentů; detekce tabulek; hluboké učení; multimodální transformer; optické rozpoznávání písma; počítačové vidění; predikce vztahů slov; Rozpozání struktury tabulky; Computer vision; Deep learning; Document Structure Analysis; Multimodal transformer; Optical Character Recognition; Table detection; Table Structure Recognition; Word relation prediction
Institution: Brno University of Technology
(web)
Document availability information: Fulltext is available in the Brno University of Technology Digital Library. Original record: http://hdl.handle.net/11012/255126