Original title:
Word embeddings and their meaning
Translated title:
Word embeddings and their meaning
Authors:
LucieDvořáčková ; MichalČerný (advisor) ; JaromírKukal (referee) ; DavidHartman (referee) Document type: Doctoral theses
Year:
2026
Language:
eng Publisher:
Vysoká škola ekonomická v Praze Abstract:
[eng][cze] Embedding models represent techniques that convert texts into numerical vectors in a low-dimensional space, capturing the semantic and syntactic properties of the text. They were developed out of the need to represent textual data numerically so that it can be efficiently processed by machine learning algorithms. Practical applications of embedding models include serving as inputs for classification and predictive models, machine translation, information retrieval, and other applications where understanding the meaning and structure of text is essential. In this dissertation, I focus on three tasks that expand the practical use of embeddings. The first part is devoted to predicting the potential citation impact of scientific articles; using embeddings derived from abstracts, I build models that estimate whether an article may achieve high citation counts in the future. The second part introduces a novel method for explaining embedding models, which precisely determines how individual words contribute to the prediction results, thereby providing a clear interpretation of the model. The third part addresses the explanation of semantic similarity, demonstrating how distributional statistics combined with lexical knowledge enable a better understanding of the relationships between words in the embedding space. The results of this work have the potential to improve the accuracy and efficiency of classification and predictive models in scientific analytics, while also contributing to a deeper understanding of the mechanisms by which text models process natural language.Embeddingové modely představují techniky, které převádějí slova a texty do číselných vektorů v nízkodimenzionálním prostoru, přičemž tyto vektory zachycují sémantické a syntaktické vlastnosti původního textu. Jejich vznik byl motivován potřebou číselně reprezentovat textová data, aby bylo možné je efektivně zpracovávat pomocí algoritmů strojového učení. Praktické využití embeddingových modelů najdeme například při tvorbě vstupů do klasifikačních a prediktivních modelů, v automatickém překladu, informačním vyhledávání a dalších aplikacích, kde je klíčové porozumět významu a struktuře textu. V této dizertaci se konkrétně věnuji třem úlohám, které rozšiřují možnosti využití embeddingů. První část se zaměřuje na predikci potenciální citovanosti vědeckých článků; pomocí embeddingů získaných z abstraktů vytvářím modely, které odhadují, zda článek může v budoucnu dosáhnout vysoké citovanosti. Druhá část práce navrhuje novou metodu pro vysvětlitelnost embeddingových modelů, která umožňuje přesně určit, jak jednotlivá slova přispívají k výsledkům predikce, a tím poskytuje srozumitelnou interpretaci modelu. Třetí část se zabývá vysvětlením sémantické podobnosti, kdy ukazuji, jak distribuční statistika ve spojení s lexikálními znalostmi umožňuje lepší pochopení vztahů mezi slovy v embeddingovém prostoru. Výsledky této práce mají potenciál zlepšit přesnost a efektivitu klasifikačních a prediktivních modelů ve vědecké analytice a zároveň přispět k hlubšímu porozumění mechanismům, kterými textové modely zpracovávají přirozený jazyk.
Keywords:
Explaining; Text data; Word embeddings; Explaining; Text data; Word embeddings
Institution: University of Economics, Prague
(web)
Document availability information: Available in the digital repository of the University of Economics, Prague. Original record: https://vskp.vse.cz/eid/98669