National Repository of Grey Literature 1 records found  Search took 0.00 seconds. 
Representation of Text and Its Influence on Categorization
Šabatka, Ondřej ; Chmelař, Petr (referee) ; Bartík, Vladimír (advisor)
The thesis deals with machine processing of textual data. In the theoretical part, issues related to natural language processing are described and different ways of pre-processing and representation of text are also introduced. The thesis also focuses on the usage of N-grams as features for document representation and describes some algorithms used for their extraction. The next part includes an outline of classification methods used. In the practical part, an application for pre-processing and creation of different textual data representations is suggested and implemented. Within the experiments made, the influence of these representations on accuracy of classification algorithms is analysed.

Interested in being notified about new results for this query?
Subscribe to the RSS feed.