Original title: Vyvozování v přirozeném jazyce s využitím obrazových dat
Translated title: Grounding Natural Language Inference on Images
Authors: Vu Trong, Hoa ; Pecina, Pavel (advisor) ; Libovický, Jindřich (referee)
Document type: Master’s theses
Year: 2018
Language: eng
Abstract: Grounding Natural Language Inference on Images Hoa Trong VU July 20, 2018 Abstract Despite the surge of research interest in problems involving linguistic and vi- sual information, exploring multimodal data for Natural Language Inference remains unexplored. Natural Language Inference, regarded as the basic step towards Natural Language Understanding, is extremely challenging due to the natural complexity of human languages. However, we believe this issue can be alleviated by using multimodal data. Given an image and its description, our proposed task is to determined whether a natural language hypothesis contra- dicts, entails or is neutral with regards to the image and its description. To address this problem, we develop a multimodal framework based on the Bilat- eral Multi-perspective Matching framework. Data is collected by mapping the SNLI dataset with the image dataset Flickr30k. The result dataset, made pub- licly available, has more than 565k instances. Experiments on this dataset show that the multimodal model outperforms the state-of-the-art textual model. References 1
Keywords: Grounding Natural Language Inference on Images; vyvozování v přirozeném jazyce

Institution: Charles University Faculties (theses) (web)
Document availability information: Available in the Charles University Digital Repository.
Original record: http://hdl.handle.net/20.500.11956/101573

Permalink: http://www.nusl.cz/ntk/nusl-387831


The record appears in these collections:
Universities and colleges > Public universities > Charles University > Charles University Faculties (theses)
Academic theses (ETDs) > Master’s theses
 Record created 2018-11-15, last modified 2022-03-04


No fulltext
  • Export as DC, NUŠL, RIS
  • Share