Papers by Özge Alaçam
WikiScenes with Descriptions: Aligning Paragraphs and Sentences with Images in Wikipedia Articles (2024.starsem-1)
Copied to clipboard
| Challenge: | Existing work on processing image-text alignment in multimodal documents has been unsupervised, facing the challenge of missing evaluation and training data. |
| Approach: | They propose to provide one of the first datasets that provides ground-truth annotations of image-text alignments in multi-paragraph multi-image articles. |
| Outcome: | The proposed dataset can be used to study phenomena of visual language grounding in longer documents and assess retrieval capabilities of language models trained on captioning data. |
Incorporating Contextual Information for Language-Independent, Dynamic Disambiguation Tasks (L18-1)
Copied to clipboard
| Challenge: | a proposed multimodal system can resolve syntactic ambiguities by exploiting external evidence, says a researcher . a parser that processes linguistic information is expected to handle syntakically unambiguous sentences, but it cannot. |
| Approach: | They propose to exploit external contextual information to resolve ambiguous sentences . they propose to use data-driven and grammar-based approaches to solve ambiguities . |
| Outcome: | The proposed system confirms this hypothesis in experiments on syntactically ambiguous sentences. |
Towards Multi-Modal Text-Image Retrieval to improve Human Reading (2021.naacl-srw)
Copied to clipboard
| Challenge: | In primary school, children's books, as well as in modern language learning apps, multi-modal learning strategies like illustrations of terms and phrases are used to support reading comprehension. |
| Approach: | They propose to use multi-modal transformers to train multi-dimensional models on text-image retrieval to support a user's reading comprehension of arbitrary text. |
| Outcome: | The proposed model performs poorly because of the short and relatively simple textual data that the current models are trained with. |