Papers by Denis Paperno
Geo-Aware Image Caption Generation (2020.coling-main)
Copied to clipboard
| Challenge: | Standard image caption generation systems do not take contextual information or world knowledge into account. |
| Approach: | They propose to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions. |
| Outcome: | The proposed system achieves significant improvements on a dataset that contains contextualized captions and geographic metadata and improves BLEU, ROUGE, METEOR and CIDEr scores. |
„Mann“ is to “Donna” as「国王」is to « Reine » Adapting the Analogy Task for Multilingual and Contextual Embeddings (2023.starsem-1)
Copied to clipboard
| Challenge: | a lack of comparable multilingual benchmarks and a consensual evaluation protocol for contextual models remains an open question. |
| Approach: | They propose a multilingual analogy dataset and evaluate human and contextual embedding performance. |
| Outcome: | The proposed dataset evaluates human and contextual embedding models on the analogy task. |
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks focus on single image settings, but some focus on multi-image settings. |
| Approach: | They introduce the TempVS benchmark which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models in image sequences. |
| Outcome: | The proposed model performs poorly compared to human models in vision and language tasks. |
Grounded and well-rounded: a methodological approach to the study of cross-modal and cross-lingual grounding (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on grounding have focused on qualitatively different generalizations, but limited empirical evidence supports either position. |
| Approach: | They propose a methodological framework for studying the effects of grounding on NLP systems . they use a sample of models trained on different input modalities to tease out qualitative differences . |
| Outcome: | The proposed framework teases out qualitative differences in model behavior between models trained on different input sources from quantifiable models. |
What Meaning-Form Correlation Has to Compose With: A Study of MFC on Artificial and Natural Language (2020.coling-main)
Copied to clipboard
| Challenge: | Compositionality is a widely discussed property of natural languages, although its exact definition has been elusive. |
| Approach: | They propose that compositionality can be measured by measuring meaning-form correlation . they analyze three sets of languages: artificial toy languages tailored to be compositional . |
| Outcome: | The proposed method can assess compositionality on three sets of languages . linguistic phenomena such as synonymy and ungrounded stop-words weigh on the results . |
How to Dissect a Muppet: The Structure of Transformer Embedding Spaces (2022.tacl-1)
Copied to clipboard
| Challenge: | Pretrained embeddings based on the Transformer architecture have taken the NLP community by storm . a novel decomposition of Transformer output embeddables is demonstrated . |
| Approach: | They propose to decompose Transformer output embeddings into a sum of vector factors . they show multi-head attentions and feed-forwards are not equally useful in downstream applications . |
| Outcome: | The proposed method outperforms recurrent architectures on a wide variety of tasks. |
Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations (2023.starsem-1)
Copied to clipboard
| Challenge: | a recent study of the effect of visual grounding on language representations has given a new life to the debate around extractability and quality of semantic information in representations trained solely on textual input. |
| Approach: | They compare word embeddings from vision-and-language models to text-only models . they identify meaning properties and relations that characterize words whose embeddements are most affected by visual grounding . |
| Outcome: | The proposed model differs from text-only models on semantic representations of language . the study is the first large-scale study of the effect of visual grounding on language representations . |