Papers by Denis Paperno

7 papers
Geo-Aware Image Caption Generation (2020.coling-main)

Copied to clipboard

Challenge: Standard image caption generation systems do not take contextual information or world knowledge into account.
Approach: They propose to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions.
Outcome: The proposed system achieves significant improvements on a dataset that contains contextualized captions and geographic metadata and improves BLEU, ROUGE, METEOR and CIDEr scores.
„Mann“ is to “Donna” as「国王」is to « Reine » Adapting the Analogy Task for Multilingual and Contextual Embeddings (2023.starsem-1)

Copied to clipboard

Challenge: a lack of comparable multilingual benchmarks and a consensual evaluation protocol for contextual models remains an open question.
Approach: They propose a multilingual analogy dataset and evaluate human and contextual embedding performance.
Outcome: The proposed dataset evaluates human and contextual embedding models on the analogy task.
Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences? (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on single image settings, but some focus on multi-image settings.
Approach: They introduce the TempVS benchmark which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models in image sequences.
Outcome: The proposed model performs poorly compared to human models in vision and language tasks.
Grounded and well-rounded: a methodological approach to the study of cross-modal and cross-lingual grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on grounding have focused on qualitatively different generalizations, but limited empirical evidence supports either position.
Approach: They propose a methodological framework for studying the effects of grounding on NLP systems . they use a sample of models trained on different input modalities to tease out qualitative differences .
Outcome: The proposed framework teases out qualitative differences in model behavior between models trained on different input sources from quantifiable models.
What Meaning-Form Correlation Has to Compose With: A Study of MFC on Artificial and Natural Language (2020.coling-main)

Copied to clipboard

Challenge: Compositionality is a widely discussed property of natural languages, although its exact definition has been elusive.
Approach: They propose that compositionality can be measured by measuring meaning-form correlation . they analyze three sets of languages: artificial toy languages tailored to be compositional .
Outcome: The proposed method can assess compositionality on three sets of languages . linguistic phenomena such as synonymy and ungrounded stop-words weigh on the results .
How to Dissect a Muppet: The Structure of Transformer Embedding Spaces (2022.tacl-1)

Copied to clipboard

Challenge: Pretrained embeddings based on the Transformer architecture have taken the NLP community by storm . a novel decomposition of Transformer output embeddables is demonstrated .
Approach: They propose to decompose Transformer output embeddings into a sum of vector factors . they show multi-head attentions and feed-forwards are not equally useful in downstream applications .
Outcome: The proposed method outperforms recurrent architectures on a wide variety of tasks.
Leverage Points in Modality Shifts: Comparing Language-only and Multimodal Word Representations (2023.starsem-1)

Copied to clipboard

Challenge: a recent study of the effect of visual grounding on language representations has given a new life to the debate around extractability and quality of semantic information in representations trained solely on textual input.
Approach: They compare word embeddings from vision-and-language models to text-only models . they identify meaning properties and relations that characterize words whose embeddements are most affected by visual grounding .
Outcome: The proposed model differs from text-only models on semantic representations of language . the study is the first large-scale study of the effect of visual grounding on language representations .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations