Papers by Aida Nematzadeh
Evaluating Theory of Mind in Question Answering (D18-1)
Copied to clipboard
| Challenge: | a dataset is proposed for question answering models with respect to their capacity to reason about beliefs. |
| Approach: | They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs. |
| Outcome: | The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others. |
Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining (2023.emnlp-main)
Copied to clipboard
| Challenge: | Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations. |
| Approach: | They propose two approaches to contextualise visual entities in a multimodal setup by using verbalised scene graphs and masked relation prediction. |
| Outcome: | The proposed models can learn better representations from weakly-supervised relations data. |
Language Learning and Processing in People and Machines (N19-5)
Copied to clipboard
| Challenge: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
| Approach: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
| Outcome: | This tutorial introduces different stages of language acquisition and their parallel problems in NLP. |
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)
Copied to clipboard
| Challenge: | Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets. |
| Approach: | They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining . |
| Outcome: | This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area. |
A Systematic Investigation of Commonsense Knowledge in Large Language Models (2022.emnlp-main)
Copied to clipboard
Xiang Lorraine Li, Adhiguna Kuncoro, Jordan Hoffmann, Cyprien de Masson d’Autume, Phil Blunsom, Aida Nematzadeh
| Challenge: | Recent large language models (LMs) have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup. |
| Approach: | They conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pre-trained language models to better understand their ability to capture commonsensical knowledge. |
| Outcome: | The proposed model can exploit surface cues and annotation artefacts without task-specific supervision and is insufficient to achieve human-level commonsense performance. |
MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot Prompting (2023.eacl-main)
Copied to clipboard
| Challenge: | Large pre-trained models have proved to be remarkable zero- and (prompt-based) few-shot learners in unimodal vision and language tasks. |
| Approach: | They propose to use frozen unimodal models to learn a lightweight mapping between the representation spaces of unimod models using aligned image-text data. |
| Outcome: | The proposed method can generalize to unseen VL tasks from a few in-context examples while training orders of magnitude fewer parameters. |
Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers (2021.tacl-1)
Copied to clipboard
| Challenge: | Recent studies suggest multimodal transformer models learn rich visual-linguistic representations. |
| Approach: | They focus on dataset noise and language similarity to their downstream task . they find that models with a multimodal attention mechanism outperform deeper models with modality-specific attention mechanisms. |
| Outcome: | The proposed models outperform models with a multimodal attention mechanism on downstream tasks. |
Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling Approaches (2023.findings-emnlp)
Copied to clipboard
| Challenge: | People rely heavily on context to enrich meaning beyond what is literally said. |
| Approach: | They analyze how task goals, environmental contexts, and communicative affordances in each work enrich linguistic meaning. |
| Outcome: | The proposed frameworks are based on linguistic goals, environmental contexts, and communicative affordances to enrich linguistic meaning. |
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization (2023.findings-eacl)
Copied to clipboard
Aishwarya Agrawal, Ivana Kajic, Emanuele Bugliarello, Elnaz Davoodi, Anita Gergely, Phil Blunsom, Aida Nematzadeh
| Challenge: | Visual question answering (VQA) is a task of answering open-ended questions about images. |
| Approach: | They evaluate two vision-and-language (V&L) models under different settings . they find they tend to learn to solve the benchmark rather than the skills required by VQA . |
| Outcome: | The proposed models exhibit poor generalization under out-of-distribution settings. |
Measuring Progress in Fine-grained Vision-and-Language Understanding (2023.acl-long)
Copied to clipboard
| Challenge: | X-VLM models lack "fine-grained" understanding of relationships, verbs and numbers in images . pretraining on large-scale image–text data from the Web has facilitated rapid progress on many vision-and-language tasks . |
| Approach: | They investigate models that outperform other baselines on fine-grained data . they highlight importance of novel losses and rich data sources for learning fine-grain skills . |
| Outcome: | The proposed model outperforms baseline models on four fine-grained benchmarks . the model outpersforms other baseline models and even degrades performance . |
Learning to Segment Actions from Observation and Narration (2020.acl-main)
Copied to clipboard
| Challenge: | a generative segmental model of task structure is applied to video training . despite its simplicity, the model performs well in unsupervised and weakly-supervised settings . |
| Approach: | They propose a generative segmental model of task structure guided by narration to video segmentation . |
| Outcome: | The proposed model performs well in unsupervised and weakly-supervised training . it allows us to vary the sources of supervision used in training despite its simplicity . |
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)
Copied to clipboard
| Challenge: | Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning. |
| Approach: | They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech. |
| Outcome: | The proposed model trains on a manually-annotated and smaller dataset does better on the task. |