Papers by Aida Nematzadeh

12 papers
Evaluating Theory of Mind in Question Answering (D18-1)

Copied to clipboard

Challenge: a dataset is proposed for question answering models with respect to their capacity to reason about beliefs.
Approach: They propose a dataset for evaluating question answering models with respect to their capacity to reason about beliefs.
Outcome: The proposed dataset is inspired by theory-of-mind experiments that examine whether children are able to reason about beliefs of others.
Weakly-Supervised Learning of Visual Relations in Multimodal Pretraining (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work in vision-and-language pretraining has investigated supervised signals from object detection data to learn better, fine-grained multimodal representations.
Approach: They propose two approaches to contextualise visual entities in a multimodal setup by using verbalised scene graphs and masked relation prediction.
Outcome: The proposed models can learn better representations from weakly-supervised relations data.
Language Learning and Processing in People and Machines (N19-5)

Copied to clipboard

Challenge: This tutorial introduces different stages of language acquisition and their parallel problems in NLP.
Approach: This tutorial introduces different stages of language acquisition and their parallel problems in NLP.
Outcome: This tutorial introduces different stages of language acquisition and their parallel problems in NLP.
Vision-Language Pretraining: Current Trends and the Future (2022.acl-tutorials)

Copied to clipboard

Challenge: Recent vision-language models are being used for downstream tasks that require large datasets and supervised datasets.
Approach: They focus on recent vision-language pretraining paradigms and their strengths and shortcomings . they compare the different family of models used for vision- language pretraining .
Outcome: This paper provides the background on image–language datasets, benchmarks, and modeling innovations before the multimodal pretraining area.
A Systematic Investigation of Commonsense Knowledge in Large Language Models (2022.emnlp-main)

Copied to clipboard

Challenge: Recent large language models (LMs) have shown impressive performance on many NLP tasks under the zero-shot and few-shot setup.
Approach: They conduct a systematic and rigorous zero-shot and few-shot commonsense evaluation of large pre-trained language models to better understand their ability to capture commonsensical knowledge.
Outcome: The proposed model can exploit surface cues and annotation artefacts without task-specific supervision and is insufficient to achieve human-level commonsense performance.
MAPL: Parameter-Efficient Adaptation of Unimodal Pre-Trained Models for Vision-Language Few-Shot Prompting (2023.eacl-main)

Copied to clipboard

Challenge: Large pre-trained models have proved to be remarkable zero- and (prompt-based) few-shot learners in unimodal vision and language tasks.
Approach: They propose to use frozen unimodal models to learn a lightweight mapping between the representation spaces of unimod models using aligned image-text data.
Outcome: The proposed method can generalize to unseen VL tasks from a few in-context examples while training orders of magnitude fewer parameters.
Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers (2021.tacl-1)

Copied to clipboard

Challenge: Recent studies suggest multimodal transformer models learn rich visual-linguistic representations.
Approach: They focus on dataset noise and language similarity to their downstream task . they find that models with a multimodal attention mechanism outperform deeper models with modality-specific attention mechanisms.
Outcome: The proposed models outperform models with a multimodal attention mechanism on downstream tasks.
Pragmatics in Language Grounding: Phenomena, Tasks, and Modeling Approaches (2023.findings-emnlp)

Copied to clipboard

Challenge: People rely heavily on context to enrich meaning beyond what is literally said.
Approach: They analyze how task goals, environmental contexts, and communicative affordances in each work enrich linguistic meaning.
Outcome: The proposed frameworks are based on linguistic goals, environmental contexts, and communicative affordances to enrich linguistic meaning.
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization (2023.findings-eacl)

Copied to clipboard

Challenge: Visual question answering (VQA) is a task of answering open-ended questions about images.
Approach: They evaluate two vision-and-language (V&L) models under different settings . they find they tend to learn to solve the benchmark rather than the skills required by VQA .
Outcome: The proposed models exhibit poor generalization under out-of-distribution settings.
Measuring Progress in Fine-grained Vision-and-Language Understanding (2023.acl-long)

Copied to clipboard

Challenge: X-VLM models lack "fine-grained" understanding of relationships, verbs and numbers in images . pretraining on large-scale image–text data from the Web has facilitated rapid progress on many vision-and-language tasks .
Approach: They investigate models that outperform other baselines on fine-grained data . they highlight importance of novel losses and rich data sources for learning fine-grain skills .
Outcome: The proposed model outperforms baseline models on four fine-grained benchmarks . the model outpersforms other baseline models and even degrades performance .
Learning to Segment Actions from Observation and Narration (2020.acl-main)

Copied to clipboard

Challenge: a generative segmental model of task structure is applied to video training . despite its simplicity, the model performs well in unsupervised and weakly-supervised settings .
Approach: They propose a generative segmental model of task structure guided by narration to video segmentation .
Outcome: The proposed model performs well in unsupervised and weakly-supervised training . it allows us to vary the sources of supervision used in training despite its simplicity .
Probing Image-Language Transformers for Verb Understanding (2021.findings-acl)

Copied to clipboard

Challenge: Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning.
Approach: They collect a dataset of image-sentence pairs consisting of 421 verbs . they evaluate pretrained image-language transformers and find they fail more in situations that require verb understanding compared to other parts of speech.
Outcome: The proposed model trains on a manually-annotated and smaller dataset does better on the task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations