Papers by Mayu Otani
Attending Self-Attention: A Case Study of Visually Grounded Supervision in Vision-and-Language Transformers (2021.acl-srw)
Copied to clipboard
| Challenge: | a growing body of research has been focused on what attention heads learn during the pre-training of visual grounded language models. |
| Approach: | They propose to use visual grounding to supervise attention directly to learn visual ground. |
| Outcome: | The proposed method improves the performance of a state-of-the-art visual grounded language model on vision-and-language tasks. |
iParaphrasing: Extracting Visually Grounded Paraphrases via an Image (C18-1)
Copied to clipboard
| Challenge: | iParaphrasing extracts visually grounded paraphrases, which are different phrasal expressions describing the same visual concept in an image. |
| Approach: | They propose a task to extract visually grounded paraphrases from images . they propose to model the similarity between the extracted VGPs using existing methods . |
| Outcome: | The proposed task extracts visually grounded paraphrases from images . the proposed method has the potential to improve multimodal language and image tasks . |