Papers by Dustin Schwenk

2 papers
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work has adapted vision-and-language models to generative tasks like image captioning.
Approach: They propose an extension to LXMERT with training refinements to generate images from text.
Outcome: The proposed model can generate images from pieces of text while still being comparable to existing models.
Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text (2021.emnlp-main)

Copied to clipboard

Challenge: Communicating with humans is challenging for AIs because of its complexity and multimodality.
Approach: They propose to use a game of drawing and guessing based on Pictionary to test AIs' understanding of the world and multi-modal gestures.
Outcome: The proposed game is a test for mixing language and visual/symbolic communication in AI.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations