Papers with CMA

1 papers
Improving Cross-modal Alignment for Text-Guided Image Inpainting (2023.eacl-main)

Copied to clipboard

Challenge: Existing methods allocate most of computation to visual encoding, while light computation on modeling modality interactions.
Approach: They propose a novel model for text-guided image inpainting by improving cross-modal alignment knowledge by using a vision-language encoder and an image generator.
Outcome: The proposed model achieves state-of-the-art performance compared with other strong competitors on two vision-language datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations