Papers with CMA
Improving Cross-modal Alignment for Text-Guided Image Inpainting (2023.eacl-main)
Copied to clipboard
| Challenge: | Existing methods allocate most of computation to visual encoding, while light computation on modeling modality interactions. |
| Approach: | They propose a novel model for text-guided image inpainting by improving cross-modal alignment knowledge by using a vision-language encoder and an image generator. |
| Outcome: | The proposed model achieves state-of-the-art performance compared with other strong competitors on two vision-language datasets. |