Papers by Takuya Narihira
Transformer-Exclusive Cross-Modal Representation for Vision and Language (2021.findings-acl)
Copied to clipboard
| Challenge: | a number of approaches to crossmodal representation have been used, but transformer architecture has taken over the recurrent neural networks in natural language processing tasks. |
| Approach: | They propose to use transformer architecture to handle cross-modal representations for vision and language with compatible performance to convolutional neural networks. |
| Outcome: | The proposed model outperforms recurrent neural networks in vision and language representations with transformer architecture. |