Papers by João Bordalo
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks (2024.acl-long)
Copied to clipboard
João Bordalo, Vasco Ramos, Rodrigo Valério, Diogo Glória-Silva, Yonatan Bitton, Michal Yarom, Idan Szpektor, Joao Magalhaes
| Challenge: | Large Vision/Language Models (LVLMs) are less capable of generating accompanying image sequences. |
| Approach: | They propose a method that integrates a Latent Diffusion Model (LDM) with an LLM to generate captions to maintain semantic coherence of the sequence. |
| Outcome: | The proposed method is preferred by humans in 46.6% of the cases against 26.6% for the second best method. |