Papers by Arkhipkin Vladimir

2 papers
Kandinsky 3: Text-to-Image Synthesis for Multifunctional Generative Framework (2024.emnlp-demo)

Copied to clipboard

Challenge: Text-to-image (T2I) diffusion models are popular for image manipulation, but also for video generation.
Approach: They propose a novel T2I diffusion model based on latent diffusion that extends the base model for various applications.
Outcome: The proposed model achieves high quality and photorealism and is 3 times faster than the base model.
Kandinsky: An Improved Text-to-Image Synthesis with Image Prior and Latent Diffusion (2023.emnlp-demo)

Copied to clipboard

Challenge: Experimental evaluations demonstrate FID score of 8.03 on the COCO-30K dataset, marking our model as the top open source performer in terms of measurable image generation quality.
Approach: They propose a latent diffusion-based model that combines image prior and latent diffusive techniques to create a text-to-image architecture.
Outcome: The proposed model achieves the highest FID score among open-source models . it is compared with the state-of-the-art models on the COCO-30K dataset .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations