Papers by Shufan Li
InstructAny2Pix: Image Editing with Multi-Modal Prompts (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing image editing methods struggle with complex instructions involving multiple objects or reference images. |
| Approach: | They propose a novel image editing model that leverages a multi-modal LLM to execute complex edit instructions. |
| Outcome: | The proposed model outperforms existing models and benchmarks in two multi-modal datasets. |