Challenge: Existing text-guided image editing methods rely on end-to-end pixel-level inpainting paradigm . existing models lack such intermediate representations and Reasoning-then-action process .
Approach: They propose a "Decompose-then-Action" paradigm that revisits image editing as an actionable interaction process within a structured environment.
Outcome: The proposed paradigm outperforms existing methods in compositional editing tasks.

Similar Papers

AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models (2025.emnlp-main)

Copied to clipboard

Challenge: Text-to-Image (T2I) models have been successful in generating images from textual descriptions, but they struggle to capture nuanced and implicit attributes inherent in action depiction.
Approach: They propose a benchmark to evaluate the performance of T2I models in generating images from action-centric prompts.
Outcome: The proposed model achieves an increase of 72% on AcT2I.
Learning to Follow Object-Centric Image Editing Instructions Faithfully (2023.findings-emnlp)

Copied to clipboard

Challenge: avrahami et al., 2022b,a): natural language instructions are often underspecified, requiring models to uncover their implicit meaning.
Approach: They propose to use paired data to model the implicit meaning of instructions . they also propose to ground the model to localize where the edit has to be performed .
Outcome: The proposed model performs better than state-of-the-art baselines on paired data, showing improvements in quality and faithfulness.
InstructAny2Pix: Image Editing with Multi-Modal Prompts (2025.findings-naacl)

Copied to clipboard

Challenge: Existing image editing methods struggle with complex instructions involving multiple objects or reference images.
Approach: They propose a novel image editing model that leverages a multi-modal LLM to execute complex edit instructions.
Outcome: The proposed model outperforms existing models and benchmarks in two multi-modal datasets.
Pioneering Reliable Assessment in Text-to-Image Knowledge Editing: Leveraging a Fine-Grained Dataset and an Innovative Criterion (2024.findings-emnlp)

Copied to clipboard

Challenge: Text-to-image models encode factual knowledge into their parameters, but they may become obsolete over time.
Approach: They propose a framework for T2I knowledge editing that integrates paraphrase and multi-object test to enable more fine-grained assessment on knowledge generalization.
Outcome: The proposed framework improves on existing models and improves their performance.
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Reasoning is a fundamental capability underpinning text-to-image (T2I) generation.
Approach: They propose a benchmark to rigorously assess reasoning-driven T2I generation.
Outcome: Experiments with 16 representative T2I models show limited reasoning performance . a strong pipeline-based framework decouples reasoning and generation .
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits (2025.acl-long)

Copied to clipboard

Challenge: Text-guided image editing is becoming increasingly widespread . current models struggle to evaluate edits comprehensively and often hallucinate when describing changes.
Approach: They propose a novel framework to evaluate edits based on human annotations . they use a template to collect human annotation data and validate the results .
Outcome: The proposed methods outperform current models in artifact detection and difference caption generation.
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent work has shown promise by incorporating pixel-level visual information into the reasoning process, enabling VLMs to access high-resolution visual details during their thought process.
Approach: They propose a framework that dynamically determines necessary pixel-level operations based on the input query.
Outcome: The proposed model achieves 73.4% accuracy on HR-Bench 4K while maintaining a tool usage ratio of only 20.1%, improving accuracy and reducing tool usage by 66.5% compared to the previous methods.
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training (2025.coling-main)

Copied to clipboard

Challenge: State-of-the-art approaches to this task resort to supervised training and labelling, masking or training.
Approach: They propose an approach that does without any task-specific supervision and offers thus a better potential for improvement.
Outcome: The proposed approach achieves very competitive performance and scales up in a way that requires no task-specific supervision.
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder (2024.naacl-long)

Copied to clipboard

Challenge: Text-to-image generative models encode factual associations that can quickly become outdated, diminishing their utility for end-users.
Approach: They propose a method for editing factual associations in text-to-image models without retraining or explicit input from end-users.
Outcome: The proposed method improves generalization and preservation of unrelated concepts on an existing dataset and compares with other methods.
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation (2026.findings-acl)

Copied to clipboard

Challenge: Text-to-image (T2I) generative models have demonstrated exceptional capability in synthesizing high-quality images from textual prompts.
Approach: They propose a benchmark to explore the knowledge-driven reasoning capabilities of T2I models.
Outcome: The proposed benchmark examines the knowledge-driven reasoning capabilities of T2I models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations