Challenge: Existing methods for multi-change captioning are difficult because it requires a higher level of cognition to reason an arbitrary number of changes.
Approach: They propose a context-aware difference distilling network to capture all genuine changes for yielding sentences.
Outcome: The proposed network captures all genuine changes for yielding sentences on three public datasets.

Similar Papers

Change Entity-guided Heterogeneous Representation Disentangling for Change Captioning (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches to describe differences between two images are highly challenging due to distractors such as illumination and viewpoint changes.
Approach: They propose a change-entity-guided disentanglement network that explicitly learns difference representations while mitigating the impact of distractors.
Outcome: The proposed method outperforms existing methods on CLEVR-Change, CLE VR-DC and Spot-the-Diff datasets and achieves state-of-the art performance.
Semantic Relation-aware Difference Representation Learning for Change Captioning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to describe semantic change in images with distractors are difficult to learn .
Approach: They propose a semantic relation-aware difference representation learning network to explicitly learn the difference representation in the existence of distractors.
Outcome: The proposed network achieves state-of-the-art performance on CLEVR-Change and Spot-the -Diff datasets.
Sequence Shortening for Context-Aware Machine Translation (2024.findings-eacl)

Copied to clipboard

Challenge: Context-aware Machine Translation aims to improve translations of sentences by incorporating surrounding sentences as context.
Approach: They propose to use latent representation of source sentence as context in a multi-encoder architecture to achieve higher accuracy on contrastive datasets.
Outcome: The proposed architectures achieve comparable BLEU and COMET scores on contrastive datasets and comparable accuracies on the single- and multi-encoder approaches.
Distillation of encoder-decoder transformers for sequence labelling (2023.findings-eacl)

Copied to clipboard

Challenge: despite the strong trend in NLP to explore the use of large language models, there is still limited work evaluating prompting and decoding mechanisms for SL tasks.
Approach: They propose a hallucination-free framework for sequence tagging that is especially suited for distillation.
Outcome: The proposed framework performs well across multiple sequence labelling datasets and in a few-shot learning scenario.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Distilling Structured Knowledge for Text-Based Relational Reasoning (2020.emnlp-main)

Copied to clipboard

Challenge: Existing text-based relational reasoning models lack a symbolic representation of text . performance gap between NLP models and structured models remains .
Approach: They first pre-train a GNN on a reasoning task using structured inputs and then incorporate its knowledge into an NLP model.
Outcome: The proposed model improves on two state-of-the-art NLP models on 13 different inductive reasoning datasets from the CLUTRR benchmark.
Distilling Linguistic Context for Language Model Compression (2021.emnlp-main)

Copied to clipboard

Challenge: Knowledge distillation is a major technique for deploying vast language models in resource-strapped environments.
Approach: They propose a method that transfers contextual knowledge via Word Relation and Layer Transforming Relation.
Outcome: The proposed method is able to transfer contextual knowledge without restrictions on architectural changes between teacher and student on language understanding tasks.
Decoding by Contrasting Knowledge: Enhancing Large Language Model Confidence on Edited Facts (2025.acl-long)

Copied to clipboard

Challenge: In-context knowledge editing (ICE) is currently the most effective method for knowledge editing, but it is constrained by the black-box modeling of LLMs and lacks interpretability.
Approach: They propose a method to decode new knowledge by comparing logits with unedited knowledge to improve the accuracy of LLMs.
Outcome: The proposed method improves the performance of LLaMA3-8B-instruct on MQuAKE by up to 219%.
CLIP4IDC: CLIP for Image Difference Captioning (2022.aacl-short)

Copied to clipboard

Challenge: Conventional approaches learn an IDC model with a pre-trained and usually frozen visual feature extractor.
Approach: They propose to transfer a CLIP model to the downstream IDC task to address two major issues: (1) a large domain gap exists between the pre-training datasets used for training such a visual feature extractor; (2) the visual feature extraction often does not effectively encode the visual changes between two images.
Outcome: Experiments on three IDC benchmark datasets show the proposed model performs well.
Learning to Describe Implicit Changes: Noise-robust Pre-training for Image Difference Captioning (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) are used to capture subtle differences between images but are noisy and coarse summaries.
Approach: They propose a noise-robust approach to image difference capture using large multimodal models . they use LMMs with structured prompts to generate fine-grained change descriptions .
Outcome: The proposed model outperforms streamlined architectures and improves inference efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations