ViPE: Visualise Pretty-much Everything (2023.emnlp-main)

Copied to clipboard

Challenge: Figure and non-literal expressions are deeply integrated in human communication . text-to-image models like Stable Diffusion struggle to depict non-figural expression .
Approach: They propose a series of lightweight and robust language models that can be used to visualise non-literal expressions.
Outcome: The proposed language models are more robust than existing models and can generate high-quality images.

Similar Papers

CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations (2026.findings-acl)

Copied to clipboard

Challenge: Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment.
Approach: They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation.
Outcome: The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations.
Computational Narrative Understanding for Expressive Text-to-Speech (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora.
Approach: They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch .
Outcome: The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model.
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons (2026.acl-long)

Copied to clipboard

Challenge: Autoregressive (AR) models rely on bidirectional attention, creating a structural mismatch with pre-trained Autoregression models.
Approach: They propose a framework that efficiently adapts autoregressive (AR) models to the diffusion paradigm.
Outcome: The proposed framework reduces training costs by orders of magnitude while maintaining state-of-the-art performance.
FLUTE: Figurative Language Understanding through Textual Explanations (2022.emnlp-main)

Copied to clipboard

Challenge: Figurative language understanding is a recognizing textual entailment task, but lacks data for figurative language.
Approach: They propose to use a dataset to analyze figurative NLI instances with explanations to improve models' performance.
Outcome: The proposed dataset can scale up models even for figurative language using human annotations.
Beyond Understanding: Evaluating the Pragmatic Gap in LLMs’ Cultural Processing of Figurative Language (2026.eacl-long)

Copied to clipboard

Challenge: Using figurative language as a proxy for cultural nuance and local knowledge, large language models struggle with connotative meaning.
Approach: They evaluate large language models' ability to process culturally grounded language . they use figurative language as a proxy for cultural nuance and local knowledge .
Outcome: The proposed models can understand and use figurative expressions that encode local knowledge and social nuance.
Learning Visually-Grounded Semantics from Contrastive Adversarial Samples (C18-1)

Copied to clipboard

Challenge: Existing frameworks for grounding distributional representations of texts on the visual domain are limited . effective and efficient grounding of distributional embeddings remains challenging .
Approach: They propose to ground distributional representations of texts on the visual domain using visual-semantic embeddings.
Outcome: The proposed model improves on a diverse set of downstream tasks and defends known-type adversarial attacks.
On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization (2022.findings-emnlp)

Copied to clipboard

Challenge: Combining visual modality with pretrained language models has been effective for descriptive tasks such as image captioning.
Approach: They ask: do multimodal models combine visual and visual adapted language models? they find that CLIP image representations and scaling of language models do not consistently improve self-rationalization in multimodal tasks.
Outcome: The proposed model types do not consistently improve self-rationalization in multimodal tasks.
MEGA: Multilingual Evaluation of Generative AI (2023.emnlp-main)

Copied to clipboard

Challenge: Large Large Models (LLMs) have shown impressive performance on many natural language processing tasks such as language understanding, reasoning, and language generation.
Approach: They present a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.
Outcome: The proposed framework evaluates generative models on 16 NLP datasets across 70 typologically diverse languages and compares them to state-of-the-art non-autoregressive models.
VisText: A Benchmark for Semantically Rich Chart Captioning (2023.acl-long)

Copied to clipboard

Challenge: Current approaches for automatically generating chart captions struggle to articulate the perceptual or cognitive features that are the hallmark of charts (e.g., complex trends and patterns).
Approach: They propose a dataset of 12,441 pairs of charts and captions that describe charts’ construction, report key statistics, and identify perceptual and cognitive phenomena.
Outcome: The proposed model generates coherent, semantically rich captions and performs on par with state-of-the-art chart captioning models across machine translation and text generation metrics.
On the Consistency of Commonsense in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluations of commonsense for large language models focus on downstream knowledge tasks, failing to probe whether LLMs truly understand and utilize knowledge or merely memorize it.
Approach: They propose to automatically construct a large benchmark named CoCo which measures LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.
Outcome: The proposed benchmark systematically assesses LLMs’ knowledge memorization, comprehension, and application and examines the consistency between these tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations