Challenge: Large language models (LLMs) are increasingly used for creative tasks such as literary translation.
Approach: They propose a paired-task framework that assesses translational creativity using Units of Creative Potential (UCPs) they benchmark 23 models and four creativity-oriented prompts to assess translational comprehension .
Outcome: The proposed framework compares 23 models and four creativity-oriented prompts on literary excerpts from 11 books.

Similar Papers

Automated Creativity Evaluation of Language Models Across Open-Ended Tasks (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating creativity are tightly coupled to specific tasks and limiting scalability and generality.
Approach: They propose a domain-agnostic framework for quantifying LLM creativity across open-ended tasks.
Outcome: The proposed framework captures key facets of creativity including novelty, diversity, and task fulfilment with over 60% improved efficiency.
Towards A “Novel” Benchmark: Evaluating Literary Fiction with Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) context windows have enabled them to process inputs over 100K tokens and generate outputs of up to 10K token.
Approach: They propose a multi-level evaluation framework that incorporates ten metrics across the Macro, Meso, and Micro levels and an annotated fiction dataset.
Outcome: The proposed framework incorporates ten metrics across the Macro, Meso, and Micro levels and is based on a human-human-AI dataset.
Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models are increasingly used in verbal creative tasks.
Approach: They propose a divergent association task that focuses on novelty, ignoring appropriateness, a core component of creativity.
Outcome: The proposed model scores are lower than baselines with no creative abilities, undermining its validity for model evaluation.
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)’s influential work on the New Yorker Cartoon Caption Contest.
Approach: They propose to decompose humor understanding into three components and improve each by enhancing visual understanding through improved annotation and utilizing LLM-generated humor reasoning and explanations.
Outcome: The proposed approach achieves 82.4% accuracy in caption ranking, significantly better than the previous 67% benchmark and matches the performance of world-renowned human experts in this domain.
Systematic Task Exploration with LLMs: A Study in Citation Text Generation (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) provide unprecedented flexibility in defining and executing complex, creative natural language generation tasks.
Approach: They propose a framework that consists of input manipulation, reference data, and output measurement to explore citation text generation.
Outcome: The proposed framework explores citation text generation, a popular scholarly NLP task that lacks consensus on the task definition and evaluation metric and has not yet been tackled within the LLM paradigm.
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies evaluate the creative capabilities of large language models (LLMs) through diverse tasks, aiming to understand their strengths and limitations.
Approach: They propose to ask LLMs to generate Parallel Chains of Associations to Evaluate their creativity.
Outcome: The proposed framework minimizes the risk of data contamination and offers a highly efficient evaluation.
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)

Copied to clipboard

Challenge: Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs.
Approach: They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria.
Outcome: The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning.
Evaluating the Creativity of LLMs in Persian Literary Text Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior research has focused primarily on English, with limited exploration of non-English literary traditions and without standardized methods for assessing creativity.
Approach: They build a dataset of user-generated Persian literary spanning 20 diverse topics and assess model outputs along four creativity dimensions .
Outcome: The proposed models generate Persian literary text enriched with culturally relevant expressions.
Exploring Precision and Recall to assess the quality and diversity of LLMs (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models are limited to specific tasks, but they are now widely available for a wide range of tasks.
Approach: They propose a framework for large language models such as Llama-2 and Mistral that imports precision and recall metrics from image generation to text generation.
Outcome: The proposed framework allows for a nuanced assessment of the quality and diversity of generated text without the need for aligned corpora.
Creative Planning with Language Models: Practice, Evaluation and Applications (2025.naacl-tutorial)

Copied to clipboard

Challenge: This tutorial explores how planning has been learned and deployed in creative workflows . many human creative tasks involve extensive planning, and actions need to be taken .
Approach: This tutorial explores how planning has been learned and deployed in creative workflows . authors discuss forward and backward learning approaches for planning in LLMs - and evaluation metrics tailored to latent plans .
Outcome: This tutorial examines how planning has been learned and deployed in creative workflows . it discusses forward and backward learning approaches for planning in LLMs - evaluation metrics tailored to latent plans .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations