Challenge: Large language models possess remarkable capacity for processing language, but it remains unclear whether they can further generate creative content.
Approach: They utilize the divergent association task (DAT) to examine the creative thinking of large language models through a cognitive perspective.
Outcome: The proposed model outperforms the greedy search strategy while outperforming the average human level.

Similar Papers

Beyond Divergent Creativity: A Human-Based Evaluation of Creativity in Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models are increasingly used in verbal creative tasks.
Approach: They propose a divergent association task that focuses on novelty, ignoring appropriateness, a core component of creativity.
Outcome: The proposed model scores are lower than baselines with no creative abilities, undermining its validity for model evaluation.
Automated Creativity Evaluation of Language Models Across Open-Ended Tasks (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for evaluating creativity are tightly coupled to specific tasks and limiting scalability and generality.
Approach: They propose a domain-agnostic framework for quantifying LLM creativity across open-ended tasks.
Outcome: The proposed framework captures key facets of creativity including novelty, diversity, and task fulfilment with over 60% improved efficiency.
Benchmarking Language Model Creativity: A Case Study on Code Generation (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies on LLM creativity evaluation focus on open-ended generation tasks . however, the degree to which LLMs possess and utilize creativity for problem-solving remains unclear .
Approach: They propose a framework for quantifying LLM creativity that incorporates design ingredients . they introduce DENIAL PROMPTING which pushes LLMs to develop more creative solutions .
Outcome: The proposed framework quantifies creativity in LLMs on Codeforces problems . it also finds that even the most creative model fails to demonstrate human-like creativity .
Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used for creative tasks such as literary translation.
Approach: They propose a paired-task framework that assesses translational creativity using Units of Creative Potential (UCPs) they benchmark 23 models and four creativity-oriented prompts to assess translational comprehension .
Outcome: The proposed framework compares 23 models and four creativity-oriented prompts on literary excerpts from 11 books.
Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies evaluate the creative capabilities of large language models (LLMs) through diverse tasks, aiming to understand their strengths and limitations.
Approach: They propose to ask LLMs to generate Parallel Chains of Associations to Evaluate their creativity.
Outcome: The proposed framework minimizes the risk of data contamination and offers a highly efficient evaluation.
Creative Problem Solving in Large Language and Vision Models - What Would it Take? (2024.findings-emnlp)

Copied to clipboard

Challenge: CC is a multi-disciplinary field that seeks to develop computational methods capable of generating creative outcomes reminiscent of creative processes in humans.
Approach: They advocate for a strong integration of Computational Creativity with research in large language and vision models to address creative problem solving.
Outcome: The proposed model can address creative problem solving, the authors argue . they show that the model can be integrated with LLVMs to address creative problems .
Evaluating the Deductive Competence of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing large language models have limited abilities to solve deductive reasoning problems . performance differences between conditions do not improve overall performance .
Approach: They investigate whether several large language models can solve a deductive reasoning problem in their conventional form.
Outcome: The proposed models can solve a classic type of deductive reasoning problem in their conventional form.
Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs (2025.coling-main)

Copied to clipboard

Challenge: a fine-tuned small language model (SLM) can generate human-like text, but it requires immense computational resources and large datasets.
Approach: They evaluate the creative writing abilities of a fine-tuned small language model, BART-large . they compare it to human writers and two large language models: GPT-3.5 and GPT-4o .
Outcome: The proposed model outperforms human writers and two large language models in two experiments . the results highlight how model size and fine-tuning influence creativity, fluency, and coherence .
Large Language Models for Generative Recommendation: A Survey and Visionary Discussions (2024.lrec-main)

Copied to clipboard

Challenge: Large language models (LLMs) have revolutionized the field of natural language processing but are not fully able to leverage the generative power of LLM.
Approach: They examine the progress, methods, and future directions of large language models . they examine what generative recommendation is, why RS should advance to generative recommendations .
Outcome: The proposed approach can be simplified to generate recommendations from the entire pool of items.
The Pragmatic Mind of Machines: Tracing the Emergence of Pragmatic Competence in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Current large language models (LLMs) have demonstrated emerging capabilities in social intelligence tasks, including implicature resolution and theory-of-mind reasoning.
Approach: They introduce a dataset grounded in the pragmatic concept of alternatives to evaluate whether large language models can accurately infer nuanced speaker intentions.
Outcome: The proposed model can infer nuanced speaker intentions by inferring the speaker’s intended meaning and explaining when and why a speaker would choose one utterance over its alternative.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations