Papers by Yufei Tian

16 papers
Detecting Machine-Generated Long-Form Content with Latent-Space Variables (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing zero-shot methods to distinguish machine-generated long-form texts from humans are vulnerable to domain shift including different decoding strategies, variations in prompts, and attacks.
Approach: They propose a method that incorporates abstract elements as key deciding factors by training a latent-space model on sequences of events or topics derived from human-written texts.
Outcome: The proposed method improves on baselines on three domains and significantly improves over existing methods.
Are Large Language Models Capable of Generating Human-Level Narratives? (2024.emnlp-main)

Copied to clipboard

Challenge: a recent HCI study has pointed to gaps in machine storytelling ability at the global level . authors show that LLMs have less suspense and less tension than human stories .
Approach: They propose a computational framework to analyze narratives through three discourse-level aspects.
Outcome: The proposed framework analyzes narratives through three discourse-level aspects . it shows that LLMs fall short of human abilities in discourse understanding .
REFFLY: Melody-Constrained Lyrics Editing Model (2025.naacl-long)

Copied to clipboard

Challenge: Automatic melody-to-lyric (M2L) generation aims to create lyrics that align with a given melody.
Approach: They propose a framework for automatic melody-to-lyric generation that allows for a more flexible approach to creating lyrics from plain text.
Outcome: The proposed framework outperforms baselines Lyra and GPT-4 in musicality and text quality.
Evaluating Large Language Models on Controlled Generation Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have looked into the ability of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc. However, few studies investigate the controllability of large languages.
Approach: They propose to compare large language models with state-of-the-start finetuned smaller models to find that large language model controls are comparable to smaller models.
Outcome: The proposed model can meet hard constraints and perform better than state-of-the-art models.
Harnessing Black-Box Control to Boost Commonsense in LM’s Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Recent years have seen remarkable progress in massively Pre-Trained Language Models such as GPT-3 . however, their generated outputs lack commonsense at times .
Approach: They propose a framework that steers a frozen Pre-Trained Language Model towards more commonsense generation by training an auxiliary model.
Outcome: The proposed framework produces plausible outputs that incorporate concepts in a meaningful way.
Zero-shot Sonnet Generation with Discourse-level Planning and Aesthetics Features (2022.naacl-main)

Copied to clipboard

Challenge: a sonnet is a fourteen-line poem with rigorous meter-and-rhyme constraints.
Approach: They propose a framework which plans the poem sketch before decoding a sonnet without training on poems . they use a rhyme module, polishing module and a constrained decoding algorithm to impose the meter-and-rhyme constraint .
Outcome: The proposed framework generates sonnets that are coherent and poetic without training on poems . the proposed framework is based on a framework that plans the poem sketch before decoding .
GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion (2025.findings-acl)

Copied to clipboard

Challenge: Existing knowledge graphs lack the ability to integrate structural information into LLMs and output predictions deterministically.
Approach: They propose a method which encodes structural information of KGs and merges it with LLMs to enhance KGC performance.
Outcome: The proposed method improves the performance of KG Completion datasets on KGs by integrating structural information with LLMs.
Go Back in Time: Generating Flashbacks in Stories with Event Temporal Prompts (2022.naacl-main)

Copied to clipboard

Challenge: Existing systems that generate *flashbacks* are monotonic and lack explicit guidance on how to insert them.
Approach: They propose to use event temporal orders to encode events as temporal prompts . they leverage a Plan-and-Write framework enhanced by reinforcement learning to generate storylines .
Outcome: The proposed method generates more interesting stories with *flashbacks* while maintaining textual diversity, fluency, and temporal coherence.
SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Language models evolve to tackle complex, multifaceted tasks, requiring granular evaluations . recent studies have focused on leaderboard and benchmark results, but limited interpretability makes it difficult to compare strengths and weaknesses of models.
Approach: They propose an unsupervised tree-structured diagnosis framework for understanding model proficiency in specific abilities with an LLM as a judge.
Outcome: The proposed framework improves model in-context learning and predicts model weaknesses with a 55% success rate compared to the framework without SkillVerse.
AmbiPun: Generating Humorous Puns with Ambiguous Context (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for generating homographic puns are heavy-weighted due to the lack of training data.
Approach: They propose a way to generate pun sentences that does not require training on existing puns.
Outcome: The proposed method outperforms baseline models and state-of-the-art models by a large margin.
MacGyver: Are Large Language Models Creative Problem Solvers? (2024.naacl-long)

Copied to clipboard

Challenge: a new study examines the creative problem-solving capabilities of modern LLMs . it provides insight into the constrained problem- solving capabilities of both humans and AI .
Approach: They use an automatically generated dataset to compare and contrast LLMs and humans to find out their creative problem-solving abilities.
Outcome: The proposed dataset compares LLMs and humans in a constrained setting . it shows that humans excel in tasks they are familiar with but struggle with domain-specific knowledge .
Paraphrase Generation as Unsupervised Machine Translation (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for paraphrase generation rely on labeled datasets or are limited in narrow domains.
Approach: They propose a paradigm for paraphrase generation by treating the task as unsupervised machine translation based on pairs of unlabeled monolingual sentences.
Outcome: The proposed paradigm can generate paraphrases on a large unlabeled monolingual corpus without relying on bilingual sentence pairs.
A Unified Framework for Pun Generation with Humor Principles (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing models for generating homophonic and homographic puns lack the linguistic attributes of successful puns to resolve the split-up in existing work.
Approach: They propose a framework to generate both homophonic and homographic puns to resolve the split-up in existing works by incorporating three linguistic attributes of puns into the language models: ambiguity, distinctiveness, and surprise.
Outcome: The proposed model over strong baselines shows that it can generate both homophonic and homographic puns.
Rethinking Creativity Evaluation: A Critical Analysis of Existing Creativity Evaluations (2026.eacl-long)

Copied to clipboard

Challenge: Creativity measures that distinguish creativity in one domain fail in others, and different metrics disagree on the same data points.
Approach: They examine, analyze, and compare four representative creativity measures across the diverse creative domains, including creative writing, unconventional problem-solving, and research ideation.
Outcome: The measures of creativity across creative domains are compared using a set of human-aligned examples and lack consistency across domains and metrics.
Unsupervised Melody-to-Lyrics Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for automatic melody-to-lyric generation are limited due to the limited amount of melody-lyrical aligned data.
Approach: They propose a method for automatic melody-to-lyric generation without training on any aligned melody-lyr data.
Outcome: The proposed model generates high-quality lyrics that are singable, intelligible, and coherent than baseline models.
HypoGen: Hyperbole Generation with Commonsense and Counterfactual Knowledge (2021.findings-emnlp)

Copied to clipboard

Challenge: despite its abundance, the computational explorations of hyperboles remain under-explored.
Approach: They propose a sentence-level hyperbole generation method that leverages commonsense and counterfactual inference to generate hyperbolic candidates based on the results.
Outcome: The proposed method generates hyperboles with high success rate, intensity, funniness, and creativity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations