Challenge: Existing methods for controlling the generation of pre-trained language models infuse domain bias into the generation process, making it difficult to generate out-of-domain texts.
Approach: They propose a retrieval-augmented generation framework that uses retrieval to generate fluent sentences with high attribute relevance.
Outcome: The proposed method can generate fluent sentences with high attribute relevance while keeping domain bias out of the model.

Similar Papers

Retrieval-Augmented Controllable Review Generation (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to generate reviews using attribute identifiers are limited and dependent on how well they can capture vector representations of attributes.
Approach: They propose to leverage attributes as inputs for review generation by using reference sets . they propose to use these references to enrich inductive biases of given attributes .
Outcome: The proposed model improves over previous approaches on automatic and human evaluation metrics.
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP (2024.findings-naacl)

Copied to clipboard

Challenge: a low-resource dataset is limited in training data, so generating task-specific data is challenging.
Approach: They propose a data augmentation technique that prompts off-the-shelf instruction-following Large Language Models to generate augmentations.
Outcome: The proposed technique outperforms baselines on 11 datasets spanning 3 tasks and 3 low-resource settings.
Tailor: A Soft-Prompt-Based Approach to Attribute-Based Controlled Text Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing work focuses on generating sentences satisfying pre-specified attributes such as topic and sentiment, yet suffers from increases in storage and inference time.
Approach: They propose a method that uses a pre-trained continuous vector to generate a fixed pre-trainable language model to satisfy a specified attribute.
Outcome: The proposed model can achieve improvements on eleven attribute-specific generation tasks with 0.08% extra training parameters.
Attribute Alignment: Controlling Text Generation from Pre-trained Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: Large language models can generate text with sentiment polarity or specific topics without changing the original model parameters.
Approach: They propose a method for controlling text generation by aligning disentangled attribute representations.
Outcome: The proposed method shows large performance gains while maintaining diversity and fluency.
CFL: Causally Fair Language Models Through Token-level Attribute Controlled Generation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to control attributes of Language Models (LMs) for text generation are not safe, as toxicity and bias goals are opposed to each other.
Approach: They propose a method to control the attributes of Language Models (LMs) for the text generation task using Causal Average Treatment Effect (ATE) scores and counterfactual augmentation.
Outcome: The proposed architecture achieves state of the art performance for toxic degeneration, which are computed using Real Toxicity Prompts.
RSA-Control: A Pragmatics-Grounded Lightweight Controllable Text Generation Framework (2024.emnlp-main)

Copied to clipboard

Challenge: RSA-Control is a training-free controllable text generation framework . existing studies rely on fine-tuning pre-trained language models . external components could hurt coherence and accuracy of the model .
Approach: They propose a training-free controllable text generation framework grounded in pragmatics that directs the generation process by recursively reasoning between imaginary speakers and listeners.
Outcome: The proposed framework achieves strong attribute control while maintaining fluency and content consistency.
RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within Generation (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on separate retrievers to fetch top-k text chunks for generating evidence, and they lack joint optimization.
Approach: They propose a framework that integrates retrieval and generation into a single, auto-regressive process, enabling LLMs to directly generate fine-grained evidence from the corpus with constrained decoding.
Outcome: Extensive experiments on five open-domain QA datasets demonstrate the proposed framework’s superior performance across both in-domain and out-of-domain tasks.
Controllable Meaning Representation to Text Generation: Linearization and Data Augmentation Strategies (2020.emnlp-main)

Copied to clipboard

Challenge: Using task-oriented dialogue generation benchmarks, we compare the effect of four input linearization strategies on controllability and faithfulness.
Approach: They compare the effect of four input linearization strategies on controllability and faithfulness . they also evaluate how a phrase-based data augmentation method can improve performance .
Outcome: The proposed model can generate utterances whose phrases follow the order of the provided plan.
Rethinking Data Augmentation in Text-to-text Paradigm (2022.coling-1)

Copied to clipboard

Challenge: Existing approaches to augment training data are limited or marginal, or even diminishing or adverse especially given original training corpus is relatively sufficient or the backbone classifiers are PLM based.
Approach: They propose to integrate text-to-text language models and construct a new two-phase framework for augmentation using two novel schemes.
Outcome: The proposed framework synthesizes new samples benefiting from the knowledge learned from pre-trained language models on two public classification datasets and shows remarkable gains.
Facts2Story: Controlling Text Generation by Key Facts (2020.coling-main)

Copied to clipboard

Challenge: Existing methods for story generation struggle with staying coherent for long periods of time.
Approach: They propose a controlled generation task which expands a sequence of facts into a longer narrative.
Outcome: The proposed model produces competitive fluency while adhering to the requested facts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations