Diffusion Guided Language Modeling (2024.findings-acl)

Copied to clipboard

Challenge: Existing guidance methods for text generation are prone to decoding errors and degrade performance.
Approach: They propose a model that steers an auto-regressive language model to generate text with desired properties.
Outcome: The proposed model outperforms existing guidance methods on a wide range of benchmark data sets.

Similar Papers

A Plug-and-Play Method for Controlled Text Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for controlling language generation are not able to produce fluent text . current methods require additional models or fine-tuning to ensure specific words are included .
Approach: They propose a plug-and-play decoding method that allows for controlled language generation . they add a shift in the probability distribution over our vocabulary towards semantically similar words .
Outcome: The proposed method outperforms competing methods in human evaluations and does not impact fluency.
Improving Adversarial Text Generation by Modeling the Distant Future (2020.acl-main)

Copied to clipboard

Challenge: Recent work has shown excellent performance on text generation tasks by combining reinforcement learning (RL) and generative models.
Approach: They propose a model-based imitation-learning approach to improve text generation performance by focusing on a long horizon.
Outcome: The proposed model improves on a number of text-generation tasks and provides intermediate rewards for generator optimization.
Segment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models (2025.acl-long)

Copied to clipboard

Challenge: Diffusion models have shown promise in text generation, but often struggle with generating long, coherent, and contextually accurate text.
Approach: They propose a framework that enhances diffusion-based text generation through text segmentation, robust representation training with adversarial and contrastive learning, and improved latent-space guidance.
Outcome: The proposed framework improves diffusion-based text generation and improves scalability and fluency.
SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control (2023.acl-long)

Copied to clipboard

Challenge: Existing diffusion models for continuous-valued domains have not been adopted for text data.
Approach: They propose a diffusion-based language model with two key design choices . semi-autoregressive model generates blocks of text and allows local context updates . they evaluate it on unconstrained text generation benchmarks .
Outcome: The proposed model outperforms autoregressive models on unconstrained text generation benchmarks on uncontrolled text generation.
Attribute Alignment: Controlling Text Generation from Pre-trained Language Models (2021.findings-emnlp)

Copied to clipboard

Challenge: Large language models can generate text with sentiment polarity or specific topics without changing the original model parameters.
Approach: They propose a method for controlling text generation by aligning disentangled attribute representations.
Outcome: The proposed method shows large performance gains while maintaining diversity and fluency.
Why Generate When You Can Discriminate? A Novel Technique for Text Classification using Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods for text classification using autoregressive language models are limited . authors propose a novel technique for text classification using autoreregressives .
Approach: They propose a two-step technique for text classification using autoregressive language models . they use a set of perplexity and log-likelihood based numeric features to elicit a text instance .
Outcome: The proposed technique eliminates parameter updates in LMs and does not limit training examples . it is evaluated across 5 datasets and compares with multiple competent baselines .
Conditional [MASK] Discrete Diffusion Language Model (2025.emnlp-main)

Copied to clipboard

Challenge: Auto-regressive models excel in natural language processing but struggle to generate diverse text and lack controllability.
Approach: They propose entropy-adaptive Gibbs sampling and entropic-based noise scheduling to counterbalance each model’s shortcomings.
Outcome: The proposed framework outperforms baseline models and achieves the best quality-diversity tradeoff, demonstrating its effectiveness in non-autoregressive text generation.
Plug-in Language Model: Controlling Text Generation with a Simple Regression Model (2024.findings-naacl)

Copied to clipboard

Challenge: Large-scale pre-trained language models have demonstrated unrivaled capacity in generating text that closely resembles human-written content.
Approach: They propose a plug-in language model that leverages reinforcement learning to adjust latent states to control text generation.
Outcome: The proposed model outperforms existing methods that rely on gradient-based, weighted decoding, or prompt-based methods.
LanguageFlow: Advancing Diffusion Language Generation with Probabilistic Flows (2024.naacl-long)

Copied to clipboard

Challenge: Recent work has demonstrated success in controlling sentence attributes and structure based on diffusion language models.
Approach: They propose a language-rectified flow method that reformulates standard probabilistic flow models to learn ordinary differential equations to transport between the source and target distributions.
Outcome: The proposed method outperforms baselines on three fine-grained control tasks and multiple high-quality text editing tasks.
DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis.
Approach: They propose a controllable data synthesis framework based on variational autoencoder which leverages diffusion models to reserve more information of original distribution and format structure in the learned latent distribution.
Outcome: The proposed framework generates high-quality data with performance exceeding that of real data by 2%–7% on seven real-world datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations