Challenge: Modern neural language models achieve high accuracy in text generation, yet precise control over generation length remains underdeveloped.
Approach: They propose a method to provide robust length control using Reverse Positional Embeddings.
Outcome: The proposed method provides stable length fidelity without degrading text accuracy . the proposed method generalizes well to unseen target lengths .

Similar Papers

Positional Encoding to Control Output Sequence Length (N19-1)

Copied to clipboard

Challenge: Neural encoder-decoder models have been successful in natural language generation tasks, but they must be limited to a specified length for abstractive summarization.
Approach: They propose a sinusoidal positional extension to preserve the length constraint so that a neural encoder-decoder model can generate a text of any length even if the target length is unseen in training data.
Outcome: The proposed method can generate a text of any length even if the target length is unseen in training data and improves ROUGE scores.
LenAtten: An Effective Length Controlling Unit For Text Summarization (2021.findings-acl)

Copied to clipboard

Challenge: Fixed length summarization (FLS) requires generating summaries with a preset number of characters or words.
Approach: They propose a length control unit called LenAtten to break this trade-off by generating a short and coherent summary with the target length.
Outcome: The proposed model improves controllability and ROGUE scores and generalizes well.
Towards More Efficient Insertion Transformer with Fractional Positional Encoding (2023.eacl-main)

Copied to clipboard

Challenge: Empirical studies on text generation tasks demonstrate the effectiveness of insertion-based models.
Approach: They propose a reusable positional encoding scheme for insertion transformers that allows reusing representations calculated in previous steps.
Outcome: Empirical studies show that the proposed model reduces the time required to generate a token and improves decoding efficiency.
DcLM: Output Length Control of Large Language Models via Dynamic Length Markers (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have limited awareness of output length, making it difficult to satisfy precise length requirements.
Approach: They propose a model-agnostic approach that introduces dynamic length markers to guide length-controllable outputs.
Outcome: The proposed method significantly reduces length deviation across multiple datasets.
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGen (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to control text length are lacking in LCTG, posing a major limitation for practical applications.
Approach: They propose a plug-and-play approach that decomposes LCTG sub-abilities with human patterns as reference and performs detailed error analysis.
Outcome: The proposed method significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability.
Length-Induced Embedding Collapse in PLM-based Models (2025.acl-long)

Copied to clipboard

Challenge: In text embeddings from PLMs are essential for many NLP applications, but performance degrades on longer texts.
Approach: They propose a method which mitigates the phenomenon of Length Collapse . they propose TempScale to ensure more consistent embeddings across different text lengths .
Outcome: The proposed method improves performance on MTEB and LongEmbed by 0.94% on short and 1.10% on long texts.
Stable Language Model Pre-training by Reducing Embedding Variability (2024.emnlp-main)

Copied to clipboard

Challenge: Stable pre-training is essential for achieving better-performing language models, but tracking pre-train stability is impractical due to high computational costs.
Approach: They propose to use Token Embedding Variability as a proxy to estimate pre-training stability.
Outcome: The proposed method improves stability and lowers perplexities even at deeper layer counts.
Plan-then-Generate: Controlled Data-to-Text Generation via Planning (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies focus on producing results that are close to the references, i.e. what to generate and in what order (the output structure) cannot be explicitly controlled by the users.
Approach: They propose a Plan-then-Generate framework to improve the controllability of neural data-to-text models.
Outcome: The proposed model can control both the intra-sentence and inter-sentent structure of the generated output.
Investigating the Effect of Relative Positional Embeddings on AMR-to-Text Generation with Structural Adapters (2023.eacl-main)

Copied to clipboard

Challenge: Recent approaches to text generation from Abstract Meaning Representation (AMR) have been based on neural-centered encoderdecoder architectures.
Approach: They propose a structure-aware adapter which injects the input graph connectivity within PLMs using Graph Neural Networks.
Outcome: The proposed adapter is robust to a variety of approaches and can be used to generate Graph-to-Text representations.
CIE: Controlling Language Model Text Generations Using Continuous Signals (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to control language models with intent are brittle and hard to scale.
Approach: They propose to use a set of LMs to fine-tune to expect a control vector that is interpolated between a "low" and a 'high' token embedding.
Outcome: The proposed method can be finetuned to expect a control vector that is interpolated between a “low” and a ‘high” token embedding.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations