Challenge: Current language models generate high-quality text, but are they copying it or have they learned generalizable linguistic abstractions?
Approach: They propose a suite of analyses for assessing the novelty of generated text . they focus on sequential structure (n-grams) and syntactic structure (syntactical structure).
Outcome: The proposed model-generated text is as novel as the baseline human-generated model- generated text, but it is copied substantially, the authors show .

Similar Papers

Language Model Evaluation Beyond Perplexity (2021.acl-long)

Copied to clipboard

Challenge: a nascent literature on probing language models has focused on studying linguistic phenomena.
Approach: They propose a framework for evaluating the fit of language models to natural language tendencies.
Outcome: The proposed framework evaluates language models to the tendencies of natural language . it shows that the models learn only a subset of the tendancies considered .
The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text (2024.findings-naacl)

Copied to clipboard

Challenge: a new study examines the effects of training language models on synthetic data generated by their predecessors.
Approach: They propose to use recursive finetuning techniques to assess linguistic diversity of models.
Outcome: The proposed metrics show a decrease in diversity of model outputs through successive iterations, especially for tasks demanding high levels of creativity.
Overestimation of Syntactic Representation in Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Several testing methodologies have been developed to probe models’ syntactic representations.
Approach: They propose a method to determine syntactic structure by training a model on strings generated according to a template and testing its ability to distinguish between similar ones with different syntax.
Outcome: The proposed method reproduces positive results with two non-syntactic baseline language models: an n-gram model and an LSTM model trained on scrambled inputs.
Evaluating n-Gram Novelty of Language Models Using Rusty-DAWG (2024.emnlp-main)

Copied to clipboard

Challenge: a new study examines how novel language models generate training text . large LMs and constrained decoding strategies both decrease novelty .
Approach: They develop a novel search tool inspired by genomic data to find n-grams in training data.
Outcome: The proposed tool can search for n-grams over a corpus in constant time w.r.t. large LMs and more constrained decoding strategies both decrease novelty.
The Amazing World of Neural Language Generation (2020.emnlp-tutorials)

Copied to clipboard

Challenge: Recent years have seen a paradigm shift in neural text generation due to advances in deep contextual language modeling and transfer learning.
Approach: They will discuss how and why NLG models succeed/fail at generating coherent text.
Outcome: This paper will discuss how and why these models succeed/fail at generating coherent text, and provide insights on several applications.
On the Multilingual Capabilities of Very Large-Scale English Language Models (2022.lrec-1)

Copied to clipboard

Challenge: Generative Pre-trained Transformers (GPTs) have been scaled to unprecedented sizes in the history of machine learning.
Approach: They investigate the potential and limits of Generative Pre-trained Transformers in three tasks . they find it can be almost as useful for many languages as it is for English .
Outcome: The proposed model can perform tasks in five different languages, and its potential is explored . it can learn from a few examples "via text interaction" and is scalable to many languages .
Synthetic Data in the Era of Large Language Models (2025.acl-tutorials)

Copied to clipboard

Challenge: 'synthetic data' is a data generated with the assistance of large language models to make dataset construction faster and cheaper.
Approach: This tutorial seeks to build a shared understanding of recent progress in synthetic data generation from NLP and related fields by grouping and describing major methods, applications, and open problems.
Outcome: This tutorial will describe methods, applications, and open problems that have been developed and are being used to improve the quality and efficiency of synthetic data generation.
Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on using LLMs to classify text as either human-written or machine-generated .
Approach: They characterize human-written and machine-generated texts using a set of linguistic features across different linguistic levels such as morphology, syntax, and semantics.
Outcome: The proposed model reveals that human-written texts exhibit simpler syntactic structures and more diverse semantic content.
Evaluating Language Models as Synthetic Data Generators (2025.acl-long)

Copied to clipboard

Challenge: Prior studies have focused on developing effective data generation methods, but lack systematic comparison of different LMs as data generators in a unified setting.
Approach: They propose to use a benchmark to compare language models' data generation abilities against a set of standardized settings and metrics.
Outcome: The proposed benchmark provides standardized settings and metrics to evaluate LMs’ data generation abilities.
Low-Perplexity LLM-Generated Sequences and Where To Find Them (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied across various domains, but the ways they leverage their training data during inference remains only partially understood.
Approach: They propose a systematic approach that analyzes low-perplexity sequences and traces them back to their sources in the training data.
Outcome: The proposed pipeline extracts low-perplexity sequences across diverse topics while avoiding degeneration, then trace them back to their sources in the training data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations