Challenge: The paper extends the Data Movement Distance (DMD) metric defined to measure the locality in computer memory to text by defining a new term designed to better characterize low-frequency tokens.
Approach: They propose to define a normalized version of the Data Movement Distance (nDMD) term is designed to better characterize low-frequency tokens.
Outcome: The proposed normalized version outperforms baselines and improves performance on the English subset of the M4 dataset and the GenAI detection shared task.

Similar Papers

A Large-Scale Comparison of Historical Text Normalization Systems (N19-1)

Copied to clipboard

Challenge: a large study of historical text normalization is done on eight languages . there is no consensus on the state-of-the-art approach to normalization .
Approach: They present a large study of historical text normalization done on eight languages . they evaluate four different systems based on supervised learning on datasets from eight different languages based in the literature .
Outcome: The proposed methods are based on supervised learning and are available online.
Linguistic and Embedding-Based Profiling of Texts Generated by Humans and Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on using LLMs to classify text as either human-written or machine-generated .
Approach: They characterize human-written and machine-generated texts using a set of linguistic features across different linguistic levels such as morphology, syntax, and semantics.
Outcome: The proposed model reveals that human-written texts exhibit simpler syntactic structures and more diverse semantic content.
EvoBench: Towards Real-world LLM-Generated Text Detection Benchmarking for Evolving Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods to detect LLM-generated texts rely on static benchmarks that neglect the evolving nature of LLMs.
Approach: They propose a benchmark to evaluate the generalization of LLM-generated text detection methods.
Outcome: The proposed benchmark measures generalization of 14 detection methods across LLMs.
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech (2025.emnlp-industry)

Copied to clipboard

Challenge: Text Normalization (TN) is a key preprocessing step in Text-to-Speech systems.
Approach: They propose a prompt-based approach to TN using Large Language Models (LLMs) they propose scalable experimentation across languages to reduce the reliance on manual rules .
Outcome: The proposed approach reduces the reliance on manual rules and enables broader linguistic applicability with minimal human intervention across eight languages.
QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings (2025.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have revolutionized natural language processing, but their success remains limited to high-resource domains.
Approach: They analyze the coverage and representativeness of specialized-domain QA datasets against large-scale reference datasets.
Outcome: The proposed methods and evaluations highlight the challenges faced by LLMs in low-resource domains.
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MarkerGen (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to control text length are lacking in LCTG, posing a major limitation for practical applications.
Approach: They propose a plug-and-play approach that decomposes LCTG sub-abilities with human patterns as reference and performs detailed error analysis.
Outcome: The proposed method significantly improves LCTG across various settings, exhibiting outstanding effectiveness and generalizability.
Learning to Rewrite: Generalized LLM-Generated Text Detection (2025.acl-long)

Copied to clipboard

Challenge: Existing detectors for Large Language Models (LLMs) struggle to generalize in open-world settings.
Approach: They propose a framework to detect LLM-generated text with exceptional generalization to unseen domains by reinforcing LLMs’ inherent rewriting tendencies.
Outcome: The proposed framework outperforms state-of-the-art detection methods by 23.04% in AUROC, 35.10% for out-of distribution tests, and 48.66% under adversarial attacks.
GPT-who: An Information Density-based Machine-Generated Text Detector (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) generate misinformation, memorized content, plagiarized content, toxic speech, and hallucinated content.
Approach: They propose a statistical detector that uses UID to model the unique statistical signature of each LLM and human author for accurate detection.
Outcome: The proposed method outperforms state-of-the-art detectors by over 20% across domains.
Out-of-Distribution Generalization in Natural Language Processing: Past, Present, and Future (2023.emnlp-main)

Copied to clipboard

Challenge: Existing literature on the generalization of machine learning models to out-of-distribution data is lacking.
Approach: They propose to present the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
Outcome: The proposed survey provides the first comprehensive review of recent progress, methods, and evaluations on the generalization challenge from an OOD perspective in natural language understanding.
HLU: Human Vs LLM Generated Text Detection Dataset for Urdu at Multiple Granularities (2025.coling-main)

Copied to clipboard

Challenge: Using large language models (LLMs) to generate human-like text has raised concerns about misuse, especially in low-resource languages like Urdu.
Approach: They propose a dataset that contains documents, paragraphs, and sentences . they conducted human evaluations and automated evaluations .
Outcome: The proposed dataset shows that distinguishing between human and machine-generated text is challenging for both humans and LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations