Challenge: a statistical analysis of arXiv paper abstracts shows a marked drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024.
Approach: They report a drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024.
Outcome: The frequency of words previously identified as overused by ChatGPT, such as “delve”, has instead kept increasing.

Similar Papers

Why Does ChatGPT “Delve” So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Scientific English is currently undergoing rapid change, with words like “delve,” “intricate,” and “underscore” appearing far more frequently than just a few years ago.
Approach: They propose a formal method to characterize scientific English linguistic changes . they propose lexical overrepresentation by reinforcement learning from human feedback .
Outcome: The proposed method yields 21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usage.
The Impact of Large Language Models in Academia: from Writing to Speaking (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are impacting human society, especially in textual information.
Approach: They propose to build an automated monitoring platform to track the impact of large language models on human expression.
Outcome: The results show that LLM-style words such as significant are used more frequently in abstracts and oral presentations.
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance .
Approach: They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy.
Outcome: The proposed model will be able to detect human-written content in real time.
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models have shown capabilities close to human performance in various analytical tasks.
Approach: They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions .
Outcome: The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks.
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers (2024.naacl-long)

Copied to clipboard

Challenge: Recent advances in language modeling have caused disruptive shifts throughout AI research, spurring discussion about how the field is changing and how it should change.
Approach: They analyze a dataset of 16,979 LLM-related arXiv papers and examine industry and academic publishing trends.
Outcome: The authors examine the impact of large language models on AI research in 2023 and 2022.
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)

Copied to clipboard

Challenge: Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity).
Approach: They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions.
Outcome: The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization.
Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) require accurate text detection, but authors' characteristics are neglected.
Approach: They investigate how author characteristics impact AI-generated text detection . they use corpus of human-authored texts and parallel AI-generated texts .
Outcome: The results show that gender, CEFR proficiency, academic field and language environment influence detector accuracy.
Human Alignment: How Much Do We Adapt to LLMs? (2025.acl-short)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming a common part of our lives, yet few studies have examined how they influence our behavior.
Approach: They propose a cooperative language game in which players aim to converge on a word and play a game in a group.
Outcome: The proposed game shows that humans notice and adapt to differences regardless of whether they are aware they are interacting with an LLM.
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have raised concerns about the potential threats large language models pose to academic integrity and copyright protection.
Approach: They propose a dataset of 46.5K synthetic text pairs that represent three major types of plagiarism: verbatim copying, paraphrasing, and summarization.
Outcome: The proposed dataset shows that GPT-3.5 Turbo can produce high-quality paraphrases and summaries without significantly increasing text complexity compared to GPT-4 Turbo.
On the Generalization of Training-based ChatGPT Detection Methods (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that training-based methods are ineffective to detect LLM generated texts from unseen tasks or topics which are not collected during training.
Approach: They propose to train classification models to distinguish LLMs from human texts by a distribution shift caused by prompts, text lengths, topics, and language tasks.
Outcome: The proposed methods can detect LLMs from black-box models, but they suffer from distribution shifts due to a wide range of factors, including prompts, text lengths, topics, and language tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations