| Challenge: | a statistical analysis of arXiv paper abstracts shows a marked drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024. |
| Approach: | They report a drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024. |
| Outcome: | The frequency of words previously identified as overused by ChatGPT, such as “delve”, has instead kept increasing. |
Similar Papers
Why Does ChatGPT “Delve” So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Scientific English is currently undergoing rapid change, with words like “delve,” “intricate,” and “underscore” appearing far more frequently than just a few years ago. |
| Approach: | They propose a formal method to characterize scientific English linguistic changes . they propose lexical overrepresentation by reinforcement learning from human feedback . |
| Outcome: | The proposed method yields 21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usage. |
The Impact of Large Language Models in Academia: from Writing to Speaking (2025.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impacting human society, especially in textual information. |
| Approach: | They propose to build an automated monitoring platform to track the impact of large language models on human expression. |
| Outcome: | The results show that LLM-style words such as significant are used more frequently in abstracts and oral presentations. |
A Survey on Detection of LLMs-Generated Content (2024.findings-emnlp)
Copied to clipboard
Xianjun Yang, Liangming Pan, Xuandong Zhao, Haifeng Chen, Linda Petzold, William Yang Wang, Wei Cheng
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
The LLM Effect: Are Humans Truly Using LLMs, or Are They Being Influenced By Them Instead? (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have shown capabilities close to human performance in various analytical tasks. |
| Approach: | They investigate the efficiency and accuracy of Large Language Models in specialized tasks . they integrate LLMs with expert annotators to observe the impact of LLM suggestions . |
| Outcome: | The proposed model improves task completion speed but introduces anchoring bias . the proposed model is not suitable for open-ended analysis, but is capable of handling specialized tasks. |
Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent advances in language modeling have caused disruptive shifts throughout AI research, spurring discussion about how the field is changing and how it should change. |
| Approach: | They analyze a dataset of 16,979 LLM-related arXiv papers and examine industry and academic publishing trends. |
| Outcome: | The authors examine the impact of large language models on AI research in 2023 and 2022. |
People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated text (2025.acl-long)
Copied to clipboard
| Challenge: | Qualitative analysis of experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity). |
| Approach: | They hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. |
| Outcome: | The annotators who frequently use LLMs for writing tasks outperform commercial and open-source detectors even without evasion tactics like paraphrasing and humanization. |
Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) require accurate text detection, but authors' characteristics are neglected. |
| Approach: | They investigate how author characteristics impact AI-generated text detection . they use corpus of human-authored texts and parallel AI-generated texts . |
| Outcome: | The results show that gender, CEFR proficiency, academic field and language environment influence detector accuracy. |
Human Alignment: How Much Do We Adapt to LLMs? (2025.acl-short)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are becoming a common part of our lives, yet few studies have examined how they influence our behavior. |
| Approach: | They propose a cooperative language game in which players aim to converge on a word and play a game in a group. |
| Outcome: | The proposed game shows that humans notice and adapt to differences regardless of whether they are aware they are interacting with an LLM. |
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have raised concerns about the potential threats large language models pose to academic integrity and copyright protection. |
| Approach: | They propose a dataset of 46.5K synthetic text pairs that represent three major types of plagiarism: verbatim copying, paraphrasing, and summarization. |
| Outcome: | The proposed dataset shows that GPT-3.5 Turbo can produce high-quality paraphrases and summaries without significantly increasing text complexity compared to GPT-4 Turbo. |
On the Generalization of Training-based ChatGPT Detection Methods (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies show that training-based methods are ineffective to detect LLM generated texts from unseen tasks or topics which are not collected during training. |
| Approach: | They propose to train classification models to distinguish LLMs from human texts by a distribution shift caused by prompts, text lengths, topics, and language tasks. |
| Outcome: | The proposed methods can detect LLMs from black-box models, but they suffer from distribution shifts due to a wide range of factors, including prompts, text lengths, topics, and language tasks. |