Why Does ChatGPT “Delve” So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Scientific English is currently undergoing rapid change, with words like “delve,” “intricate,” and “underscore” appearing far more frequently than just a few years ago. |
| Approach: | They propose a formal method to characterize scientific English linguistic changes . they propose lexical overrepresentation by reinforcement learning from human feedback . |
| Outcome: | The proposed method yields 21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usage. |
Similar Papers
Human-LLM Coevolution: Evidence from Academic Writing (2025.findings-acl)
Copied to clipboard
| Challenge: | a statistical analysis of arXiv paper abstracts shows a marked drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024. |
| Approach: | They report a drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024. |
| Outcome: | The frequency of words previously identified as overused by ChatGPT, such as “delve”, has instead kept increasing. |
Lexical Semantics with Large Language Models: A Case Study of English “break” (2023.findings-eacl)
Copied to clipboard
| Challenge: | Large neural language models (LLMs) can be powerful tools for research in lexical semantics. |
| Approach: | They argue that large neural language models can be powerful tools for research in lexical semantics by capturing known sense distinctions and identifying informative new sense combinations. |
| Outcome: | The proposed models capture many of the sense distinctions found in the English verb break and can be used to identify informative new sense combinations for further analysis. |
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment. |
| Approach: | They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge. |
| Outcome: | The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses. |
Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) excel on public benchmarks, but high scores may mask overreliance on dataset-specific surface cues rather than true language understanding. |
| Approach: | They propose a meta-evaluation framework that systematically rephrases benchmark inputs to detect overfitting. |
| Outcome: | The proposed framework detects performance degradation indicative of superficial pattern reliance on dataset-specific cues and distortion levels. |
Geographical Erasure in Language Generation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models encode vast amounts of world knowledge but are at risk of inordinately capturing information about dominant groups. |
| Approach: | They propose to operationalise a form of geographical erasure wherein language models underpredict certain countries. |
| Outcome: | The proposed model underpredicts certain countries by a factor 3 . the model is based on large datasets and is able to mitigate the effects . |
Unveiling the Generalization Power of Fine-Tuned Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, but the comprehensive effects of fine-tuning on the LLMs’ generalization ability are not fully understood. |
| Approach: | They conduct extensive experiments across five distinct language tasks on different datasets to investigate whether fine-tuning affects the generalization ability intrinsic to LLMs. |
| Outcome: | The proposed model can generalize to different domains and tasks by integrating the in-context learning strategy during fine-tuning on generation tasks. |
Lost in the Distance: Large Language Models Struggle to Capture Long-Distance Relational Knowledge (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent large language models have demonstrated impressive capabilities in handling long contexts . however, as context length increases, LLMs struggle more with filtering out irrelevant information . |
| Approach: | They propose to use unrelated sentences to capture relational knowledge over long contexts . they find that LLMs can handle edge noise with little impact, but can reason about distant relationships . |
| Outcome: | The proposed model can handle edge noise with little impact, but its ability to reason about distant relationships declines as the noise grows. |
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)
Copied to clipboard
Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schuetze, Kemal Oflazer, David Mortensen
| Challenge: | Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English. |
| Approach: | They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages. |
| Outcome: | The proposed model massively underperforms purpose-built systems, particularly in English. |
Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods that edit large language models with updated knowledge can cause side effects on the general abilities of LLMs such as reasoning, natural language inference, and question answering. |
| Approach: | They propose to regularize the edit update weights by imposing constraints on their complexity based on the RElative Change in weighT. |
| Outcome: | The proposed method can significantly mitigate the side effects while maintaining over 94% editing performance. |
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing benchmarks have exposed patterns and may not truly assess generalization ability of Large Language Models (LLMs). |
| Approach: | They propose a “Generalization Stress Test” to assess Large Language Models’ generalization ability under slight and controlled perturbations, including option length, problem types, and irrelevant noun replacements. |
| Outcome: | The proposed test shows that LLMs exhibit severe accuracy drops and unexpected biases when faced with minor but content-preserving modifications. |