Challenge: Scientific English is currently undergoing rapid change, with words like “delve,” “intricate,” and “underscore” appearing far more frequently than just a few years ago.
Approach: They propose a formal method to characterize scientific English linguistic changes . they propose lexical overrepresentation by reinforcement learning from human feedback .
Outcome: The proposed method yields 21 focal words whose increased occurrence in scientific abstracts is likely the result of LLM usage.

Similar Papers

Human-LLM Coevolution: Evidence from Academic Writing (2025.findings-acl)

Copied to clipboard

Challenge: a statistical analysis of arXiv paper abstracts shows a marked drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024.
Approach: They report a drop in the frequency of several words previously identified as overused by ChatGPT, such as “delve”, starting soon after they were pointed out in early 2024.
Outcome: The frequency of words previously identified as overused by ChatGPT, such as “delve”, has instead kept increasing.
Lexical Semantics with Large Language Models: A Case Study of English “break” (2023.findings-eacl)

Copied to clipboard

Challenge: Large neural language models (LLMs) can be powerful tools for research in lexical semantics.
Approach: They argue that large neural language models can be powerful tools for research in lexical semantics by capturing known sense distinctions and identifying informative new sense combinations.
Outcome: The proposed models capture many of the sense distinctions found in the English verb break and can be used to identify informative new sense combinations for further analysis.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.
Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) excel on public benchmarks, but high scores may mask overreliance on dataset-specific surface cues rather than true language understanding.
Approach: They propose a meta-evaluation framework that systematically rephrases benchmark inputs to detect overfitting.
Outcome: The proposed framework detects performance degradation indicative of superficial pattern reliance on dataset-specific cues and distortion levels.
Geographical Erasure in Language Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models encode vast amounts of world knowledge but are at risk of inordinately capturing information about dominant groups.
Approach: They propose to operationalise a form of geographical erasure wherein language models underpredict certain countries.
Outcome: The proposed model underpredicts certain countries by a factor 3 . the model is based on large datasets and is able to mitigate the effects .
Unveiling the Generalization Power of Fine-Tuned Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated exceptional multitasking abilities, but the comprehensive effects of fine-tuning on the LLMs’ generalization ability are not fully understood.
Approach: They conduct extensive experiments across five distinct language tasks on different datasets to investigate whether fine-tuning affects the generalization ability intrinsic to LLMs.
Outcome: The proposed model can generalize to different domains and tasks by integrating the in-context learning strategy during fine-tuning on generation tasks.
Lost in the Distance: Large Language Models Struggle to Capture Long-Distance Relational Knowledge (2025.findings-naacl)

Copied to clipboard

Challenge: Recent large language models have demonstrated impressive capabilities in handling long contexts . however, as context length increases, LLMs struggle more with filtering out irrelevant information .
Approach: They propose to use unrelated sentences to capture relational knowledge over long contexts . they find that LLMs can handle edge noise with little impact, but can reason about distant relationships .
Outcome: The proposed model can handle edge noise with little impact, but its ability to reason about distant relationships declines as the noise grows.
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English.
Approach: They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages.
Outcome: The proposed model massively underperforms purpose-built systems, particularly in English.
Model Editing Harms General Abilities of Large Language Models: Regularization to the Rescue (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods that edit large language models with updated knowledge can cause side effects on the general abilities of LLMs such as reasoning, natural language inference, and question answering.
Approach: They propose to regularize the edit update weights by imposing constraints on their complexity based on the RElative Change in weighT.
Outcome: The proposed method can significantly mitigate the side effects while maintaining over 94% editing performance.
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks have exposed patterns and may not truly assess generalization ability of Large Language Models (LLMs).
Approach: They propose a “Generalization Stress Test” to assess Large Language Models’ generalization ability under slight and controlled perturbations, including option length, problem types, and irrelevant noun replacements.
Outcome: The proposed test shows that LLMs exhibit severe accuracy drops and unexpected biases when faced with minor but content-preserving modifications.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations