Challenge: Existing metrics to quantify lexical diversity have been proposed.
Approach: They propose to examine how generic language characteristics are impacted by text alterations.
Outcome: The proposed models show that lexical features are more sensitive to text modifications than syntactic ones.

Similar Papers

Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models excel in varied tasks, but their mechanisms remain unclear . current LLMs' development has put this assumption in jeopardy, authors say .
Approach: They examine whether LLMs learn generalized linguistic abstraction or rely on surface-level features that match their pre-training data.
Outcome: The proposed model can learn generalized linguistic abstraction or rely on surface-level features that match their pre-training data.
Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change Detection (2021.eacl-main)

Copied to clipboard

Challenge: Lexical semantic change detection is a new and innovative research field.
Approach: They propose to pre-train on large corpora and refine on diachronic target corpors to improve performance.
Outcome: The proposed models improve on large corpora and diachronic target corpors . the proposed models are compared with existing models in a variety of learning scenarios .
The impact of lexical and grammatical processing on generating code from natural language (2022.findings-acl)

Copied to clipboard

Challenge: Yin and Neubig (2018) identify four key components of importance for natural language to code translation.
Approach: They propose a seq2seq-based architecture that relies on a grammar-based decoder and a lexical substitution component for natural language to code translation.
Outcome: The proposed architecture relies on a grammar-based decoder and a BERT encoder . the proposed architecture is based on lexical substitutions in natural language to code translation .
Inducing Generalizable and Interpretable Lexica (2022.findings-emnlp)

Copied to clipboard

Challenge: Lexica are widely used as generalizable language features to predict sentiment, emotions, mental health, and personality.
Approach: They propose to induce lexica using context-oblivious and context-aware approaches and compare their performance using crowd-worker assessment.
Outcome: The proposed models can be induced using context-oblivious and context-aware approaches and evaluate their quality using crowd-worker assessment.
Exploring hybrid approaches to readability: experiments on the complementarity between linguistic features and transformers (2024.findings-eacl)

Copied to clipboard

Challenge: Linguistic features have been a key component of the automatic assessment of text readability (ARA) with the development in the ARA field, the research moved to Deep Learning (DL)
Approach: They compare 6 hybrid approaches to Machine Learning and DL on 4 corpora and found they are the most robust on smaller datasets and across languages.
Outcome: The proposed approaches perform better on smaller datasets and across languages and tasks.
Genre Matters: How Text Types Interact with Decoding Strategies and Lexical Predictors in Shaping Reading Behavior (2025.emnlp-main)

Copied to clipboard

Challenge: eMTeC is the first eye-tracking corpus of LLM-generated texts . it shows that text type strongly modulates cognitive effort during reading .
Approach: They use the first eye-tracking corpus of LLM-generated texts to study eye movements during reading and how decoding strategies interact with text types to shape reading behavior.
Outcome: The first eye-tracking corpus of LLM-generated texts shows that text type strongly modulates cognitive effort during reading and that word-level psycholinguistic effects vary systematically across genres.
Investigating Ensemble Methods for Model Robustness Improvement of Text Classifiers (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to reduce model's reliance on bias features ignore the learnability of these features.
Approach: They propose to reduce models' reliance on bias features by first training models with fixed low-capacity models which ignore the learnability of the bias features.
Outcome: The proposed models can perform better on out-of-distribution datasets than baseline models with a more sophisticated model design.
TinyAttack: Exploring Stylistic Vulnerabilities in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on robustness of large language models has focused on text-based perturbations and the use of invisible characters and homoglyphs.
Approach: They propose a framework to exploit weaknesses in Large Language Models (LLMs) by changing their stylistic structure using Unicode.
Outcome: The proposed framework exploits vulnerabilities in large language models through Unicode-based stylistic transformations without altering its semantic or syntactic structure.
When Format Changes Meaning: Investigating Semantic Inconsistency of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Large language models are vulnerable to semantic inconsistency, a study finds . minor formatting variations result in divergent predictions for semantically equivalent inputs.
Approach: They evaluate LLMs for semantic inconsistency and find they remain vulnerable . they propose to use mechanistic analysis to develop models that improve their reliability .
Outcome: The proposed model is vulnerable to semantic inconsistency, the authors show . their model is brittle even in state-of-the-art models, they say .
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks.
Approach: They propose a benchmark for textual adversarial defence that evaluates state-of-the-art defence mechanisms across diverse datasets, models, and tasks.
Outcome: The proposed benchmark incorporates a wide range of datasets and evaluates state-of-the-art defence mechanisms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations