Lexical Features Are More Vulnerable, Syntactic Features Have More Predictive Power (D19-55)
Copied to clipboard
| Challenge: | Existing metrics to quantify lexical diversity have been proposed. |
| Approach: | They propose to examine how generic language characteristics are impacted by text alterations. |
| Outcome: | The proposed models show that lexical features are more sensitive to text modifications than syntactic ones. |
Similar Papers
Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance (2026.eacl-long)
Copied to clipboard
| Challenge: | Large Language Models excel in varied tasks, but their mechanisms remain unclear . current LLMs' development has put this assumption in jeopardy, authors say . |
| Approach: | They examine whether LLMs learn generalized linguistic abstraction or rely on surface-level features that match their pre-training data. |
| Outcome: | The proposed model can learn generalized linguistic abstraction or rely on surface-level features that match their pre-training data. |
Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change Detection (2021.eacl-main)
Copied to clipboard
| Challenge: | Lexical semantic change detection is a new and innovative research field. |
| Approach: | They propose to pre-train on large corpora and refine on diachronic target corpors to improve performance. |
| Outcome: | The proposed models improve on large corpora and diachronic target corpors . the proposed models are compared with existing models in a variety of learning scenarios . |
The impact of lexical and grammatical processing on generating code from natural language (2022.findings-acl)
Copied to clipboard
| Challenge: | Yin and Neubig (2018) identify four key components of importance for natural language to code translation. |
| Approach: | They propose a seq2seq-based architecture that relies on a grammar-based decoder and a lexical substitution component for natural language to code translation. |
| Outcome: | The proposed architecture relies on a grammar-based decoder and a BERT encoder . the proposed architecture is based on lexical substitutions in natural language to code translation . |
Inducing Generalizable and Interpretable Lexica (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Lexica are widely used as generalizable language features to predict sentiment, emotions, mental health, and personality. |
| Approach: | They propose to induce lexica using context-oblivious and context-aware approaches and compare their performance using crowd-worker assessment. |
| Outcome: | The proposed models can be induced using context-oblivious and context-aware approaches and evaluate their quality using crowd-worker assessment. |
Exploring hybrid approaches to readability: experiments on the complementarity between linguistic features and transformers (2024.findings-eacl)
Copied to clipboard
| Challenge: | Linguistic features have been a key component of the automatic assessment of text readability (ARA) with the development in the ARA field, the research moved to Deep Learning (DL) |
| Approach: | They compare 6 hybrid approaches to Machine Learning and DL on 4 corpora and found they are the most robust on smaller datasets and across languages. |
| Outcome: | The proposed approaches perform better on smaller datasets and across languages and tasks. |
Genre Matters: How Text Types Interact with Decoding Strategies and Lexical Predictors in Shaping Reading Behavior (2025.emnlp-main)
Copied to clipboard
| Challenge: | eMTeC is the first eye-tracking corpus of LLM-generated texts . it shows that text type strongly modulates cognitive effort during reading . |
| Approach: | They use the first eye-tracking corpus of LLM-generated texts to study eye movements during reading and how decoding strategies interact with text types to shape reading behavior. |
| Outcome: | The first eye-tracking corpus of LLM-generated texts shows that text type strongly modulates cognitive effort during reading and that word-level psycholinguistic effects vary systematically across genres. |
Investigating Ensemble Methods for Model Robustness Improvement of Text Classifiers (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to reduce model's reliance on bias features ignore the learnability of these features. |
| Approach: | They propose to reduce models' reliance on bias features by first training models with fixed low-capacity models which ignore the learnability of the bias features. |
| Outcome: | The proposed models can perform better on out-of-distribution datasets than baseline models with a more sophisticated model design. |
TinyAttack: Exploring Stylistic Vulnerabilities in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing research on robustness of large language models has focused on text-based perturbations and the use of invisible characters and homoglyphs. |
| Approach: | They propose a framework to exploit weaknesses in Large Language Models (LLMs) by changing their stylistic structure using Unicode. |
| Outcome: | The proposed framework exploits vulnerabilities in large language models through Unicode-based stylistic transformations without altering its semantic or syntactic structure. |
When Format Changes Meaning: Investigating Semantic Inconsistency of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models are vulnerable to semantic inconsistency, a study finds . minor formatting variations result in divergent predictions for semantically equivalent inputs. |
| Approach: | They evaluate LLMs for semantic inconsistency and find they remain vulnerable . they propose to use mechanistic analysis to develop models that improve their reliability . |
| Outcome: | The proposed model is vulnerable to semantic inconsistency, the authors show . their model is brittle even in state-of-the-art models, they say . |
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks (2025.coling-main)
Copied to clipboard
| Challenge: | Recent advances in natural language processing have highlighted the vulnerability of deep learning models to adversarial attacks. |
| Approach: | They propose a benchmark for textual adversarial defence that evaluates state-of-the-art defence mechanisms across diverse datasets, models, and tasks. |
| Outcome: | The proposed benchmark incorporates a wide range of datasets and evaluates state-of-the-art defence mechanisms. |