Challenge: lexically-independent subject-verb number agreement (NA) is performed by transformer-based neural language models (NLMs) . but when as little as one attractor is present, the model fails to perform lexical generalization .
Approach: They propose to disrupt lexical patterns found in naturally occurring stimuli for each targeted structure in a novel fine-grained analysis of BERT's behavior.
Outcome: The proposed model generalizes well for simple templates, but fails to perform lexically-independent generalization when as little as one attractor is present.

Similar Papers

Enhancing BERT for Lexical Normalization (D19-55)

Copied to clipboard

Challenge: Pre-trained contextual language models have improved performance of many NLP tasks.
Approach: They propose to use a pre-trained language model to perform lexical normalisation without UGC resources.
Outcome: The proposed model can perform lexical normalisation without the need for training sentences and 3,000 tokens.
A Primer in BERTology: What We Know About How BERT Works (2020.tacl-1)

Copied to clipboard

Challenge: a new study examines the current state of knowledge about the BERT model . the model is a stack of transformer encoder layers that are based on multiple self-attention ''heads''
Approach: They present a survey of over 150 studies of the popular Transformer-based model BERT . they discuss the current state of knowledge about how BERT works and how it is represented .
Outcome: The proposed model is based on the Transformer-based model with state-of-the-art results . the proposed model has little cognitive motivation and is too small to perform ablation studies .
Frequency Effects on Syntactic Rule Learning in Transformers (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models perform well on a variety of linguistic tasks that require symbolic reasoning, raising the question of whether such models implicitly represent abstract symbols and rules.
Approach: They investigate the performance of BERT on English subject–verb agreement by analyzing word frequency and absolute frequency of verb forms.
Outcome: The proposed model generalizes well to subject–verb pairs that never occurred in training, suggesting a degree of rule-governed behavior.
Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies have not investigated the relationship between a token's frequency in the training corpus and syntactic properties models learn about it.
Approach: They develop controlled experiments that probe models’ syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training.
Outcome: The proposed models can make syntactic generalizations for tokens seen as few as two times during training and transfer them to transformed contexts.
Are Transformers a Modern Version of ELIZA? Observations on French Object Verb Agreement (2021.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that unsupervised sentence representations of neural networks encode syntactic information by observing that neural language models are able to predict the agreement between a verb and its subject.
Approach: They propose to take an alternative look at these results by studying whether neural networks are able to build an abstract sentence representation rather than capture surface statistical regularities.
Outcome: The proposed model can achieve high accuracy on the long-range French object-verb agreement, indicating a possible flaw in the model's syntactic ability.
Do Neural Language Models Show Preferences for Syntactic Formalisms? (2020.acl-main)

Copied to clipboard

Challenge: Recent work on interpretability of deep neural language models concludes that many properties of natural language syntax are encoded in their representational spaces.
Approach: They propose to examine whether syntactic structure adheres to a surface-syntactical or deep syntaktic style of analysis.
Outcome: The proposed model prefers Universal Dependencies (UD) over Surface-Syntactic Universal Dependency (SUD) with interesting variations across languages and layers.
Syntactic Data Augmentation Increases Robustness to Inference Heuristics (2020.acl-main)

Copied to clipboard

Challenge: Pretrained neural models lack sensitivity to word order on controlled challenge sets . augmentation methods that improve accuracy on standard training sets may be a problem .
Approach: They propose to augment standard training sets with syntactically informative examples by applying syntastic transformations to sentences from the MNLI corpus.
Outcome: The proposed method improved BERT’s accuracy on controlled examples that diagnose sensitivity to word order from 0.28 to 0.73 without affecting performance on the MNLI test set.
Limited Generalizability in Argument Mining: State-Of-The-Art Models Learn Datasets, Not Arguments (2025.acl-long)

Copied to clipboard

Challenge: Identifying arguments is a prerequisite for various tasks in automated discourse analysis.
Approach: They evaluate four BERT-like transformers on 17 English sentence-level datasets . they find that they tend to rely on lexical shortcuts tied to content words .
Outcome: The proposed models perform best on 17 English sentence-level datasets on common tasks, but their performance drops when applied to unseen datasets.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (N19-1)

Copied to clipboard

Challenge: Existing language representation models pre-train deep bidirectional representations from unlabeled text without significant task-specific architecture modifications.
Approach: They propose a language representation model that pre-trains bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers.
Outcome: The proposed model achieves state-of-the-art results on eleven natural language processing tasks, pushing the GLUE score to 80.5 (7.7 point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement)
When classifying grammatical role, BERT doesn’t care about word order... except when it matters (2022.acl-short)

Copied to clipboard

Challenge: Recent work has shown large language models are surprisingly word order invariant . however, word order knowledge is crucial in defining later-layer representations of words .
Approach: They probe grammatical role representations in English BERT and GPT-2 to find word order crucial . they find word orders are crucial in defining later-layer representations of words in non-prototypical positions .
Outcome: The proposed model is based on natural prototypical inputs where word order is crucial for correct classification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations