Challenge: Language models learn rare syntactic phenomena by generalization vs. memorization, a study finds . aannalysis experiments show that humans learn rare grammatical structures by generalizing from less rare phenomena.
Approach: They iteratively trained transformer language models on a systematically manipulated corpus and evaluated their learning of a rare grammatical phenomenon.
Outcome: The results show that language models learn rare grammatical phenomena by generalization vs. memorization . human-scale corpora are used to train the models and compare their learning to counterfactual corpors .

Similar Papers

Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent evidence suggests that language models with human-scale pretraining data may possess a similar generalization ability by generalizing from frequent to rare constructions.
Approach: They construct a synthetic benchmark that targets syntactic and semantic properties of the English Let-Alone construction and compare it with a human-scale transformer language model.
Outcome: The proposed model can generalize from frequent to rare constructions, but human-scale models do not make correct generalizations about Let-Alone’s meaning.
Memorisation versus Generalisation in Pre-trained Language Models (2022.acl-long)

Copied to clipboard

Challenge: State-of-the-art pre-trained language models have been shown to memorise facts and perform well with limited amounts of training data.
Approach: They propose to extend pre-trained language models to generalise and memorise facts in noisy and low-resource scenarios.
Outcome: The proposed extension improves performance in low-resource named entity recognition tasks.
Do Neural Language Models Overcome Reporting Bias? (2020.coling-main)

Copied to clipboard

Challenge: Recent studies show that pre-trained language models can overcome reporting bias by estimating the plausibility of rare but unspoken facts.
Approach: They revisit the experiments conducted by Gordon and Van Durme (2013) . they find that pre-trained language models overestimate the very rare .
Outcome: The proposed approach overestimates the rare at the expense of the rare, while minimizing reporting bias.
Methods for Estimating and Improving Robustness of Language Models (2022.naacl-srw)

Copied to clipboard

Challenge: Large language models suffer from weak generalisation ability due to shallow textual relations over full semantic complexity of the problem.
Approach: They propose to incorporate some of these measures into training objectives to enhance distributional robustness of LLMs.
Outcome: The proposed models outperform human models on complex tasks and outperformed other models on deep networks.
How Much Do Language Models Copy From Their Training Data? Evaluating Linguistic Novelty in Text Generation Using RAVEN (2023.tacl-1)

Copied to clipboard

Challenge: Current language models generate high-quality text, but are they copying it or have they learned generalizable linguistic abstractions?
Approach: They propose a suite of analyses for assessing the novelty of generated text . they focus on sequential structure (n-grams) and syntactic structure (syntactical structure).
Outcome: The proposed model-generated text is as novel as the baseline human-generated model- generated text, but it is copied substantially, the authors show .
On the Ability and Limitations of Transformers to Recognize Formal Languages (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies on LSTMs have not revealed their ability to model syntactic properties.
Approach: They propose to build a Transformers model for a subclass of counter languages and find that their learning mechanism strongly correlates with their construction.
Outcome: The proposed model generalizes well on counter languages and its learned mechanism correlates with it.
Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies have not investigated the relationship between a token's frequency in the training corpus and syntactic properties models learn about it.
Approach: They develop controlled experiments that probe models’ syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training.
Outcome: The proposed models can make syntactic generalizations for tokens seen as few as two times during training and transfer them to transformed contexts.
Quantity doesn’t buy quality syntax with neural language models (D19-1)

Copied to clipboard

Challenge: Recurrent neural network language models can learn to predict upcoming words with remarkably low perplexity . but in syntactically complex contexts, they often assign unexpectedly high probabilities to ungrammatical words .
Approach: They investigate whether recurrent neural networks can learn to predict upcoming words with remarkably low perplexity.
Outcome: The proposed models perform worse than GPT and BERT in some constructions than LSTMs in other contexts.
Adaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve? (2024.emnlp-main)

Copied to clipboard

Challenge: In the last decade, the generalization and adaptation abilities of deep learning models were evaluated on fixed training and test distributions.
Approach: They propose to train large language models on unlabeled text corpora and train them online.
Outcome: The proposed model training on a text domain could degrade its perplexity on the test portion of the same domain.
Blackbird language matrices (BLM), a new task for rule-like generalization in neural networks: Can Large Language Models pass the test? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to evaluate large language models for generalization lack generalization ability . current methods for evaluating LLMs are based on tests of human intelligence .
Approach: They propose to use a language task to evaluate large language models' generalisation ability . they propose to ask LLMs to solve simple variants of the RAVEN IQ test .
Outcome: The proposed task can be used to evaluate the generalisation ability of large language models . it shows that current generative models can handle the task in the sense that they understand instructions .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations