Challenge: In this paper, we test the hypothesis that deeper transformers generalize more compositionally.
Approach: They propose to add layers to transformers to generalize more compositionally . they propose to fine-tune the models so that the total number of parameters is constant .
Outcome: The proposed model generalizes more compositionally than shallower models, but returns diminish . the proposed model can be made shallower without sacrificing performance .

Similar Papers

Analyzing the Inner Workings of Transformers in Compositional Generalization (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on compositional generalization abilities of neural models have focused on benchmarks, but the results do not reflect the underlying competence of the model.
Approach: They propose to find an existing subnetwork that contributes to the generalization performance and perform causal analyses on how the model utilizes syntactic features.
Outcome: The proposed model relies on syntactic features but the subnetwork with better generalization performance relies mainly on a non-compositional algorithm .
When Can Transformers Ground and Compose: Insights from Compositional Generalization Benchmarks (2022.emnlp-main)

Copied to clipboard

Challenge: Recent benchmarks like ReaSCAN use navigation tasks grounded in a grid world to assess whether neural models exhibit compositional behaviour.
Approach: They propose a transformer-based model that outperforms specialized architectures on ReaSCAN and a modified version of gSCAN to test their performance.
Outcome: The proposed model outperforms specialized architectures on ReaSCAN and gSCAN on a grid world and can generalize to deeper input structures.
Making Transformers Solve Compositional Tasks (2022.acl-long)

Copied to clipboard

Challenge: Several studies have reported the inability of Transformer models to generalize compositionally . a key aspect of natural language is the ability to learn basic primitives .
Approach: They propose to use Transformers to generalize compositionally in a large range of tasks . they find that Transformers generalize significantly better than previous models .
Outcome: The proposed models generalize compositionally significantly better than previous models . a set of 12 datasets shows that the proposed models can be improved .
Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models struggle on out-of-distribution compositional generalization . recent work shows considerable improvements on many NLP tasks from model scaling .
Approach: They evaluate encoder-decoder models up to 11B parameters and decoder-only models up 540B parameters . they compare scaling curves for fine-tuning, prompt tuning, and in-context learning methods .
Outcome: The proposed scaling methods improve compositional generalization on many tasks . fine-tuning generally has flat or negative scaling curves on out-of-distribution compositional . larger models are better at modeling the syntax of the output space, the study finds .
Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers (2025.tacl-1)

Copied to clipboard

Challenge: Inductive biases in transformers can cause hierarchical generalization without explicitly encoding structural bias.
Approach: They investigate sources of inductive bias in transformer models and their training that could cause such preference for hierarchical generalization.
Outcome: The proposed model can generalize to novel syntactic forms without explicit bias . the proposed model is able to generalize on a dataset with a hierarchical grammar .
Data Factors for Better Compositional Generalization (2023.emnlp-main)

Copied to clipboard

Challenge: Recent diagnostic datasets on compositional generalization expose severe problems . state-of-the-art models trained on larger and more general datasets show better generalization ability .
Approach: They conduct an empirical analysis by training Transformer models on a variety of training sets with different data factors including dataset scale, pattern complexity, example difficulty, etc.
Outcome: The proposed model training on larger datasets improves on compositional generalization tasks.
Grokking of Hierarchical Structure in Vanilla Transformers (2023.acl-short)

Copied to clipboard

Challenge: a recent study has shown that neural sequence models like transformers can generalize hierarchically when training for extended periods.
Approach: They show that transformers can learn to generalize hierarchically after long training periods . they call this phenomenon structural grokking, which exhibits inverted U-shaped scaling in model depth .
Outcome: The proposed model generalizes better than both very deep and very shallow models on multiple datasets.
Exploring Compositional Generalization of Large Language Models (2024.naacl-srw)

Copied to clipboard

Challenge: a recent study has found that large language models can generalize compositional instructions from simple instructions to complex ones.
Approach: They study the generalization ability of large language models with respect to compositional instructions . they first construct a dataset with the help of ChatGPT guided by the self-instruct technique .
Outcome: The proposed model can generalize from simple instructions to more intricate ones, the authors show . their results show that training LLMs on higher-order compositional instructions improves performance on lower-order ones, but not on higher order ones.
How Abstract Is Linguistic Generalization in Large Language Models? Experiments with Argument Structure (2023.tacl-1)

Copied to clipboard

Challenge: Competent speakers of a language know how likely a word w is to appear in a specific context .
Approach: They use transformer-based large language models to generalize a novel noun argument . they show a bias to generalise based on linear order, instead of a linear order .
Outcome: The proposed models perform well in generalizing the distribution of a novel noun argument between related contexts that were seen during pre-training.
Optimizing Deeper Transformers on Small Datasets (2021.acl-long)

Copied to clipboard

Challenge: a common belief that training deep transformers from scratch requires large datasets is wrong . however, with proper initialization and optimization, the benefits of very deep transformer can carry over to challenging tasks with small datasets.
Approach: They train 48 layers of transformers from pre-trained RoBERTa and 24 relation-aware layers from scratch.
Outcome: The proposed scheme achieves state-of-the-art performance on a text-to-sql parsing benchmark . it uses 24 fine-tuned layers from pre-trained RoBERTa and 24 relation-aware layers from scratch .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations