Challenge: Existing benchmarks and models focus on systematicity of representations, but they focus on the systematicity in behaviour.
Approach: They argue that systematicity is a desirable property in ML models as it enables strong generalization to novel contexts.
Outcome: The proposed benchmarks and models focus on the systematicity of behaviour, while existing models focus primarily on language and vision.

Similar Papers

Combine to Describe: Evaluating Compositional Generalization in Image Captioning (2022.acl-srw)

Copied to clipboard

Challenge: Recent work on compositionality has focused on the ability to combine simpler concepts to understand & generate arbitrarily more complex conceptual structures.
Approach: They propose to use a set of image captioning models to benchmark their compositional generalization properties.
Outcome: The proposed models do not generalize in terms of systematicity and productivity, but are robust to synonym substitutions.
Toward Compositional Behavior in Neural Models: A Survey of Current Views (2024.emnlp-main)

Copied to clipboard

Challenge: Compositionality is a core property of natural language, and it is regarded as a key goal for modern NLP systems.
Approach: They propose a conceptual framework to address compositionality in NLP . they propose to use this framework to survey researchers active in this area .
Outcome: The proposed framework finds consensus on key points and suggests that scale alone is unlikely to achieve the desired behavior.
Probing Linguistic Systematicity (2020.acl-main)

Copied to clipboard

Challenge: Existing evidence that deep natural language understanding models do not learn systematically is lacking.
Approach: They examine whether deep natural language understanding models exhibit systematicity . they find that network architectures can generalize non-systematically .
Outcome: The proposed model generalizes non-systematically, but is unsatisfactory, the authors argue . they show that the current state-of-the-art models do not generalize systematically .
Systematicity, Compositionality and Transitivity of Deep NLP Models: a Metamorphic Testing Perspective (2022.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on robustness-like metamorphic relations, which limit the scope of linguistic properties they can test.
Approach: They propose three new classes of metamorphic relations which address the properties of systematicity, compositionality and transitivity.
Outcome: The proposed methods show that metamorphic models do not always behave according to expected linguistic properties.
Systematic Generalization in Language Models Scales with Information Entropy (2025.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for assessing compositional behavior are unclear on how to measure the difficulty of a systematic generalization problem.
Approach: They propose a framework for measuring entropy in a sequence-to-sequence task and a method for measuring it.
Outcome: The proposed framework scales with the entropy of the distribution of component parts in the training data.
Defending Compositionality in Emergent Languages (2022.naacl-srw)

Copied to clipboard

Challenge: a recent paper has suggested that compositionality is a key factor in language productivity, but some research has questioned this.
Approach: They argue that compositionality is essential for successful generalization . they run a two-agent communication game to test this hypothesis .
Outcome: The proposed results show that ANNs can generalize well even without compositional behavior . authors argue that the results are incomplete and weak .
The Learnability of Model-Theoretic Interpretation Functions in Artificial Neural Networks (2026.findings-acl)

Copied to clipboard

Challenge: Entity vectors improve scores on basic event, while gated architectures benefit most.
Approach: They extend entity-level semantic representations, modern architectures, principled competing event generation, extended systematicity tests and a two-dimensional difficulty analysis disaggregating results by modifier complexity.
Outcome: The proposed model-theoretic interpretation functions generalize systematically to out-of-training-sample sentences.
Meaning to Form: Measuring Systematicity as Information (P19-1)

Copied to clipboard

Challenge: A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between word forms and their meaning, or does some systematic phenomenon pervade?
Approach: They propose to quantify the systematicity of the sign using mutual information and recurrent neural networks to examine 106 languages.
Outcome: The proposed model reduces entropy in a word form conditioned on its semantic representation and recovers English examples of systematic affixes.
On Evaluating Multilingual Compositional Generalization with Translated Datasets (2023.acl-long)

Copied to clipboard

Challenge: a growing amount of research investigating compositional generalization in NLP is done on English . a critical semantic distortion is a limitation of the translation of datasets .
Approach: They propose to translate a dataset for evaluating compositional generalization in semantic parsing.
Outcome: The proposed benchmarks show that the translation of the MCWQ dataset suffers from semantic distortion.
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on compositional generalization in data-to-text generation focus on one manifestation, Systematicity, Productivity, Order invariance, and Rule learnability.
Approach: They propose a method for evaluation of compositional generalization in data-to-text generation that includes four aspects of manifestations and allows high-quality evaluation without additional manual annotations.
Outcome: The proposed method is based on two datasets and evaluates existing language models including LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations