Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey (2025.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks and models focus on systematicity of representations, but they focus on the systematicity in behaviour. |
| Approach: | They argue that systematicity is a desirable property in ML models as it enables strong generalization to novel contexts. |
| Outcome: | The proposed benchmarks and models focus on the systematicity of behaviour, while existing models focus primarily on language and vision. |
Similar Papers
Combine to Describe: Evaluating Compositional Generalization in Image Captioning (2022.acl-srw)
Copied to clipboard
| Challenge: | Recent work on compositionality has focused on the ability to combine simpler concepts to understand & generate arbitrarily more complex conceptual structures. |
| Approach: | They propose to use a set of image captioning models to benchmark their compositional generalization properties. |
| Outcome: | The proposed models do not generalize in terms of systematicity and productivity, but are robust to synonym substitutions. |
Toward Compositional Behavior in Neural Models: A Survey of Current Views (2024.emnlp-main)
Copied to clipboard
| Challenge: | Compositionality is a core property of natural language, and it is regarded as a key goal for modern NLP systems. |
| Approach: | They propose a conceptual framework to address compositionality in NLP . they propose to use this framework to survey researchers active in this area . |
| Outcome: | The proposed framework finds consensus on key points and suggests that scale alone is unlikely to achieve the desired behavior. |
Probing Linguistic Systematicity (2020.acl-main)
Copied to clipboard
| Challenge: | Existing evidence that deep natural language understanding models do not learn systematically is lacking. |
| Approach: | They examine whether deep natural language understanding models exhibit systematicity . they find that network architectures can generalize non-systematically . |
| Outcome: | The proposed model generalizes non-systematically, but is unsatisfactory, the authors argue . they show that the current state-of-the-art models do not generalize systematically . |
Systematicity, Compositionality and Transitivity of Deep NLP Models: a Metamorphic Testing Perspective (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing studies focus on robustness-like metamorphic relations, which limit the scope of linguistic properties they can test. |
| Approach: | They propose three new classes of metamorphic relations which address the properties of systematicity, compositionality and transitivity. |
| Outcome: | The proposed methods show that metamorphic models do not always behave according to expected linguistic properties. |
Systematic Generalization in Language Models Scales with Information Entropy (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for assessing compositional behavior are unclear on how to measure the difficulty of a systematic generalization problem. |
| Approach: | They propose a framework for measuring entropy in a sequence-to-sequence task and a method for measuring it. |
| Outcome: | The proposed framework scales with the entropy of the distribution of component parts in the training data. |
Defending Compositionality in Emergent Languages (2022.naacl-srw)
Copied to clipboard
| Challenge: | a recent paper has suggested that compositionality is a key factor in language productivity, but some research has questioned this. |
| Approach: | They argue that compositionality is essential for successful generalization . they run a two-agent communication game to test this hypothesis . |
| Outcome: | The proposed results show that ANNs can generalize well even without compositional behavior . authors argue that the results are incomplete and weak . |
The Learnability of Model-Theoretic Interpretation Functions in Artificial Neural Networks (2026.findings-acl)
Copied to clipboard
| Challenge: | Entity vectors improve scores on basic event, while gated architectures benefit most. |
| Approach: | They extend entity-level semantic representations, modern architectures, principled competing event generation, extended systematicity tests and a two-dimensional difficulty analysis disaggregating results by modifier complexity. |
| Outcome: | The proposed model-theoretic interpretation functions generalize systematically to out-of-training-sample sentences. |
Meaning to Form: Measuring Systematicity as Information (P19-1)
Copied to clipboard
| Challenge: | A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between word forms and their meaning, or does some systematic phenomenon pervade? |
| Approach: | They propose to quantify the systematicity of the sign using mutual information and recurrent neural networks to examine 106 languages. |
| Outcome: | The proposed model reduces entropy in a word form conditioned on its semantic representation and recovers English examples of systematic affixes. |
On Evaluating Multilingual Compositional Generalization with Translated Datasets (2023.acl-long)
Copied to clipboard
| Challenge: | a growing amount of research investigating compositional generalization in NLP is done on English . a critical semantic distortion is a limitation of the translation of datasets . |
| Approach: | They propose to translate a dataset for evaluating compositional generalization in semantic parsing. |
| Outcome: | The proposed benchmarks show that the translation of the MCWQ dataset suffers from semantic distortion. |
SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation (2024.acl-long)
Copied to clipboard
| Challenge: | Existing studies on compositional generalization in data-to-text generation focus on one manifestation, Systematicity, Productivity, Order invariance, and Rule learnability. |
| Approach: | They propose a method for evaluation of compositional generalization in data-to-text generation that includes four aspects of manifestations and allows high-quality evaluation without additional manual annotations. |
| Outcome: | The proposed method is based on two datasets and evaluates existing language models including LLMs. |