Challenge: Existing approaches to generalize compositionally are inadequate, but there is no evidence for this.
Approach: They propose a model-agnostic algorithm for subsampling instances with diverse structures from a labeled instance pool with structured outputs.
Outcome: The proposed algorithm leads to comparable or better generalization than prior algorithms in 9 out of 10 dataset-split type pairs.

Similar Papers

Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional Generalization (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research shows that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored.
Approach: They propose to generate synthetic utterance-program pairs for improving compositional generalization in semantic parsing by using structurally-diverse examples.
Outcome: The proposed approach leads to dramatic improvements in compositional generalization and moderate improvements in the traditional i.i.d setup.
Diverse Demonstrations Improve In-context Compositional Generalization (2023.acl-long)

Copied to clipboard

Challenge: In-context learning has shown great success in i.i.d semantic parsing splits . however, in compositional generalization, selecting similar demonstrations is insufficient .
Approach: They propose a method to select diverse demonstrations that collectively cover all the structures required in the output program and encourage the model to generalize to new structures from these demonstrations.
Outcome: The proposed method improves performance across three compositional generalization datasets and finetuning.
Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both? (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches to semantic parsing only evaluated on synthetic datasets that are not representative of natural language variation.
Approach: They propose a semantic parsing approach that handles both natural language variation and compositional generalization.
Outcome: The proposed model outperforms existing models across compositional generalization challenges on non-synthetic datasets while being competitive with the state-of-the-art on standard evaluations.
Improving Compositional Generalization in Classification Tasks via Structure Annotations (2021.acl-short)

Copied to clipboard

Challenge: Compositional generalization is the ability to generalize systematically to a new data distribution by combining known components.
Approach: They propose to convert a natural language sequence-to-sequence dataset into a classification dataset that requires compositional generalization.
Outcome: The proposed model can generalize compositionally by providing hints on the structure of the input.
SLOG: A Structural Generalization Benchmark for Semantic Parsing (2023.emnlp-main)

Copied to clipboard

Challenge: Existing compositional generalization benchmarks focus on lexical generalisation, the interpretation of novel lexicals in syntactic structures familiar from training.
Approach: They propose a semantic parsing dataset that extends COGS with 17 structural generalization cases to evaluate how well models generalize to new complex linguistic expressions.
Outcome: The proposed model generalization accuracy is far below the near-perfect accuracy of existing models on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models’ lexical and structural generalization capacities.
Unobserved Local Structures Make Compositional Generalization Hard (2022.emnlp-main)

Copied to clipboard

Challenge: Recent studies show sequence-to-sequence models struggle to generalize to new compositions . little is known on what makes generalization hard on a particular test instance .
Approach: They propose a criterion for the difficulty of an example that is hard if it contains a local structure that was not observed at training time.
Outcome: The proposed rule predicts instance-level generalization well across 5 different datasets.
Improving Compositional Generalization in Semantic Parsing (2020.findings-emnlp)

Copied to clipboard

Challenge: Generalization of models to out-of-distribution data has sparked substantial interest . compositional generalization is the ability to systematically generalize to test examples composed of components seen during training .
Approach: They propose to extend compositional generalization in semantic parsing by using contextual representations and training attention to agree with pre-computed token alignments.
Outcome: The proposed extensions improve compositional generalization on OOD compositions.
Mixture Content Selection for Diverse Sequence Generation (D19-1)

Copied to clipboard

Challenge: Generating diverse sequences exhibit semantically one-to-many relationships between source and target sequences.
Approach: They propose to separate diversification from generation using a general plug-and-play module that wraps around and guides an existing encoder-decoder model.
Outcome: The proposed method shows that diversification and generation are separate steps in the same model and that the model is robust.
Structural generalization in COGS: Supertagging is (almost) all you need (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that neural networks fail to generalize on out-of-distribution examples.
Approach: They extend a neural graph-based parsing framework to address compositional generalization limitations . they introduce a supertagging step with valency constraints and reduce the graph prediction problem .
Outcome: The proposed approach improves results on COGS datasets that require structural generalization.
Learning to Substitute Spans towards Improving Compositional Generalization (2023.acl-long)

Copied to clipboard

Challenge: despite the rising prevalence of neural sequence models, there is a deficiency in compositional generalization.
Approach: They propose a compositional augmentation strategy that enables multi-grained composition of substructures in the whole training set.
Outcome: The proposed strategy outperforms existing strategies on three compositional generalization benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations