Challenge: Existing approaches to grammar induction have resorted to manually-engineered features and auxiliary objectives to induce the desired structures.
Approach: They propose a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar.
Outcome: Experiments on English and Chinese show that the proposed approach is more efficient than other methods.

Similar Papers

Categorial grammar induction from raw data (2023.findings-acl)

Copied to clipboard

Challenge: a new model for categorial grammar induction is based on raw data without part-of-speech information.
Approach: They propose a grammar induction model that learns from raw data without part-of-speech information.
Outcome: a new model for inducing a basic categorial grammar is developed . the model attains a recall-homogeneity of 0.33 on average, and a bias toward forward function application is added .
PCFGs Can Do Better: Inducing Probabilistic Context-Free Grammars with Many Symbols (2021.naacl-main)

Copied to clipboard

Challenge: Recent work shows that probabilistic context-free grammars with neural parameterization can be effective in unsupervised constituency parsing.
Approach: They propose a parameterization form of PCFGs based on tensor decomposition which has at most quadratic computational complexity in the symbol number.
Outcome: The proposed model improves unsupervised constituency parsing performance across ten languages.
The Return of Lexical Dependencies: Neural Lexicalized PCFGs (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to grammar induction focus on discovering constituents or dependencies.
Approach: They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies .
Outcome: The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods.
Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models face expressiveness bottlenecks, resulting in unnecessarily large yet underperforming grammars.
Approach: They propose a method to reduce the expressiveness bottleneck of unsupervised neural grammar induction by leveraging neural parameterization to estimate prob-ability distributions.
Outcome: The proposed approach significantly improves parsing performance while enabling the use of significantly more compact grammars across a wide range of languages.
Unsupervised Learning of PCFGs with Normalizing Flow (P19-1)

Copied to clipboard

Challenge: Existing induction models unable to incorporate semantics and morphology into induction . current models lack a robust model for generating morphologically rich sentences .
Approach: They propose a PCFG inducer which uses context embeddings to generalize over rare, morphologically rich forms.
Outcome: The proposed model produces grammars with state-of-the-art accuracy on a variety of languages.
Consistent Unsupervised Estimators for Anchored PCFGs (2020.tacl-1)

Copied to clipboard

Challenge: a novel approach for learning probabilistic context-free grammars from strings is proposed . strong learning means that there can be many structurally different PCFGs that define the same distribution over strings.
Approach: They propose an algorithm that is a consistent estimator for a class of PCFGs that are anchored . they show that if the grammar is anchored, the parameters can be directly related to distributional properties of the anchoring strings.
Outcome: The proposed algorithm is consistent for a large class of probabilistic context-free grammars . it shows that the proposed algorithm has good finite sample behavior .
Character-based PCFG Induction for Modeling the Syntactic Acquisition of Morphologically Rich Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for syntactic acquisition are word-based and do not inspect functional affixes.
Approach: They propose a computer-based induction model that allows a clean ablation of the influence of subword information in grammar induction.
Outcome: The proposed model is more accurate in morphologically richer languages with subword information than word-based models.
Categorial Grammar Induction with Stochastic Category Selection (2024.lrec-main)

Copied to clipboard

Challenge: categorial grammar inducers have been used to learn from raw data, but they use shortcuts to ensure branching behavior.
Approach: They propose a grammar inducer that learns from raw data and does not rely on bias terms . they show a recall-homogeneity of 0.48 on a corpus of English child-directed speech .
Outcome: The proposed model achieves a recall-homogeneity of 0.48 on a corpus of English child-directed speech .
Grammar Induction with Neural Language Models: An Unusual Replication (D18-1)

Copied to clipboard

Challenge: Recent work on latent tree learning attempts to develop models with parse-valued latent variables and train them on non-parsing tasks.
Approach: They propose a model with parse-valued latent variables and a strong latent tree learning result on constituency parsing.
Outcome: The proposed model outperforms all baselines and performs competitively with symbolic grammar induction systems.
Video-aided Unsupervised Grammar Induction (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods of multi-modal grammar induction focus on grammar inducing from text-image pairs, but videos provide even richer information, such as static objects and actions.
Approach: They propose a video-aided grammar induction model which learns a constituency parser from unlabeled text and its corresponding video.
Outcome: The proposed model outperforms existing systems on three benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations