| Challenge: | Existing approaches to grammar induction have resorted to manually-engineered features and auxiliary objectives to induce the desired structures. |
| Approach: | They propose a formalization of the grammar induction problem that models sentences as being generated by a compound probabilistic context free grammar. |
| Outcome: | Experiments on English and Chinese show that the proposed approach is more efficient than other methods. |
Similar Papers
Categorial grammar induction from raw data (2023.findings-acl)
Copied to clipboard
| Challenge: | a new model for categorial grammar induction is based on raw data without part-of-speech information. |
| Approach: | They propose a grammar induction model that learns from raw data without part-of-speech information. |
| Outcome: | a new model for inducing a basic categorial grammar is developed . the model attains a recall-homogeneity of 0.33 on average, and a bias toward forward function application is added . |
PCFGs Can Do Better: Inducing Probabilistic Context-Free Grammars with Many Symbols (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent work shows that probabilistic context-free grammars with neural parameterization can be effective in unsupervised constituency parsing. |
| Approach: | They propose a parameterization form of PCFGs based on tensor decomposition which has at most quadratic computational complexity in the symbol number. |
| Outcome: | The proposed model improves unsupervised constituency parsing performance across ten languages. |
The Return of Lexical Dependencies: Neural Lexicalized PCFGs (2020.tacl-1)
Copied to clipboard
| Challenge: | Existing approaches to grammar induction focus on discovering constituents or dependencies. |
| Approach: | They propose to model lexical dependencies using context free grammars instead of lexicals . they show that this unified framework induces both constituents and dependencies . |
| Outcome: | The proposed model overcomes sparsity problems and induces constituents and dependencies better than the current methods. |
Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing models face expressiveness bottlenecks, resulting in unnecessarily large yet underperforming grammars. |
| Approach: | They propose a method to reduce the expressiveness bottleneck of unsupervised neural grammar induction by leveraging neural parameterization to estimate prob-ability distributions. |
| Outcome: | The proposed approach significantly improves parsing performance while enabling the use of significantly more compact grammars across a wide range of languages. |
Unsupervised Learning of PCFGs with Normalizing Flow (P19-1)
Copied to clipboard
| Challenge: | Existing induction models unable to incorporate semantics and morphology into induction . current models lack a robust model for generating morphologically rich sentences . |
| Approach: | They propose a PCFG inducer which uses context embeddings to generalize over rare, morphologically rich forms. |
| Outcome: | The proposed model produces grammars with state-of-the-art accuracy on a variety of languages. |
Consistent Unsupervised Estimators for Anchored PCFGs (2020.tacl-1)
Copied to clipboard
| Challenge: | a novel approach for learning probabilistic context-free grammars from strings is proposed . strong learning means that there can be many structurally different PCFGs that define the same distribution over strings. |
| Approach: | They propose an algorithm that is a consistent estimator for a class of PCFGs that are anchored . they show that if the grammar is anchored, the parameters can be directly related to distributional properties of the anchoring strings. |
| Outcome: | The proposed algorithm is consistent for a large class of probabilistic context-free grammars . it shows that the proposed algorithm has good finite sample behavior . |
Character-based PCFG Induction for Modeling the Syntactic Acquisition of Morphologically Rich Languages (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for syntactic acquisition are word-based and do not inspect functional affixes. |
| Approach: | They propose a computer-based induction model that allows a clean ablation of the influence of subword information in grammar induction. |
| Outcome: | The proposed model is more accurate in morphologically richer languages with subword information than word-based models. |
Categorial Grammar Induction with Stochastic Category Selection (2024.lrec-main)
Copied to clipboard
| Challenge: | categorial grammar inducers have been used to learn from raw data, but they use shortcuts to ensure branching behavior. |
| Approach: | They propose a grammar inducer that learns from raw data and does not rely on bias terms . they show a recall-homogeneity of 0.48 on a corpus of English child-directed speech . |
| Outcome: | The proposed model achieves a recall-homogeneity of 0.48 on a corpus of English child-directed speech . |
Grammar Induction with Neural Language Models: An Unusual Replication (D18-1)
Copied to clipboard
| Challenge: | Recent work on latent tree learning attempts to develop models with parse-valued latent variables and train them on non-parsing tasks. |
| Approach: | They propose a model with parse-valued latent variables and a strong latent tree learning result on constituency parsing. |
| Outcome: | The proposed model outperforms all baselines and performs competitively with symbolic grammar induction systems. |
Video-aided Unsupervised Grammar Induction (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods of multi-modal grammar induction focus on grammar inducing from text-image pairs, but videos provide even richer information, such as static objects and actions. |
| Approach: | They propose a video-aided grammar induction model which learns a constituency parser from unlabeled text and its corresponding video. |
| Outcome: | The proposed model outperforms existing systems on three benchmarks. |