Papers by Aaron Courville

8 papers
Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic Oracle (2021.naacl-main)

Copied to clipboard

Challenge: Failing to capture the structure of input language could lead to generalization problems and over-parametrization.
Approach: They propose a new syntax-aware language model that explicitly models the structure with an incremental parser and maintains the conditional probability setting of a standard language model.
Outcome: The proposed model can achieve strong results in language modeling, parsing, and syntactic generalization tests while using fewer parameters than other models.
Supervised Seeded Iterated Learning for Interactive Language Learning (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work has focused on word-based conversational agents that tend to invent their language rather than leveraging natural language.
Approach: They propose two methods to counter language drift by combining S2P and Seeded Iterated Learning to minimize their weaknesses.
Outcome: The proposed methods reduce late-stage training collapses and higher negative likelihood when evaluated on human corpus.
Recursive Top-Down Production for Sentence Generation with Latent Trees (2020.findings-emnlp)

Copied to clipboard

Challenge: Various studies have shown that incorporating syntactic structures into recursive encoders can be beneficial for various natural language tasks.
Approach: They propose a dynamic programming algorithm that marginalises over latent binary tree structures with N leaves to train a recursive neural function.
Outcome: The proposed model outperforms previous models on the LENGTH split and English question formation tasks on the Multi30k dataset.
StructFormer: Joint Unsupervised Induction of Dependency and Constituency Structure from Masked Language Modeling (2021.acl-long)

Copied to clipboard

Challenge: Existing models that induce grammar structures from data focus on constituency or dependency structures alone.
Approach: They propose a model that can induce dependency and constituency structure at the same time.
Outcome: The proposed model can induce both constituency and dependency structures at the same time.
Sparse Universal Transformer (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models that use VTs as their backbone model are based on UTs that share parameters across layers and have better compositional generalization.
Approach: They propose to use Sparse Mixture of Experts to reduce UT's computation complexity while retaining its parameter efficiency and generalization ability.
Outcome: The proposed model achieves strong generalization results on formal language tasks and impressive parameter and computation efficiency on standard natural language benchmarks.
Straight to the Tree: Constituency Parsing with Neural Syntactic Distance (P18-1)

Copied to clipboard

Challenge: Compared to traditional shift-reduce parsing schemes, our approach is free from the potentially disastrous compounding error.
Approach: They propose a model that predicts a scalar for each split position in a sentence and then determines the topology of grammar tree based on syntactic distances.
Outcome: The proposed model achieves the state-of-the-art single model F1 score of 92.1 on PTB and 86.4 on CTB dataset, surpassing the previous single model results by a large margin.
Unsupervised Dependency Graph Network (2022.acl-long)

Copied to clipboard

Challenge: Recent work has identified properties of pretrained self-attention models that mirror those of dependency parse structures.
Approach: They propose a model that encourages attention heads to model different dependency relations from raw corpora and a masked language modeling task.
Outcome: The proposed model can induce dependency structures from raw corpora and the masked language modeling task without gold POS tags and any external information.
Understanding by Understanding Not: Modeling Negation in Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Negation is a core construction in natural language, but state-of-the-art pre-trained language models often handle it incorrectly.
Approach: They propose to augment language modeling objective with unlikelihood objective based on negated generic sentences from a raw text corpus.
Outcome: The proposed approach reduces the top 1 error rate to 4% on negated LAMA dataset and improves on negating NLI benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations