Challenge: Existing work on hierarchical structure in neural networks has not captured human intuitions about hierarchic structures.
Approach: They propose to add an extra constraint to attention heads of the bidirectional Transformer encoder to encourage attention heads to follow tree structures.
Outcome: The proposed model improves language modeling and learning more explainable attention scores.

Similar Papers

You Only Need Attention to Traverse Trees (P19-1)

Copied to clipboard

Challenge: Recent research has focused on sentence representations.
Approach: They propose a tree-based model that captures phrase-level syntax and word-level dependencies by doing recursive traversal with attention.
Outcome: a new model captures phrase-level syntax and word-level dependencies with attention.
Syntax-Based Attention Masking for Neural Machine Translation (2021.naacl-srw)

Copied to clipboard

Challenge: Existing approaches to extend transformers to source-side trees are linearized into sequences, but they are limited by positional encodings.
Approach: They propose a method for extending transformers to source-side trees by using masks based on tree positions . they define a number of masks that limit self-attention based upon relationships among tree nodes .
Outcome: The proposed method improves on translations from English to germany and English to english and germany by +2.1 BLEU.
Recursive Tree-Structured Self-Attention for Answer Sentence Selection (2021.acl-long)

Copied to clipboard

Challenge: Recent top-performing models in Answer Sentence Selection use self-attention and transfer learning, but not syntactic structure.
Approach: They propose a recursive, tree-structured self-attention model that can represent all levels of syntactic parse trees with only one additional layer.
Outcome: The proposed model can represent all levels of syntactic parse trees with only one additional layer without transfer learning.
Do Syntax Trees Help Pre-trained Transformers Extract Information? (2021.eacl-main)

Copied to clipboard

Challenge: Recent work suggests that incorporating syntax information from dependency trees can improve task-specific transformer models.
Approach: They propose to incorporate dependency tree information into pre-trained transformers for three tasks . they propose a late fusion approach and a joint fusion technique to infuses syntax structure into attention layers.
Outcome: The proposed models obtain state-of-the-art results on SRL and relation extraction tasks.
Learning Sentence Representations over Tree Structures for Target-Dependent Classification (N18-1)

Copied to clipboard

Challenge: Existing work on tree structures uses syntactic parsers or Treebank annotations to perform target-dependent classifications.
Approach: They propose a reinforcement learning based approach which automatically induces target-specific sentence representations over tree structures.
Outcome: The proposed model gives superior performance on two benchmark tasks compared to previous work on parsed trees .
RealFormer: Transformer Likes Residual Attention (2021.findings-acl)

Copied to clipboard

Challenge: Existing techniques to create Residual Attention Layer Transformer networks outperform the canonical Transformer on a wide spectrum of tasks.
Approach: They propose a technique to create Residual Attention Layer Transformer networks that outperform the canonical Transformer on a wide spectrum of tasks.
Outcome: The proposed technique outperforms the canonical Transformer on a wide spectrum of tasks including Masked Language Modeling, GLUE, SQUAD, Neural Machine Translation, WikiHop, HotpotQA, Natural Questions, and OpenKP.
Code Summarization with Structure-induced Transformer (2021.findings-acl)

Copied to clipboard

Challenge: Code summarization (CS) is a promising area in recent language understanding . previous work using structurebased traversal or non-sequential models to learn structural program semantics has shown no performance gain .
Approach: They propose to use a structure-based traversal model to learn structural program semantics to generate human language automatically for programming language in the format of source code.
Outcome: Experiments show that the proposed method achieves state-of-the-art on benchmarks.
Going “Deeper”: Structured Sememe Prediction via Transformer with Tree Attention (2022.findings-acl)

Copied to clipboard

Challenge: Existing studies ignore hierarchical structures of sememes in sememe-based semantic description systems.
Approach: They propose a structured sememe prediction problem to predict a sememes tree with hierarchical structures rather than a set of sememas.
Outcome: The proposed model outperforms baseline models and shows its effectiveness . it predicts a sememe tree with hierarchical structures rather than a set of sememes .
Roles and Utilization of Attention Heads in Transformer-based Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Sentence encoders based on transformer architectures have shown promising results on various natural language understanding tasks.
Approach: They propose a sentence representation method that takes advantage of most influential attention heads.
Outcome: The proposed method improves performance on the downstream tasks.
Multiformer: A Head-Configurable Transformer-Based Model for Direct Speech Translation (2022.naacl-srw)

Copied to clipboard

Challenge: Existing approaches to address speech tasks with a self-attention mechanism are expensive and lead to information loss.
Approach: They propose a Transformer-based model which uses different attention mechanisms on each head to bias the self-attention towards the extraction of more diverse token interactions.
Outcome: The proposed model outperforms baseline models by 0.7 BLEU in the speech task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations