Attention Understands Semantic Relations (2022.lrec-1)

Copied to clipboard

Challenge: Present-day monopoly of foundation language models in most tasks forces researchers and practitioners to rely on popular large models without genuinely understanding the models' behaviour.
Approach: They propose a probing pipeline to study the representedness of semantic relations in transformer language models and propose 'attention mechanisms' that focus on syntactic relational information and semantic one.
Outcome: The proposed pipeline shows that attention scores are expressive as output activations on this task, despite their lesser ability to represent surface cues.

Similar Papers

Attention as Grounding: Exploring Textual and Cross-Modal Attention on Entities and Relations in Language-and-Vision Transformer (2022.findings-acl)

Copied to clipboard

Challenge: Existing work has focused on what is captured by multi-modal architectures.
Approach: They propose a multi-modal transformer that learns syntactic and semantic representations about entities and relations grounded in objects at the level of masked self-attention and cross-modal attention.
Outcome: The proposed model learns syntactic and semantic representations about objects and relations cross-modally and unimodally.
Attention Can Reflect Syntactic Structure (If You Let It) (2021.eacl-main)

Copied to clipboard

Challenge: a recent study has attempted to decode linguistic structure from the Transformer . but, much of the work focused on English, a language with rigid word order and a lack of inflectional morphology.
Approach: They propose to fine-tune a feature encoder for BERT to learn linguistic structure from its multi-head attention mechanism.
Outcome: The proposed model can decode full trees above baseline accuracy from single attention heads across languages.
Plausibility Processing in Transformer Language Models: Focusing on the Role of Attention Heads in GPT (2023.findings-emnlp)

Copied to clipboard

Challenge: Using attention heads, we can explore how Transformer language models process semantic knowledge, especially regarding the plausibility of noun-verb relations.
Approach: They propose to investigate how Transformer language models process semantic knowledge, especially regarding the plausibility of noun-verb relations.
Outcome: The proposed model exhibits a higher degree of similarity with humans in plausibility processing compared to other Transformer language models.
Probing for Bridging Inference in Transformer Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Pre-trained transformer language models are capable of bridging inference, but they lack the commonsense knowledge to capture syntactic information.
Approach: They investigate whether pre-trained transformer language models capture bridging inference . they use a masked token prediction task to investigate attention heads in BERT .
Outcome: The proposed model significantly captures bridging inference, the authors show . the distance between anaphor-antecedent and context plays an important role in the inference .
Roles and Utilization of Attention Heads in Transformer-based Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Sentence encoders based on transformer architectures have shown promising results on various natural language understanding tasks.
Approach: They propose a sentence representation method that takes advantage of most influential attention heads.
Outcome: The proposed method improves performance on the downstream tasks.
Broad-Coverage Semantic Parsing as Transduction (D19-1)

Copied to clipboard

Challenge: Existing approaches to broad-coverage semantic parsing are not applicable to all frameworks because of the lack of explicit alignments between tokens in the sentence and nodes in the semantic graph.
Approach: They propose a transduction parsing paradigm that unifies different broad-coverage semantic parsers into a paradigm that leverages multiple attention mechanisms to build meaning representation.
Outcome: The proposed approach improves state-of-the-art on AMR, SDP and UCCA and is competitive with the state- of-the art on SDP.
A Transformer with Stack Attention (2024.findings-naacl)

Copied to clipboard

Challenge: Recent research suggests that transformer-based language models fail to learn basic algorithmic patterns.
Approach: They propose to augment transformer-based language models with a differentiable stack-based attention mechanism that adds a level of interpretability to the model.
Outcome: The proposed model can model some, but not all, deterministic context-freelanguages.
VISIT: Visualizing and Interpreting the Semantic Information Flow of Transformers (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work in interpretability suggests we can project weights and hidden states of transformer-based language models (LMs) to their vocabulary space, a transformation that makes them more human interpretable.
Approach: They propose a tool to visualize a forward pass of Generative Pre-trained Transformers as an interactive flow graph with nodes representing neurons or hidden states and edges representing interactions between them.
Outcome: The proposed visualization simplifies huge amounts of data into easy-to-read graphs that can reflect the models’ internal processing, uncovering the contribution of each component to the models' final prediction.
Identifying Semantic Induction Heads to Understand In-Context Learning (2024.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated remarkable performance, but lack of transparency in their inference logic raises concerns about their trustworthiness.
Approach: They conduct a detailed analysis of the operations of attention heads to understand their in-context learning of LLMs.
Outcome: The proposed analysis of attention heads reveals that they increase the output logits of object tokens and recall objects . the proposed model is a novel approach to understand the in-context learning of large language models.
Chain and Causal Attention for Efficient Entity Tracking (2024.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to handle entity tracking require at least log2 (n+1) layers to handle n state changes.
Approach: They propose an efficient enhancement to the standard attention mechanism to handle long-term dependencies with a single layer.
Outcome: The proposed model can handle entity tracking with n state changes with a single layer.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations