Papers by Willem Zuidema
Feature Interactions Reveal Linguistic Structure in Language Models (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing features attribution methods for post-hoc interpretability ignore the existence of interactions between the effects of features on the prediction. |
| Approach: | They propose a grey box method to train models to perfection on a formal language classification task using PCFGs. |
| Outcome: | The proposed methods are able to uncover the grammatical rules acquired by the model under specific configurations and provide novel insights into the linguistic structure of the target models. |
Quantifying Context Mixing in Transformers (2023.eacl-main)
Copied to clipboard
| Challenge: | Self-attention weights and their transformed variants have been used for analyzing token-to-token interactions in Transformer-based models, but they are not faithful to the models’ decisions as they are only one part of an encoder block. |
| Approach: | They propose a new context mixing score customized for Transformers that provides us with a deeper understanding of how information is mixed at each encoder layer. |
| Outcome: | The proposed score outperforms other methods in linguistically informed rationales, probing, and faithfulness analysis. |
Transparency at the Source: Evaluating and Interpreting Language Models With Access to the True Distribution (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new approach to train, evaluate and interpret neural language models uses artificial, language-like data. |
| Approach: | They propose a setup for training, evaluating and interpreting neural language models that uses artificial, language-like data. |
| Outcome: | The proposed model is based on a massive probabilistic grammar and a large natural language corpus, and provides complete control over the generative process. |
DoLFIn: Distributions over Latent Features for Interpretability (2020.coling-main)
Copied to clipboard
| Challenge: | Existing approaches to interpret neural networks face a trade-off between a model's usefulness and its complexity. |
| Approach: | They propose a novel approach to achieve interpretability that avoids this trade-off by using probability as the central quantity instead of a fixed quantity. |
| Outcome: | The proposed approach outperforms the classical CNN and BiLSTM classifiers on the SST2 and AG-news datasets. |
Transformer-specific Interpretability (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | Transformers are dominant play-ers in various scientific fields, but their inner workings remain opaque. |
| Approach: | This tutorial presents a trending approach to interpreting Transformers . it uses specific features of the Transformer architecture to quantify context- mixing interactions . |
| Outcome: | This tutorial aims to show how a new trending approach can be applied to Transformer-based models. |
Homophone Disambiguation Reveals Patterns of Context Mixing in Speech Transformers (2023.emnlp-main)
Copied to clipboard
| Challenge: | 'context mixing' is a feature of Transformers that is used to build up representations of acoustic and linguistic structure in speech models. |
| Approach: | They propose to use a French spelling quirk to probe context mixing in speech models to find out how to translate spoken words into written equivalents. |
| Outcome: | The proposed model incorporates cues to identify correct transcription, whereas encoder-decoder models relegate task to decoder modules. |
Quantifying Attention Flow in Transformers (2020.acl-main)
Copied to clipboard
| Challenge: | In the Transformer model, “self-attention” combines information from attended embeddings into the representation of the focal embeddable in the next layer. |
| Approach: | They propose two methods to quantify flow of information through self-attention using attention weights as relative relevance of input tokens. |
| Outcome: | The proposed methods give complementary views on the flow of information and yield higher correlations with importance scores of input tokens. |
Do Language Models Exhibit Human-like Structural Priming Effects? (2024.findings-acl)
Copied to clipboard
| Challenge: | a recent exposure to a structure facilitates processing of the same structure, a study finds . structural priming is well attested in humans, for both language production and comprehension . |
| Approach: | They use the structural priming paradigm to investigate where priming effects manifest . they find that rarer elements within a prime increase priming effect . |
| Outcome: | The findings provide an important piece in the puzzle of understanding how properties within their context affect structural prediction in language models. |
Structural Persistence in Language Models: Priming as a Window into Abstract Language Representations (2022.tacl-1)
Copied to clipboard
| Challenge: | a rich literature has emerged in the last few years addressing these questions, including whether specific LMs have acquired specific linguistic constructions. |
| Approach: | They introduce a novel metric and release Prime-LM, a large corpus where they control for various linguistic factors that interact with priming strength. |
| Outcome: | The proposed model can learn abstract structural information independent of the structure of a sentence and is able to perform tasks that require natural language understanding skills. |
DecoderLens: Layerwise Interpretation of Encoder-Decoder Transformers (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing interpretability methods have been proposed to interpret the inner workings of Transformer models at different levels of precision and complexity. |
| Approach: | They propose a method to analyze encoder-decoder Transformers by using the decoder module Model Output encoder to cross-attend representations of intermediate encoder activations instead of using the default output. |
| Outcome: | The proposed method maps uninterpretable representations to human-interpreted sequences of words or symbols, shedding new light on the information flow in this popular but understudied class of models. |