Papers by David Mareček

6 papers
Universal Dependencies According to BERT: Both More Specific and More General (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that individual BERT heads encode particular dependency relation types, but they do not match one-to-one.
Approach: They propose a method for relation identification and syntactic tree construction that can be applied with minimal supervision and generalizes well across languages.
Outcome: The proposed method produces significantly more consistent dependency trees than previous work and can be applied with only a minimal amount of supervision and generalizes well across languages.
Introducing Orthogonal Constraint in Structural Probes (2021.acl-long)

Copied to clipboard

Challenge: Recent studies have focused on interpreting pre-trained models' representations and analyzing their structures.
Approach: They propose a new type of structural probing where a linear projection is decomposed into two types.
Outcome: The proposed method is tested on two novel tasks and shows that lexical and syntactic information is separated in the representations.
Exploring Interpretability of Independent Components of Word Embeddings with Automated Word Intruder Test (2024.lrec-main)

Copied to clipboard

Challenge: Independent Component Analysis (ICA) is an algorithm for finding separate sources in a mixed signal.
Approach: They propose to use ICA to analyze word embeddings to quantify interpretability . they propose to automate word intruder test to quantify the components .
Outcome: The proposed algorithm can be used to find semantic features of words . it can be combined to find words that have features associated with the components .
Tokenization Impacts Multilingual Language Modeling: Assessing Vocabulary Allocation and Overlap Across Languages (2023.findings-acl)

Copied to clipboard

Challenge: Multilingual language models perform surprisingly well in a variety of NLP tasks for diverse languages.
Approach: They propose to evaluate the quality of lexical representation and vocabulary overlap observed in sub-word tokenizers.
Outcome: The proposed criteria show that the overlap of vocabulary across languages can be detrimental to certain downstream tasks.
The Functional Relevance of Probed Information: A Case Study (2023.eacl-main)

Copied to clipboard

Challenge: Recent studies have shown that transformer models like BERT rely on number information encoded in their representations of sentences’ subjects and head verbs when performing subject-verb agreement.
Approach: They propose to use probing to find out which words contain functionally relevant information encoded in the representations of subject plurality and words that agree with it in number in BERT.
Outcome: The proposed model only uses the subject plurality information encoded in its representations of the subject and words that agree with it in number.
Examining Cross-lingual Contextual Embeddings with Orthogonal Structural Probes (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies on whether multilingual embeddings can be aligned in a shared space across languages are lacking.
Approach: They propose to learn a projection based on monolingual annotated datasets and evaluate syntactic and lexical information encoded in a shared cross-lingual embedding space.
Outcome: The proposed model can be used to learn representations for languages with low resources.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations