Papers by Damián Blasi
Modeling the Unigram Distribution (2021.findings-acl)
Copied to clipboard
| Challenge: | a novel model for estimating the unigram distribution is proposed for a language . the model is based on the word type and word type distributions, but does not consider contextual information. |
| Approach: | They propose a neuralization of Goldwater's (2011) model for estimating the unigram distribution in a language. |
| Outcome: | The proposed model outperforms character-level models in estimating the unigram distribution in a language. |
On the Interpretability and Significance of Bias Metrics in Texts: a PMI-based Approach (2023.acl-short)
Copied to clipboard
| Challenge: | Word embeddings have been used to quantify biases in texts for years, but their statistical properties and advantages have not been studied. |
| Approach: | They propose to use PMI-based metric to quantify bias in corpora by conditional probabilities and odds ratio to approximate it. |
| Outcome: | The proposed measure can be approximated by an odds ratio, which makes statistical inferences cost-effective and meaningful. |
Finding Concept-specific Biases in Form–Meaning Associations (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to detect cross-linguistic associations are not effective, but their effects are minor. |
| Approach: | They propose a method to measure cross-linguistic associations by controlling for the influence of language family and geographic proximity within a large concept-aligned, cross-lingual lexicon. |
| Outcome: | The proposed method shows that it is small, but it is unsurprisingly small (less than 0.5% on average). |
How (Non-)Optimal is the Lexicon? (2021.naacl-main)
Copied to clipboard
| Challenge: | lexical meanings are mapped to wordforms by usage pressures and constraints on sequences of symbols. |
| Approach: | They propose a coding-theoretic view of the lexicon and a novel generative statistical model to quantify its compressibility under various constraints. |
| Outcome: | The proposed model shows that (compositional) morphology and graphotactics can account for most of the complexity of natural codes—as measured by code length. |
On the Relationships Between the Grammatical Genders of Inanimate Nouns and Their Co-Occurring Adjectives and Verbs (2021.tacl-1)
Copied to clipboard
| Challenge: | In many languages, nouns possess grammatical genders. |
| Approach: | They use large-scale corpora and tools from NLP and information theory to test whether there is a relationship between grammatical genders of inanimate nouns and adjectives used to describe them. |
| Outcome: | The results show that there is a statistically significant relationship between the grammatical genders of inanimate nouns and adjectives used to describe them in all six languages. |
Evaluating Word Embeddings with Categorical Modularity (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing word embeddings use different bilingual supervision signals with varying levels of strength. |
| Approach: | They propose a graph modularity metric to measure word embedding quality . they use a set of 500 words belonging to 59 neurobiologically motivated semantic categories . |
| Outcome: | The proposed metric measures word embedding quality on monolingual and cross-lingual tasks. |
A surprisal–duration trade-off across and within the world’s languages (2021.emnlp-main)
Copied to clipboard
| Challenge: | Throughout human evolution, countless languages have evolved, each with unique features. |
| Approach: | They analysed a corpus of 600 languages to find strong evidence for a surprisal–duration trade-off between languages and languages. |
| Outcome: | The proposed model shows that phones are produced faster in languages where they are less surprising and vice versa. |