Papers by Damian Blasi

6 papers
On the Distribution of Deep Clausal Embeddings: A Large Cross-linguistic Study (P19-1)

Copied to clipboard

Challenge: Empirical evidence on the prevalence and limits of embeddings has been based on either laboratory setups or corpus data of relatively limited size.
Approach: They use large, dependency-parsed corpora to capture clausal embedding through dependency graphs and assess their distribution.
Outcome: The results show that there is no evidence for hard constraints on embedding depth . they also show that sentences with many embeddable clauses do not display a bias towards less deep embedded sentences.
Meaning to Form: Measuring Systematicity as Information (P19-1)

Copied to clipboard

Challenge: A longstanding debate in semiotics centers on the relationship between linguistic signs and their corresponding semantics: is there an arbitrary relationship between word forms and their meaning, or does some systematic phenomenon pervade?
Approach: They propose to quantify the systematicity of the sign using mutual information and recurrent neural networks to examine 106 languages.
Outcome: The proposed model reduces entropy in a word form conditioned on its semantic representation and recovers English examples of systematic affixes.
Systematic Inequalities in Language Technology Performance across the World’s Languages (2022.acl-long)

Copied to clipboard

Challenge: Recent studies have revealed that NLP is limited to a subset of the world’s 6,500 languages.
Approach: They propose a framework for estimating the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Outcome: The proposed framework estimates the global utility of language technologies as revealed in a comprehensive snapshot of recent publications in NLP.
Is Word Segmentation Child’s Play in All Languages? (P19-1)

Copied to clipboard

Challenge: Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development .
Approach: They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants.
Outcome: The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid .
Speakers Fill Lexical Semantic Gaps with Context (2020.emnlp-main)

Copied to clipboard

Challenge: Lexical ambiguity is widespread in language, allowing for the reuse of economical word forms and thus making language more efficient.
Approach: They propose two ways to estimate lexical ambiguity as the entropy of meanings a word can take . they validate this hypothesis by using WordNet and BERT .
Outcome: The proposed method shows that on six high-resource languages, there are significant correlations between the estimate and the number of synonyms a word has in WordNet.
Quantifying the Semantic Core of Gender Systems (D19-1)

Copied to clipboard

Challenge: a large number of languages employ grammatical gender on the lexeme, but is it truly arbitrary? a recent study shows that the relationship between grammamatical gender and lexical semantics is opaque.
Approach: They propose a method to correlating inanimate nouns' gender with lexical semantics . they find that the gender systems of 18 languages exhibit a significant correlation with a definition .
Outcome: a new study shows that the gender assignments of 18 languages are arbitrary . the authors show that the correlation between gender and semantics is significant .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations