Papers by Lawrence Wolf-Sonkin
On the Distribution of Deep Clausal Embeddings: A Large Cross-linguistic Study (P19-1)
Copied to clipboard
| Challenge: | Empirical evidence on the prevalence and limits of embeddings has been based on either laboratory setups or corpus data of relatively limited size. |
| Approach: | They use large, dependency-parsed corpora to capture clausal embedding through dependency graphs and assess their distribution. |
| Outcome: | The results show that there is no evidence for hard constraints on embedding depth . they also show that sentences with many embeddable clauses do not display a bias towards less deep embedded sentences. |
On the Relationships Between the Grammatical Genders of Inanimate Nouns and Their Co-Occurring Adjectives and Verbs (2021.tacl-1)
Copied to clipboard
| Challenge: | In many languages, nouns possess grammatical genders. |
| Approach: | They use large-scale corpora and tools from NLP and information theory to test whether there is a relationship between grammatical genders of inanimate nouns and adjectives used to describe them. |
| Outcome: | The results show that there is a statistically significant relationship between the grammatical genders of inanimate nouns and adjectives used to describe them in all six languages. |
A Structured Variational Autoencoder for Contextual Morphological Inflection (P18-1)
Copied to clipboard
| Challenge: | morphological inflectors typically trained on fully supervised, type-level data, but how can we improve their performance? et al., 2016: a novel latent-variable model for semi-supervised learning of inflection generation. |
| Approach: | They propose a latent-variable model for semi-supervised learning of inflection generation . they use a wake-sleep algorithm to enable posterior inference over latent variables . |
| Outcome: | The proposed model improves on 23 languages and shows 10% accuracy improvement . the proposed model is based on the wake-sleep algorithm . |
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)
Copied to clipboard
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, Keith Hall
| Challenge: | a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet . |
| Approach: | They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages. |
| Outcome: | The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset . |
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)
Copied to clipboard
| Challenge: | a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization. |
| Approach: | They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities. |
| Outcome: | The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri. |
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)
Copied to clipboard
| Challenge: | a library for low-level processing of brahmic scripts is available for free. |
| Approach: | They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts. |
| Outcome: | The proposed library supports low-level processing of ten major south Asian Brahmic scripts. |
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling (P19-1)
Copied to clipboard
| Challenge: | a recent study has focused on the ways in which language is gendered . positive adjectives used to describe women are more often related to their bodies . |
| Approach: | They propose a model that models adjective choice and its sentiment given the natural gender of a head noun. |
| Outcome: | The proposed model shows that positive adjectives used to describe women are more often related to their bodies than positive adjective words used to explain men. |
Combining Sentiment Lexica with a Multi-View Variational Autoencoder (N19-1)
Copied to clipboard
| Challenge: | a new model of sentiment lexica is being developed to combine disparate scales into a common representation. |
| Approach: | They propose a model that unifies disparate scales into a common latent representation . they evaluate a text classification task using nine English-Language sentiment datasets . |
| Outcome: | The proposed model outperforms six individual sentiment lexica and a simple combination thereof. |
Quantifying the Semantic Core of Gender Systems (D19-1)
Copied to clipboard
| Challenge: | a large number of languages employ grammatical gender on the lexeme, but is it truly arbitrary? a recent study shows that the relationship between grammamatical gender and lexical semantics is opaque. |
| Approach: | They propose a method to correlating inanimate nouns' gender with lexical semantics . they find that the gender systems of 18 languages exhibit a significant correlation with a definition . |
| Outcome: | a new study shows that the gender assignments of 18 languages are arbitrary . the authors show that the correlation between gender and semantics is significant . |