Papers by Lawrence Wolf-Sonkin

9 papers
On the Distribution of Deep Clausal Embeddings: A Large Cross-linguistic Study (P19-1)

Copied to clipboard

Challenge: Empirical evidence on the prevalence and limits of embeddings has been based on either laboratory setups or corpus data of relatively limited size.
Approach: They use large, dependency-parsed corpora to capture clausal embedding through dependency graphs and assess their distribution.
Outcome: The results show that there is no evidence for hard constraints on embedding depth . they also show that sentences with many embeddable clauses do not display a bias towards less deep embedded sentences.
On the Relationships Between the Grammatical Genders of Inanimate Nouns and Their Co-Occurring Adjectives and Verbs (2021.tacl-1)

Copied to clipboard

Challenge: In many languages, nouns possess grammatical genders.
Approach: They use large-scale corpora and tools from NLP and information theory to test whether there is a relationship between grammatical genders of inanimate nouns and adjectives used to describe them.
Outcome: The results show that there is a statistically significant relationship between the grammatical genders of inanimate nouns and adjectives used to describe them in all six languages.
A Structured Variational Autoencoder for Contextual Morphological Inflection (P18-1)

Copied to clipboard

Challenge: morphological inflectors typically trained on fully supervised, type-level data, but how can we improve their performance? et al., 2016: a novel latent-variable model for semi-supervised learning of inflection generation.
Approach: They propose a latent-variable model for semi-supervised learning of inflection generation . they use a wake-sleep algorithm to enable posterior inference over latent variables .
Outcome: The proposed model improves on 23 languages and shows 10% accuracy improvement . the proposed model is based on the wake-sleep algorithm .
Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset (2020.lrec-1)

Copied to clipboard

Challenge: a new resource is available for 12 South Asian languages that use the Latin script for text entry . the Latin-script system is not widely used in South Asian language writing, despite the Latin alphabet .
Approach: They describe the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages.
Outcome: The Dakshina dataset includes text in both the Latin and native scripts for 12 languages . the authors provide baseline results on several tasks made possible by the dataset .
Extensions to Brahmic script processing within the Nisaba library: new scripts, languages and utilities (2022.lrec-1)

Copied to clipboard

Challenge: a brahmic script is used to record endangered languages such as Dogri and Bengali for low-resource languages such that do not require visual normalization.
Approach: They propose to extend Brahmic script functionality within the Nisaba library of finite-state script normalization and processing utilities.
Outcome: The proposed extensions extend coverage from the original ten scripts to an additional ten of South Asia and beyond, including some used to record endangered languages such as Dogri.
Finite-state script normalization and processing utilities: The Nisaba Brahmic library (2021.eacl-demos)

Copied to clipboard

Challenge: a library for low-level processing of brahmic scripts is available for free.
Approach: They propose an open-source library for efficient low-level processing of ten major South Asian Brahmic scripts.
Outcome: The proposed library supports low-level processing of ten major south Asian Brahmic scripts.
Unsupervised Discovery of Gendered Language through Latent-Variable Modeling (P19-1)

Copied to clipboard

Challenge: a recent study has focused on the ways in which language is gendered . positive adjectives used to describe women are more often related to their bodies .
Approach: They propose a model that models adjective choice and its sentiment given the natural gender of a head noun.
Outcome: The proposed model shows that positive adjectives used to describe women are more often related to their bodies than positive adjective words used to explain men.
Combining Sentiment Lexica with a Multi-View Variational Autoencoder (N19-1)

Copied to clipboard

Challenge: a new model of sentiment lexica is being developed to combine disparate scales into a common representation.
Approach: They propose a model that unifies disparate scales into a common latent representation . they evaluate a text classification task using nine English-Language sentiment datasets .
Outcome: The proposed model outperforms six individual sentiment lexica and a simple combination thereof.
Quantifying the Semantic Core of Gender Systems (D19-1)

Copied to clipboard

Challenge: a large number of languages employ grammatical gender on the lexeme, but is it truly arbitrary? a recent study shows that the relationship between grammamatical gender and lexical semantics is opaque.
Approach: They propose a method to correlating inanimate nouns' gender with lexical semantics . they find that the gender systems of 18 languages exhibit a significant correlation with a definition .
Outcome: a new study shows that the gender assignments of 18 languages are arbitrary . the authors show that the correlation between gender and semantics is significant .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations