Papers by Yamini Chandrashekar

4 papers
Language Diversity: Visible to Humans, Exploitable by Machines (2022.acl-demo)

Copied to clipboard

Challenge: Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over two thousand languages.
Approach: Universal Knowledge Core is a large multilingual lexical database with a focus on language diversity and covering over two thousand languages.
Outcome: the database lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings and lexicon similarities.
Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in Kinship (2022.lrec-1)

Copied to clipboard

Challenge: a method to enrich lexical resources with content relating to linguistic diversity is proposed . Typology-based approaches are being used to improve cross-lingual NLP tasks .
Approach: They propose a method to enrich lexical resources with content relating to linguistic diversity based on lexica.
Outcome: The proposed method can be used to improve cross-lingual NLP tasks by removing the need for parallel textual corpora or cross-linguistic transfer from high-to-low-resourced languages.
A Major Wordnet for a Minority Language: Scottish Gaelic (2020.lrec-1)

Copied to clipboard

Challenge: a new wordnet resource is available for Scottish Gaelic, a minority language spoken by 60,000 speakers . weak online presence of minority languages is a problem due to lack of digital corpora, authors say .
Approach: They propose a new wordnet resource for Scottish Gaelic, a Celtic minority language . the wordnet contains over 15 thousand word senses and is among the 30 largest in the world . authors hope to contribute to long-term preservation of Scottish Gaels as a living language - offline and on the Web .
Outcome: The new wordnet is for Scottish Gaelic, a minority language spoken by 60,000 speakers . the wordnet contains over 15 thousand word senses and is among the 30 largest in the world . authors hope it will contribute to the long-term preservation of the language, both offline and on the Web .
IndoUKC: A Concept-Centered Indian Multilingual Lexical Resource (2022.lrec-1)

Copied to clipboard

Challenge: a new multilingual lexical database for Indian languages is proposed . the database provides words and crosslingually mapped word meanings specific to Indian languages and cultures.
Approach: They propose to create a multilingual lexical database for Indian languages called IndoUKC . the database is based on existing IndoWordNet resources and is available for browsing .
Outcome: The proposed database is based on the existing IndoWordNet resource and is available for download through the LiveLanguage data catalogue.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations