Papers by Aryaman Arora

8 papers
Estimating the Entropy of Linguistic Distributions (2022.acl-short)

Copied to clipboard

Challenge: Shannon entropy is a quantity of interest for linguists studying the communicative capacity of human language.
Approach: They propose to use entropy estimators to estimate linguistic effects from observed data . they propose to recommend better entropic estimators for future linguistic studies .
Outcome: The proposed estimators are over-estimated due to poor entropy estimators, the authors argue . they recommend the same estimators be used in future linguistic studies.
pyvene: A Library for Understanding and Improving PyTorch Models via Interventions (2024.naacl-demo)

Copied to clipboard

Challenge: Existing libraries are often project-based, but pyvene provides a unified and extensible framework for performing interventions on neural models and sharing the intervened upon models with others.
Approach: They propose an open-source Python library that supports customizable interventions on a range of different PyTorch modules.
Outcome: The proposed framework provides a unified and extensible framework for performing interventions on neural models and sharing the intervened upon models with others.
Computational Historical Linguistics and Language Diversity in South Asia (2022.acl-long)

Copied to clipboard

Challenge: South Asian languages are underdocumented, and many with even official administrative status are lowresource (if not data-scarce) linguists must be willing to wrangle data from annotated corpora and grammatical descriptions for endangered languages.
Approach: They argue that data scatteredness is the primary obstacle in the development of South Asian language technology and propose new strategies to break the data barrier.
Outcome: The proposed approach is based on the findings of a recent study on comparative, contact, and historical linguistics in South Asia.
IruMozhi: Automatically classifying diglossia in Tamil (2024.findings-naacl)

Copied to clipboard

Challenge: Literary Tamil is highly diglossic, with two very different registers in everyday use . Spoken Tamil is under-studied in modern NLP systems compared to Literary Tamil written in the Tamil script .
Approach: They present a human-translated dataset of parallel text in Literary and Spoken Tamil.
Outcome: The proposed model trains classifiers on the task of identifying which Tamil variety a text belongs to.
MASALA: Modelling and Analysing the Semantics of Adpositions in Linguistic Annotation of Hindi (2022.lrec-1)

Copied to clipboard

Challenge: Existing work on SNACS annotation for a variety of typologically diverse languages focuses on semantic role labelling and upstream applications in related languages.
Approach: They propose to use the multilingual SNACS annotation scheme to attempt automatic labelling of SNAC supersenses in Hindi.
Outcome: The proposed method is competitive with previous work on English and Gujarati.
Universal Dependencies for Punjabi (2022.lrec-1)

Copied to clipboard

Challenge: UD is a community project that maintains a standard scheme for the annotation of grammar in a cross-lingually consistent manner.
Approach: They propose a Universal Dependencies treebank for Punjabi written in the Gurmukhi script and discuss corpus design and linguistic phenomena encountered in annotation.
Outcome: The proposed treebank covers a variety of genres and has been annotated for POS tags, dependency relations, and graph-based Enhanced Dependencies.
Supervised Grapheme-to-Phoneme Conversion of Orthographic Schwas in Hindi and Punjabi (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to predict schwa deletion in Hindi are based on prosodic or phonetic analysis.
Approach: They propose to use Hindi grapheme-to-phoneme (G2P) conversion to predict whether a schwa represented in the orthography is pronounced or unpronounced (deleted).
Outcome: The proposed model outperforms existing models on a newly-compiled pronunciation lexicon extracted from various online dictionaries.
UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations