Papers by Gábor Bella
Language Diversity: Visible to Humans, Exploitable by Machines (2022.acl-demo)
Copied to clipboard
Gábor Bella, Erdenebileg Byambadorj, Yamini Chandrashekar, Khuyagbaatar Batsuren, Danish Cheema, Fausto Giunchiglia
| Challenge: | Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over two thousand languages. |
| Approach: | Universal Knowledge Core is a large multilingual lexical database with a focus on language diversity and covering over two thousand languages. |
| Outcome: | the database lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings and lexicon similarities. |
Using Linguistic Typology to Enrich Multilingual Lexicons: the Case of Lexical Gaps in Kinship (2022.lrec-1)
Copied to clipboard
Temuulen Khishigsuren, Gábor Bella, Khuyagbaatar Batsuren, Abed Alhakim Freihat, Nandu Chandran Nair, Amarsanaa Ganbold, Hadi Khalilia, Yamini Chandrashekar, Fausto Giunchiglia
| Challenge: | a method to enrich lexical resources with content relating to linguistic diversity is proposed . Typology-based approaches are being used to improve cross-lingual NLP tasks . |
| Approach: | They propose a method to enrich lexical resources with content relating to linguistic diversity based on lexica. |
| Outcome: | The proposed method can be used to improve cross-lingual NLP tasks by removing the need for parallel textual corpora or cross-linguistic transfer from high-to-low-resourced languages. |
Using Crowd Agreement for Wordnet Localization (L18-1)
Copied to clipboard
| Challenge: | Lexical-semantic resources like WordNet are a fundamental resource for many NLP and semantic applications. |
| Approach: | They propose a crowdsourcing workflow that consists of synset localization and validation . they use inter-rater agreement metrics to estimate the precision of the results . |
| Outcome: | The proposed method is cost-effective and provides a good trade-off between quality and speed of progress. |
Exploring the Language of Data (2020.coling-main)
Copied to clipboard
| Challenge: | Structured data, such as database tables or XML trees, often contain short natural language labels that describe the data structure itself or provide content (attribute values). Conventional NLP tools, such supervised sequence labellers or embeddings trained on full sentences, do not perform well on structured data. |
| Approach: | They propose to design a type of abbreviated grammar that is called the Language of Data and to investigate the grammatical properties of such labels. |
| Outcome: | The proposed model outperforms models trained on standard text on tokenisation, part-of-speech tagging, and named entity recognition over real-world structured data. |
UniMorph 4.0: Universal Morphology (2022.lrec-1)
Copied to clipboard
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
| Challenge: | The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages. |
| Approach: | They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. |
| Outcome: | The proposed schema has added 66 new languages, including 24 endangered languages. |
A Major Wordnet for a Minority Language: Scottish Gaelic (2020.lrec-1)
Copied to clipboard
Gábor Bella, Fiona McNeill, Rody Gorman, Caoimhin O Donnaile, Kirsty MacDonald, Yamini Chandrashekar, Abed Alhakim Freihat, Fausto Giunchiglia
| Challenge: | a new wordnet resource is available for Scottish Gaelic, a minority language spoken by 60,000 speakers . weak online presence of minority languages is a problem due to lack of digital corpora, authors say . |
| Approach: | They propose a new wordnet resource for Scottish Gaelic, a Celtic minority language . the wordnet contains over 15 thousand word senses and is among the 30 largest in the world . authors hope to contribute to long-term preservation of Scottish Gaels as a living language - offline and on the Web . |
| Outcome: | The new wordnet is for Scottish Gaelic, a minority language spoken by 60,000 speakers . the wordnet contains over 15 thousand word senses and is among the 30 largest in the world . authors hope it will contribute to the long-term preservation of the language, both offline and on the Web . |
IndoUKC: A Concept-Centered Indian Multilingual Lexical Resource (2022.lrec-1)
Copied to clipboard
| Challenge: | a new multilingual lexical database for Indian languages is proposed . the database provides words and crosslingually mapped word meanings specific to Indian languages and cultures. |
| Approach: | They propose to create a multilingual lexical database for Indian languages called IndoUKC . the database is based on existing IndoWordNet resources and is available for browsing . |
| Outcome: | The proposed database is based on the existing IndoWordNet resource and is available for download through the LiveLanguage data catalogue. |