Papers by Ximena Gutierrez-Vasques

4 papers
From characters to words: the turning point of BPE merges (2021.eacl-main)

Copied to clipboard

Challenge: morphological complexity is still a major challenge for NLP and the study of language.
Approach: They perform a cross-linguistic comparison following incremental merges of BPE for 47 diverse languages.
Outcome: The results show that language distributions are similar under specific levels of tokenization.
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)

Copied to clipboard

Challenge: a deep understanding of language is being achieved through increased access to data from minority and low-resource languages.
Approach: They present a diversity sample of text data for language comparison and multilingual natural language processing.
Outcome: The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures .
Challenges of language technologies for the indigenous languages of the Americas (C18-1)

Copied to clipboard

Challenge: Indigenous languages of the American continent are highly diverse, but have received little attention from the technological perspective.
Approach: They review the research, the digital resources and the available NLP systems for indigenous languages of the American continent . they stress the need of developing language resources and NLP tools for these languages .
Outcome: The authors review the research and the available NLP systems on indigenous languages of the Americas . they argue that the lack of resources and tools can have a negative impact on the communities which depend on these languages .
Interpretability for Morphological Inflection: from Character-level Predictions to Subword-level Rules (2021.eacl-main)

Copied to clipboard

Challenge: Neural models for morphological inflection have recently attained very high results, but their interpretation remains challenging.
Approach: They propose a linguistically-motivated variant to the encoder-decoder model with attention that incorporates a character-level cross-attention mechanism and a self-attention module over substrings of the input.
Outcome: The proposed model performs well on three typologically-different languages and is highly interpretable.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations