Challenge: a corpus of Spanish newswire rich in unassimilated lexical borrowings is used to identify the language of a word.
Approach: They propose to annotate a corpus of Spanish newswire rich in unassimilated lexical borrowings and evaluate how models perform on this task.
Outcome: The proposed model outperforms models fed with subword embeddings and Transformer-based embeddables on the Spanish newswire corpus.

Similar Papers

Borrowing or Codeswitching? Annotating for Finer-Grained Distinctions in Language Mixing (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of tweets annotated for codeswitching and borrowing between Spanish and English is presented . the annotation does not treat common “internet-speak” as codeswitched when used in an otherwise monolingual context.
Approach: They present a new corpus of tweets annotated for codeswitching and borrowing between Spanish and English.
Outcome: The proposed corpus contains 9,500 tweets annotated with codeswitches, borrowings, and named entities.
Unsupervised Cross-Lingual Adaptation of Dependency Parsers Using CRF Autoencoders (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing work on cross-lingual adaptation of dependency parsers without annotated target corpora focuses on discriminative source parser ignoring unannotated corporata .
Approach: They propose to use unsupervised discriminative parsers to adapt dependency parser to unannotated target corpora without a supervised generative parsing method.
Outcome: The proposed method significantly outperforms previous methods.
A review of Spanish corpora annotated with negation (C18-1)

Copied to clipboard

Challenge: Existing corpora annotated with negation information are small and not always compatible . negation is a linguistic phenomenon that is not addressed in English .
Approach: They review existing corpora annotated with negation in Spanish and analyze compatibility . they propose to develop a supervised negation processing system for Spanish .
Outcome: The proposed system will not be able to merge the small corpora in Spanish due to lack of compatibility in annotations.
I Speak for the Árboles: Developing a Dependency Treebank for Spanish L2 and Heritage Speakers (2025.acl-srw)

Copied to clipboard

Challenge: Existing dependency treebanks for learner writing are limited due to morphosyntactic features.
Approach: They propose to use a dependency treebank for Spanish learner writing from the UC Davis COWSL2H corpus to incorporate lemmatization, POS tagging, and syntactic dependencies.
Outcome: The proposed treebanks are openly accessible to motivate future development of learner-oriented language technologies.
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)

Copied to clipboard

Challenge: Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power .
Approach: They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus.
Outcome: The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences.
Sequence-to-Sequence Spanish Pre-trained Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Spanish language models have demonstrated proficiency in natural language understanding and generation, but there is a scarcity of encoder-decoder models specifically designed for sequence-to-sequence tasks.
Approach: They propose to implement encoder-decoder architectures pre-trained on Spanish corpora . they use them to assess sequence-to-sequence tasks including summarization, question answering .
Outcome: The proposed models outperform models on sequence-to-sequence tasks in Spanish . the models show that they perform well across all tasks, the authors note .
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
GATITOS: Using a New Multilingual Lexicon for Low-resource Machine Translation (2023.emnlp-main)

Copied to clipboard

Challenge: a new study explores the effectiveness of bilingual lexica in machine translation models . cross-lingual vocabulary alignment is still highly imperfect in these models, despite the success of supervised and self-supervised training.
Approach: They use a resource to improve translation performance on 200-language models . they show that lexica is more reliable than human-translated data .
Outcome: The proposed approach improves on 200-language translation models with lexical data augmentation . the proposed approach is open-source and has 168 tail languages .
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for unsupervised bilingual lexicon induction depend on good quality static or contextual embeddings for both languages.
Approach: They propose a method for unsupervised bilingual lexicon induction between a related LRL and a high-resource language that only requires inference on a masked language model of the HRL.
Outcome: The proposed method performs well on low-resource languages with 5M tokens against Hindi . it is compared with existing methods on (mid-resourced) Marathi and Nepali .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations