| Challenge: | Existing resources and tools for the Galician language are lacking for other less-resourced languages, such as statistical tools for lemmatization and Named Entity Recognition. |
| Approach: | They propose to develop a manually revised corpus for POS tagging and lemmatization, and a new manually annotated corpus to train existing statistical tools for the Galician language. |
| Outcome: | The proposed resources include a new corpus for POS tagging and lemmatization, and a manually annotated corpus to handle Named Entity recognition. |
Similar Papers
Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | This work presents a corpus manually annotated with named entities for six Slavic languages . |
| Approach: | They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models . |
| Outcome: | The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier. |
Developing NLP Tools with a New Corpus of Learner Spanish (2020.lrec-1)
Copied to clipboard
Sam Davidson, Aaron Yamada, Paloma Fernandez Mira, Agustina Carando, Claudia H. Sanchez Gutierrez, Kenji Sagae
| Challenge: | Currently, there is little research on the development of effective NLP tools for the L2 classroom. |
| Approach: | They propose to use an annotated corpus of Spanish learner text to analyze developmental patterns and to develop a grammatical error correction system for Spanish learners. |
| Outcome: | The proposed system is based on annotated learner corpus of Spanish learners and includes error annotations and corrected text. |
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata (2024.lrec-main)
Copied to clipboard
| Challenge: | ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages . |
| Approach: | They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation . |
| Outcome: | The proposed resource is the largest of its type to date and performs well on 10 languages. |
A Diverse Set of Freely Available Linguistic Resources for Turkish (2023.acl-long)
Copied to clipboard
| Challenge: | despite the abundance of Turkish speakers, linguistic resources for natural language processing remain scarce. |
| Approach: | They propose a set of freely available linguistic resources for Turkish natural language processing . they provide corpora and pretrained models to help practitioners build their own applications . |
| Outcome: | The proposed linguistic resources are first of their kind and easy to use in a broad range of implementations. |
Give your Text Representation Models some Love: the Case for Basque (2020.lrec-1)
Copied to clipboard
Rodrigo Agerri, Iñaki San Vicente, Jon Ander Campos, Ander Barrena, Xabier Saralegi, Aitor Soroa, Eneko Agirre
| Challenge: | Word embeddings and pre-trained language models are expensive to train and are often used by small companies and research groups to build their own. |
| Approach: | They propose to use word embeddings and pre-trained language models to build rich representations of text and improve NLP tasks. |
| Outcome: | The proposed models perform better than publicly available versions in downstream NLP tasks for Basque. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)
Copied to clipboard
Špela Arhar Holdt, Jaka Čibej, Kaja Dobrovoljc, Tomaž Erjavec, Polona Gantar, Simon Krek, Tina Munda, Nejc Robida, Luka Terčon, Slavko Zitnik
| Challenge: | a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years. |
| Approach: | They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene. |
| Outcome: | The revised corpus, built on its predecessor, doubles in size and depth of annotation layers. |
Advances in Pre-Training Distributed Word Representations (L18-1)
Copied to clipboard
| Challenge: | Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications. |
| Approach: | They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations. |
| Outcome: | The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The Interplay between Metaphors and NLP (2026.acl-tutorials)
Copied to clipboard
| Challenge: | This tutorial will provide an overview of the metaphor processing field. |
| Approach: | This tutorial will provide an overview of the metaphor processing field . it will focus on recent directions opened by LLMs for metaphor interpretation . |
| Outcome: | The tutorial will discuss the influence of various metaphor theories on the creation of annotated resources and models. |