Challenge: Existing resources and tools for the Galician language are lacking for other less-resourced languages, such as statistical tools for lemmatization and Named Entity Recognition.
Approach: They propose to develop a manually revised corpus for POS tagging and lemmatization, and a new manually annotated corpus to train existing statistical tools for the Galician language.
Outcome: The proposed resources include a new corpus for POS tagging and lemmatization, and a manually annotated corpus to handle Named Entity recognition.

Similar Papers

Cross-lingual Named Entity Corpus for Slavic Languages (2024.lrec-main)

Copied to clipboard

Challenge: This work presents a corpus manually annotated with named entities for six Slavic languages .
Approach: They propose to manually annotate a corpus of names for six Slavic languages . they use a transformer-based neural network architecture to train multilingual models .
Outcome: The corpus consists of 5,017 documents on seven topics . each entity is described by a category, a lemma, and a unique cross-lingual identifier.
Developing NLP Tools with a New Corpus of Learner Spanish (2020.lrec-1)

Copied to clipboard

Challenge: Currently, there is little research on the development of effective NLP tools for the L2 classroom.
Approach: They propose to use an annotated corpus of Spanish learner text to analyze developmental patterns and to develop a grammatical error correction system for Spanish learners.
Outcome: The proposed system is based on annotated learner corpus of Spanish learners and includes error annotations and corrected text.
ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages Using Wikidata (2024.lrec-main)

Copied to clipboard

Challenge: ParaNames is a massively multilingual parallel name resource . it provides names for 16.8 million entities in over 400 languages .
Approach: They propose a massively multilingual parallel name resource with 140 million names . they use Wikidata to standardize the data and perform canonical name translation .
Outcome: The proposed resource is the largest of its type to date and performs well on 10 languages.
A Diverse Set of Freely Available Linguistic Resources for Turkish (2023.acl-long)

Copied to clipboard

Challenge: despite the abundance of Turkish speakers, linguistic resources for natural language processing remain scarce.
Approach: They propose a set of freely available linguistic resources for Turkish natural language processing . they provide corpora and pretrained models to help practitioners build their own applications .
Outcome: The proposed linguistic resources are first of their kind and easy to use in a broad range of implementations.
Give your Text Representation Models some Love: the Case for Basque (2020.lrec-1)

Copied to clipboard

Challenge: Word embeddings and pre-trained language models are expensive to train and are often used by small companies and research groups to build their own.
Approach: They propose to use word embeddings and pre-trained language models to build rich representations of text and improve NLP tasks.
Outcome: The proposed models perform better than publicly available versions in downstream NLP tasks for Basque.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
SUK 1.0: A New Training Corpus for Linguistic Annotation of Modern Standard Slovene (2024.lrec-main)

Copied to clipboard

Challenge: a training corpus for linguistic annotation of modern standard Slovene has been in continuous development for 15 years.
Approach: They introduce an upgrade of a training corpus for linguistic annotation of modern standard Slovene.
Outcome: The revised corpus, built on its predecessor, doubles in size and depth of annotation layers.
Advances in Pre-Training Distributed Word Representations (L18-1)

Copied to clipboard

Challenge: Pre-trained word representations are a building block of many Natural Language Processing and Machine Learning applications.
Approach: They propose to combine known tricks and a set of publicly available pre-trained word vector representations to train high-quality representations.
Outcome: The proposed models outperform the current state of the art on a number of tasks while maintaining a high training speed to scale to massive amount of data.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The Interplay between Metaphors and NLP (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an overview of the metaphor processing field.
Approach: This tutorial will provide an overview of the metaphor processing field . it will focus on recent directions opened by LLMs for metaphor interpretation .
Outcome: The tutorial will discuss the influence of various metaphor theories on the creation of annotated resources and models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations