Agettivu, Aggitivu o Aghjettivu? POS Tagging Corsican Dialects (2024.lrec-main)

Copied to clipboard

Challenge: a series of experiments towards POS tagging Corsican are presented . POS tags are used to tag a less-resourced language spoken in corsica .
Approach: They present a series of experiments towards POS tagging Corsican . they first contribute to the first gold standard POS-tagged corpus for Corsica .
Outcome: The proposed model is the first POS tagger for Corsican, reaching an accuracy of 93.38%.

Similar Papers

Part-of-Speech Tagging on an Endangered Language: a Parallel Griko-Italian Resource (C18-1)

Copied to clipboard

Challenge: a recent study examines POS tagging techniques on endangered languages . most natural language processing applications have been tested on only a handful of languages - a problem that is compounded by the lack of standard orthography.
Approach: They evaluate POS tagging techniques on an endangered language, Griko . they use a semi-supervised method with cross-lingual transfer to achieve better accuracy .
Outcome: The proposed method achieves 72.9% accuracy on a sample of 114 narratives in a language . the proposed method improves by 21 percentage points over previous methods .
Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Using a crowdsourcing platform, we collected 18,917 annotations for a less-resourced French regional language, Alsatian.
Approach: They developed a platform that allows people to gather part-of-speech annotations on a variety of corpora and train a first tagger specific to Alsatian.
Outcome: The proposed method is valid for Alsatian and can be adapted to other languages.
POS Tagging for the Endangered Dagur Language (2024.lrec-main)

Copied to clipboard

Challenge: a recent study has focused on the so-called "dominant" languages, but it has not been inclusive in terms of language equality.
Approach: They propose to use POS tagging to automatically annotate Dagur, an endangered Mongolic language . they use a manually annotated corpus to test transfer of models from other languages .
Outcome: The proposed method can be used to document and revitalize endangered languages . the proposed model can be trained on Buryat, the only Mongolic language included in the corpus .
What data should I include in my POS tagging training set? (2025.findings-emnlp)

Copied to clipboard

Challenge: POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages .
Approach: They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources .
Outcome: The proposed training set for Indigenous and endangered languages performs better than random sampling.
Towards a Corsican Basic Language Resource Kit (2020.lrec-1)

Copied to clipboard

Challenge: a roadmap has been set out for the development of a basic language resource kit for the Corsican language . the goal is to improve the availability of resources and tools for the language based on the Banque de Données Langue Corse project .
Approach: a team of researchers from univ-corse is developing a basic language resource kit for the corsican language . they aim to collect corpora, set up a concordancer, set-up language detection tool, build an electronic dictionary and add a part-of-speech tagger .
Outcome: the goal is to improve the availability of resources and tools for the Corsican language . the roadmap sets out the actions to be undertaken: collection of corpora, setting up of a concordancer, language detection tool, electronic dictionary and part-of-speech tagger.
A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages (N19-1)

Copied to clipboard

Challenge: Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but taggers need to ground their clusters as well.
Approach: They propose an approach for low-resource unsupervised part of speech (POS) tagging that yields fully grounded output and requires no labeled training data.
Outcome: The proposed method achieves reasonable performance across languages, including Sinhalese and Kinyarwanda, with no labeled training data.
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)

Copied to clipboard

Challenge: a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages.
Approach: They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data .
Outcome: The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality .
Multi-Dialect Arabic POS Tagging: A CRF Approach (L18-1)

Copied to clipboard

Challenge: Existing work on dialectal POS tagging is rather scant with POS tags for most dialects being nonexistent or of limited availability.
Approach: They propose a dataset of POS-tagged Arabic tweets in four major dialects and a tagging guideline for each dialect.
Outcome: The proposed model can tag four different dialects with an average accuracy of 89.3%.
Introducing a Large-Scale Dataset for Vietnamese POS Tagging on Conversational Texts (2020.lrec-1)

Copied to clipboard

Challenge: POS taggers are trained on informal texts which contain many informal inputs such as acronyms, abbreviations, out-of-vocabulary words, etc.
Approach: They propose a large-scale human-labeled dataset for the Vietnamese POS tagging task on conversational texts and develop an annotation guideline to manually annotate 16.310K sentences using this guideline.
Outcome: The proposed tagging scheme achieved 93.36% accuracy score and higher than the model with handcrafted features and fine-tuning BERT.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations