Challenge: POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages .
Approach: They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources .
Outcome: The proposed training set for Indigenous and endangered languages performs better than random sampling.

Similar Papers

To POS Tag or Not to POS Tag: The Impact of POS Tags on Morphological Learning in Low-Resource Settings (2021.acl-long)

Copied to clipboard

Challenge: Part-of-Speech (POS) tags are routinely included in many NLP tasks.
Approach: They propose to use POS tags to examine morphological learning in low-resource languages . they find that POS tagging improves joint segmentation and glossing .
Outcome: The proposed task is tested on two identical datasets with the Transformer architecture.
Make the Best of Cross-lingual Transfer: Evidence from POS Tagging with over 100 Languages (2022.acl-long)

Copied to clipboard

Challenge: Existing studies on cross-lingual generalisability of large pre-trained models use English training data and test data in unseen languages.
Approach: They propose to use multilingual pre-trained models to model cross-lingual transfer in a selection of target languages.
Outcome: The proposed model can be used to improve cross-lingual transfer performance in low-resource languages with no labeled training data.
Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text (L18-1)

Copied to clipboard

Challenge: a new approach to POS tagging noisy user generated text is proposed . word embeddings are trained on a noisy corpus to address both normalization and POS.
Approach: They propose to use word embeddings to normalize text before tagging it, while a gated neural network based tagger handles the remaining errors.
Outcome: The proposed approach normalizes some errors before tagging, while a gated neural network handles the remaining errors.
Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech Tagging (D18-1)

Copied to clipboard

Challenge: Low-resource languages lack manual annotated data to learn basic models such as part-of-speech (POS) taggers.
Approach: They propose a cross-lingual neural part-of-speech tagger that learns from disparate sources of distant supervision in a uniform framework.
Outcome: The proposed model scales to hundreds of low-resource languages without access to gold annotated data.
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have examined the quality of labeled data in non-English languages.
Approach: They annotate how datasets are created, input text and label sources, tools used to build them and what they study.
Outcome: The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability.
A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages (N19-1)

Copied to clipboard

Challenge: Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but taggers need to ground their clusters as well.
Approach: They propose an approach for low-resource unsupervised part of speech (POS) tagging that yields fully grounded output and requires no labeled training data.
Outcome: The proposed method achieves reasonable performance across languages, including Sinhalese and Kinyarwanda, with no labeled training data.
What can we learn from Semantic Tagging? (D18-1)

Copied to clipboard

Challenge: a recent study shows that multi-task learning improves performance of NLP tasks by exploiting similarities between tasks.
Approach: They employ semantic tagging as an auxiliary task for three NLP tasks . they compare full neural network sharing, partial neural network shared and learning what to share .
Outcome: The proposed model improves for part-of-speech tagging, universal dependency parsing and natural language inference.
Data, Data Everywhere: A Guide for Pretraining Dataset Construction (2024.emnlp-main)

Copied to clipboard

Challenge: Recent language models have impressive capabilities on a number of evaluation areas.
Approach: They conduct systematic analysis of pretraining set construction to identify which methods yield the greatest gains in model accuracy.
Outcome: The proposed method can be used to refine and improve a pretraining set.
Parsing linearizations appreciate PoS tags - but some are fussy about errors (2022.aacl-short)

Copied to clipboard

Challenge: Recent work on the impact of PoS tags on graph- and transition-based parsers suggests that they are only useful when tagging accuracy is prohibitively high or in low-resource scenarios.
Approach: They examine the impact of PoS tags on graph- and transition-based parsers and propose to use them in a new paradigm for sequence labeling.
Outcome: The proposed model is best when tagging accuracy and resource availability are high.
Diversifying language models for lesser-studied languages and language-usage contexts: A case of second language Korean (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing morpheme parsers/taggers do not work reliably and optimally for L2 data.
Approach: They train a neural network model on varying L2 datasets and measure its morpheme parsing/POS tagging performance on L2 test sets.
Outcome: The proposed model excels in domain-specific tokenization and POS tagging compared to the baseline model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations