Platforms for Non-speakers Annotating Names in Any Language (P18-4)

Copied to clipboard

Challenge: Traditionally, native speakers of a language have been asked to annotate a corpus in that language.
Approach: They propose two annotation platforms that allow an English speaker to annotate names for any language without knowing the language.
Outcome: The proposed annotations achieved state-of-the-art performance on two surprise languages and ten languages at TAC-KBP EDL2017.

Similar Papers

Dragonfly: Advances in Non-Speaker Annotation for Low Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using semantic and contextual information, non-speakers of a language familiar with the Latin script can produce high quality named entity annotations to support construction of . name tagger.
Approach: They propose a procedure for annotating low resource languages using Dragonfly that others can use.
Outcome: The proposed procedure improves the performance of NER models on native speaker and non-speaker annotations in low resource languages.
Zero-Shot Cross-lingual Name Retrieval for Low-Resource Languages (D19-61)

Copied to clipboard

Challenge: a novel name retrieval method is proposed for languages with no annotations or training data.
Approach: They propose a method which relies on zero annotation or resources from the target language . they pre-train an orthographic encoder using Wikipedia inter-lingual links from dozens of languages .
Outcome: The proposed method shows 11.6% improvement over state-of-the-art methods.
Should All Cross-Lingual Embeddings Speak English? (2020.acl-main)

Copied to clipboard

Challenge: lexicon induction evaluation dictionaries are mostly between English and another language, and the English hub is selected by default as the hub . lexiconic embeddings are often learned with a two-step process, whether under bilingual or multilingual settings.
Approach: They propose to use English as the hub language for lexicon induction evaluation . they also expand a standard English-centered evaluation dictionary collection to include all language pairs .
Outcome: The proposed method can significantly improve lexicon induction performance over multiple languages.
Rethinking Annotation: Can Language Learners Contribute? (2023.acl-long)

Copied to clipboard

Challenge: Researchers have traditionally recruited native speakers to provide annotations for benchmark datasets, but there are languages for which recruiting native speakers is difficult.
Approach: They recruit 36 language learners and provide two types of additional resources and perform mini-tests to measure their language proficiency.
Outcome: The proposed method improves learners' language proficiency in terms of vocabulary and grammar.
Cross-lingual Multi-Level Adversarial Transfer to Enhance Low-Resource Name Tagging (N19-1)

Copied to clipboard

Challenge: Low-resource language name tagging is an important but challenging task.
Approach: They propose a neural architecture that leverages multi-level adversarial transfer to improve name tagging for low-resource languages.
Outcome: The proposed approach outperforms previous approaches on CoNLL data sets.
Distant Supervision from Disparate Sources for Low-Resource Part-of-Speech Tagging (D18-1)

Copied to clipboard

Challenge: Low-resource languages lack manual annotated data to learn basic models such as part-of-speech (POS) taggers.
Approach: They propose a cross-lingual neural part-of-speech tagger that learns from disparate sources of distant supervision in a uniform framework.
Outcome: The proposed model scales to hundreds of low-resource languages without access to gold annotated data.
What data should I include in my POS tagging training set? (2025.findings-emnlp)

Copied to clipboard

Challenge: POS tagging is a crucial task for descriptive linguistics and language documentation . POS tags are not available in all languages, but are used for training sets for understudied languages .
Approach: They compare POS tagging with in-context learning, active learning, and random sampling . they find that POS can deliver reasonable results for communities with limited resources .
Outcome: The proposed training set for Indigenous and endangered languages performs better than random sampling.
Unsupervised Cross-Lingual Representation Learning (P19-4)

Copied to clipboard

Challenge: a comprehensive survey of cutting-edge weakly-supervised and unsupervised cross-lingual word representations is presented .
Approach: This tutorial provides a comprehensive survey of recent work on weakly-supervised and unsupervised cross-lingual word representations.
Outcome: This tutorial provides a comprehensive survey of cutting-edge weakly-supervised and unsupervised word representations.
Error Analysis of Uyghur Name Tagging: Language-specific Techniques and Remaining Challenges (L18-1)

Copied to clipboard

Challenge: despite efforts at name tagging, there is limited understanding on the performance ceiling . despite the high-resource language, there are very few natural language processing tools available .
Approach: They propose to use a machine learning model to identify Uyghur name tagger errors . they conclude that such a model is unlikely to be effective for Uygur, or low-resource languages .
Outcome: The proposed model is unlikely to be effective for Uyghur, or low-resource languages in general, the authors argue . they show that the proposed model can be used for high-res languages with superficial features .
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations