A Framework for Shared Agreement of Language Tags beyond ISO 639 (2020.lrec-1)

Copied to clipboard

Challenge: Identification and annotation of languages in an unambiguous and standardized way is essential for the description of linguistic data.
Approach: They propose a pattern that extends the BCP 47 sub-tag ‘privateuse’ and is able to overcome the limits of BCP47 and ISO 639.
Outcome: The proposed pattern overcomes the limitations of BCP 47 and ISO 639 for the identification of lesser-known languages, endangered languages, regional varieties or historical stages of a language.

Similar Papers

MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)

Copied to clipboard

Challenge: MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement.
Approach: They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs.
Outcome: The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs.
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Currently, existing systems cannot accurately identify most of the world's 7000 languages due to lack of data and computational challenges.
Approach: They propose a misprediction-resolution hierarchical model, LIMIT, that reduces error by 55% on a children's stories dataset and by 40% on 'fLORES-200' benchmark.
Outcome: The proposed model reduces error by 55% on the MCS-350 and 40% on the FLORES-200 benchmarks.
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)

Copied to clipboard

Challenge: Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources.
Approach: They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied .
Outcome: The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem .
Related Works in the Linguistic Data Consortium Catalog (2020.lrec-1)

Copied to clipboard

Challenge: Existing metadata standards for Related Works are used to define relations between language resources.
Approach: They describe the development and implementation of a Related Works schema and the steps to implementation.
Outcome: The proposed schema has been implemented in the Linguistic Data Consortium's (LDC) Catalog.
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)

Copied to clipboard

Challenge: Language is a powerful means of communication and should be regarded as more than just a collection of tokens.
Approach: They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices.
Outcome: The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices.
Interoperability of Language-related Information: Mapping the BLL Thesaurus to Lexvo and Glottolog (L18-1)

Copied to clipboard

Challenge: The Bibliography of Linguistic Literature (BLL Thesaurus) has been used since 2013 in the context of the Lin gu is tik portal, a hub for linguistically relevant information.
Approach: They propose to use Lexvo and Glottolog to facilitate interoperability between the BLL Thesaurus and terminological repositories in the Linguistic Linked Open Data cloud.
Outcome: The proposed model is based on Lexvo and Glottolog and is able to connect to the Linguistic Linked Open Data cloud.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)

Copied to clipboard

Challenge: Using ISOCat successor solutions, annotation standards have been developed since 2010 .
Approach: They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies .
Outcome: The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes.
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)

Copied to clipboard

Challenge: a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE.
Approach: They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation.
Outcome: The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations