| Challenge: | Identification and annotation of languages in an unambiguous and standardized way is essential for the description of linguistic data. |
| Approach: | They propose a pattern that extends the BCP 47 sub-tag ‘privateuse’ and is able to overcome the limits of BCP47 and ISO 639. |
| Outcome: | The proposed pattern overcomes the limitations of BCP 47 and ISO 639 for the identification of lesser-known languages, endangered languages, regional varieties or historical stages of a language. |
Similar Papers
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs (2026.tacl-1)
Copied to clipboard
| Challenge: | MultiBLiMP 1.0 is a massively multilingual benchmark of linguistic minimal pairs covering 101 languages and 2 types of subject-verb agreement. |
| Approach: | They propose to use multilingual benchmarks to evaluate linguistic minimal pairs in 101 languages and 2 types of subject-verb agreement to create the minimal pairs. |
| Outcome: | The proposed benchmark covers 101 languages and 2 types of subject-verb agreement, and contains more than 128,000 minimal pairs. |
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
LIMIT: Language Identification, Misidentification, and Translation using Hierarchical Models in 350+ Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Currently, existing systems cannot accurately identify most of the world's 7000 languages due to lack of data and computational challenges. |
| Approach: | They propose a misprediction-resolution hierarchical model, LIMIT, that reduces error by 55% on a children's stories dataset and by 40% on 'fLORES-200' benchmark. |
| Outcome: | The proposed model reduces error by 55% on the MCS-350 and 40% on the FLORES-200 benchmarks. |
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)
Copied to clipboard
Michael Rosner, Sina Ahmadi, Elena-Simona Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Gilles Sérasset, Ciprian-Octavian Truică
| Challenge: | Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources. |
| Approach: | They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied . |
| Outcome: | The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem . |
Related Works in the Linguistic Data Consortium Catalog (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing metadata standards for Related Works are used to define relations between language resources. |
| Approach: | They describe the development and implementation of a Related Works schema and the steps to implementation. |
| Outcome: | The proposed schema has been implemented in the Linguistic Data Consortium's (LDC) Catalog. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Interoperability of Language-related Information: Mapping the BLL Thesaurus to Lexvo and Glottolog (L18-1)
Copied to clipboard
| Challenge: | The Bibliography of Linguistic Literature (BLL Thesaurus) has been used since 2013 in the context of the Lin gu is tik portal, a hub for linguistically relevant information. |
| Approach: | They propose to use Lexvo and Glottolog to facilitate interoperability between the BLL Thesaurus and terminological repositories in the Linguistic Linked Open Data cloud. |
| Outcome: | The proposed model is based on Lexvo and Glottolog and is able to connect to the Linguistic Linked Open Data cloud. |
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models. |
| Approach: | They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key . |
| Outcome: | The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key . |
Annotation Interoperability for the Post-ISOCat Era (2020.lrec-1)
Copied to clipboard
| Challenge: | Using ISOCat successor solutions, annotation standards have been developed since 2010 . |
| Approach: | They describe ISOCat successor solutions and annotation standardization efforts since 2010 . they describe low-cost harmonization of post-ISOCat vocabularies by means of linked ontologies . |
| Outcome: | The proposed ontologies are linked with the Ontologie of Linguistic Annotation and ISOCat, the GOLD ontology, the Typological Database Systems ontological and a large number of annotation schemes. |
Lightweight Grammatical Annotation in the TEI: New Perspectives (L18-1)
Copied to clipboard
| Challenge: | a small set of descriptive devices have been made available for lightweight linguistic annotation . merit of a predefined TEI tagset is the homogeneity of tagging and better interoperability of simple linguistic resources encoded in the TE. |
| Approach: | They propose a new attribute class that would gather token-level attributes facilitating simple linguistic annotation. |
| Outcome: | The proposed attribute class addresses community feedback on the lack of a specific tagset for lightweight linguistic annotation within the TEI. |