Towards Flexible Cross-Resource Exploitation of Heterogeneous Language Documentation Data (2020.lrec-1)
Copied to clipboard
| Challenge: | a paper on language resource overarching data analysis aims at addressing a complex resource landscape . major challenges arise from the need for cross-resource data analysis and a rather complex resource environment . |
| Approach: | a paper aims to develop methods for language resource overarching data analysis in the field of language documentation. |
| Outcome: | The proposed methods aim to solve the tension between unification of data sets and vocabularies and maximum openness for the integration of future resources and adaption of external information. |
Similar Papers
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)
Copied to clipboard
Michael Rosner, Sina Ahmadi, Elena-Simona Apostol, Julia Bosque-Gil, Christian Chiarcos, Milan Dojchinovski, Katerina Gkirtzou, Jorge Gracia, Dagmar Gromann, Chaya Liebeskind, Giedrė Valūnaitė Oleškevičienė, Gilles Sérasset, Ciprian-Octavian Truică
| Challenge: | Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources. |
| Approach: | They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied . |
| Outcome: | The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem . |
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)
Copied to clipboard
| Challenge: | LinguaMeta is a unified repository of language metadata for thousands of languages. |
| Approach: | They introduce LinguaMeta, a unified resource for language metadata for thousands of languages. |
| Outcome: | The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages. |
Unifying Cross-Lingual Transfer across Scenarios of Resource Scarcity (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to deal with resource scarcity have not been developed to deal effectively with the problem. |
| Approach: | They propose to use a set of tools to harness data from one or more high-resource "source" languages to compensate for a shortage of data in low-resourced "target" languages. |
| Outcome: | The proposed technique can be easily adapted to unseen languages, extending the range of the proposed technique and translation-based transfer more broadly. |
Local Languages, Third Spaces, and other High-Resource Scenarios (2022.acl-long)
Copied to clipboard
| Challenge: | In one view, languages exist on a resource continuum and the challenge is to scale existing solutions, bringing under-resourced languages into the high-resource world. |
| Approach: | They propose to scale existing solutions to bring under-resourced languages into the high-resource world by bringing standardised languages into high-level global information society. |
| Outcome: | The proposed language technology agendas address the diverse situations of the world's languages. |
Beyond Counting Datasets: A Survey of Multilingual Dataset Construction and Necessary Resources (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have examined the quality of labeled data in non-English languages. |
| Approach: | They annotate how datasets are created, input text and label sources, tools used to build them and what they study. |
| Outcome: | The results show that language-proficient NLP researchers' estimated availability correlates with dataset availability. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
A Survey on Cross-Lingual Summarization (2022.tacl-1)
Copied to clipboard
| Challenge: | Cross-lingual summarization is a task of generating a summary in one language for a given document in a different language. |
| Approach: | They present a systematic review of the literature on cross-lingual summarization . they summarize previous efforts and compare them with each other . |
| Outcome: | The proposed approach is compared with previous approaches and summarizes them to provide a deeper analysis. |
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)
Copied to clipboard
| Challenge: | despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested. |
| Approach: | They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs. |
| Outcome: | The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code. |
Linghub2: Language Resource Discovery Tool for Language Technologies (2022.lrec-1)
Copied to clipboard
| Challenge: | Linghub is a platform for language resources that can be used to find and retrieve data . the platform is based on a popular open source data management system, DSpace . |
| Approach: | This work describes a rejuvenation and modernisation of the 2015 platform into using a popular open source data management system, DSpace, as foundation. |
| Outcome: | Linghub2 1 aims to help language resources and technology users find and retrieve relevant data . the new platform, Ling hub2, contains updated and extended resources and more languages offered . |
A Corpus for Multilingual Document Classification in Eight Languages (L18-1)
Copied to clipboard
| Challenge: | a subset of the Reuters corpus volume 2 is used to evaluate cross-lingual document classification . current best practice is to evaluate document classification on resources in one language and transfer it to another without additional resources. |
| Approach: | They propose to use a subset of the Reuters corpus to evaluate cross-lingual document classification . they propose to add Italian, Russian, Japanese and Chinese to the subset . |
| Outcome: | The proposed subset of the Reuters corpus has balanced class priors for eight languages. |