Challenge: Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets.
Approach: They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies.
Outcome: The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies.

Similar Papers

From Linguistic Linked Data to Big Data (2024.lrec-main)

Copied to clipboard

Challenge: Language data on the LOD cloud has grown in number, size, and variety . Linked (Open) Data (LLOD) is a standardized way of representing and sharing linguistic datasets .
Approach: They propose to combine LLOD and Big Data to improve interoperability of linguistic datasets . they propose to use a machine-readable format to represent and share linguistic data .
Outcome: This paper examines the use cases of Linked (Open) Data and Big Data in language data.
Towards a Linked Open Data Edition of Sumerian Corpora (L18-1)

Copied to clipboard

Challenge: Linguistic Linked Open Data (LLOD) is a flourishing line of research in the language resource community . existing LLOD standards and vocabularies are not widely used in this community despite its popularity .
Approach: They propose to use Linguistic Linked Open Data to link a Sumerian corpus with lexical resources . they use a linguistically annotated archive to create a corpus of cuneiform texts .
Outcome: The proposed LLOD framework is used in assyriology, with philological resources underrepresented . the proposed framework is based on a linguistically annotated corpus of Sumerian texts .
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)

Copied to clipboard

Challenge: Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources.
Approach: They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied .
Outcome: The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem .
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages .
Approach: They provide an overview of some of the major online data sources available for thousands of languages.
Outcome: The proposed language technologies are based on the data available for thousands of languages.
Social Intelligence Data Infrastructure: Structuring the Present and Navigating the Future (2024.findings-acl)

Copied to clipboard

Challenge: Existing work on social intelligence in NLP does not provide a coherent subfield for researchers to analyze and identify research gaps and future directions.
Approach: They build a social AI taxonomy and a data library of 480 NLP datasets to analyze existing datasets and evaluate language models’ performance in different social intelligence aspects.
Outcome: The proposed infrastructure analyzes existing dataset efforts and evaluates language models’ performance in different social intelligence aspects.
Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan (2024.lrec-main)

Copied to clipboard

Challenge: Aina Project aims to provide Catalan with the resources needed to keep its relevance in AI/NLP applications.
Approach: They propose a set of strategies to consider when improving technology support for a mid- or low-resource language . they propose annotated datasets and a framework to make models ready to use .
Outcome: The Aina Project aims to provide Catalan with the necessary resources to keep its relevance in AI/NLP-related industry and research.
Managing Public Sector Data for Multilingual Applications Development (L18-1)

Copied to clipboard

Challenge: eTranslation is a digital service that enables multilingual communication across public administrations in 30 European countries.
Approach: They propose to develop a repository infrastructure specifically tailored to the needs of the eTranslation service of the European Commission.
Outcome: The ELRC-SHARE repository is designed and developed specifically for the eTranslation service of the European Commission.
Language Data Sharing in European Public Services – Overcoming Obstacles and Creating Sustainable Data Sharing Infrastructures (2020.lrec-1)

Copied to clipboard

Challenge: Data is key in training modern language technologies.
Approach: They summarise findings of first pan-European study on barriers to language data sharing . they identify structural challenges, disposition towards CAT tools and lack of digital skills . overcoming language barriers is one of the main challenges european citizens face .
Outcome: The paper summarises the findings of the first pan-European study on barriers to language data sharing . the findings highlight the barriers and recommend solutions to overcome them .
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment.
Approach: They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge.
Outcome: The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations