Challenge: XML data formats that represent a semantic data model are difficult to convert into other formats . this difficulty causes a problem in the reuse of data especially in a personal data management environment.
Approach: They propose a new approach to transform a link structure into an instance structure on a marked-up scheme.
Outcome: The proposed format is based on a so-called standoff-style data format in XML . the proposed format injures the mobility of data because it is hard to convert it into other formats .

Similar Papers

Multi-lingual Entity Discovery and Linking (P18-5)

Copied to clipboard

Challenge: This tutorial reviews the framework of cross-lingual EL and motivates it as a broad paradigm for the Information Extraction task.
Approach: This tutorial will review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Outcome: The aim of this tutorial is to review the framework of cross-lingual EL and motivate it as a broad paradigm for the Information Extraction task.
Mitigating Data Scarcity in Semantic Parsing across Languages with the Multilingual Semantic Layer and its Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have advanced significantly in understanding human text, but semantic representations remain crucial for various applications.
Approach: They introduce a multilingual semantic layer which decouples from disambiguation and external inventories and simplifies the task.
Outcome: The proposed model reduces performance gap between languages and annotators by enabling them to understand semantic relations between concepts in any language.
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)

Copied to clipboard

Challenge: Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets.
Approach: They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies.
Outcome: The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies.
Managing Public Sector Data for Multilingual Applications Development (L18-1)

Copied to clipboard

Challenge: eTranslation is a digital service that enables multilingual communication across public administrations in 30 European countries.
Approach: They propose to develop a repository infrastructure specifically tailored to the needs of the eTranslation service of the European Commission.
Outcome: The ELRC-SHARE repository is designed and developed specifically for the eTranslation service of the European Commission.
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)

Copied to clipboard

Challenge: a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research .
Approach: a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary .
Outcome: a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web .
Placing multi-modal, and multi-lingual Data in the Humanities Domain on the Map: the Mythotopia Geo-tagged Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Using mythology as a starting point, visitors of Northern Greece will have a multi-faceted experience using a corpus of textual data supplemented with images, and video.
Approach: They propose to integrate a multi-lingual corpus with a dedicated database with advanced indexing, linking and search functionalities into a platform for scholarly research in the digital humanities.
Outcome: The proposed infrastructure will be integrated into a platform aimed at providing a multi-faceted experience to visitors of Northern Greece using mythology as a starting point.
Understanding Cross-Lingual Alignment—A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Cross-lingual alignment is the meaningful similarity of representations across languages in multilingual language models.
Approach: They propose a taxonomy of methods to improve cross-lingual alignment . they argue that an effective trade-off between language-neutral and language-specific information is key .
Outcome: The proposed methods can be applied to encoder models and encoder-decoder-only models . they show that language-neutral and language-specific information is key .
Connecting Language Technologies with Rich, Diverse Data Sources Covering Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: Existing data sources for many thousands of languages are rich and diverse . Efforts are ongoing to extend technology to many more of the world's languages .
Approach: They provide an overview of some of the major online data sources available for thousands of languages.
Outcome: The proposed language technologies are based on the data available for thousands of languages.
LinguaMeta: Unified Metadata for Thousands of Languages (2024.lrec-main)

Copied to clipboard

Challenge: LinguaMeta is a unified repository of language metadata for thousands of languages.
Approach: They introduce LinguaMeta, a unified resource for language metadata for thousands of languages.
Outcome: The proposed resource is intended for use by researchers and organizations who aim to extend technology to thousands of languages.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations