Papers by Jorge Gracia

6 papers
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)

Copied to clipboard

Challenge: Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets.
Approach: They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies.
Outcome: The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies.
Building MUSCLE, a Dataset for MUltilingual Semantic Classification of Links between Entities (2024.lrec-main)

Copied to clipboard

Challenge: In this paper we present a dataset for MUltilingual Lexical Relation Classification (LRC) systems with 27K pairs of universal concepts selected from Wikidata, a large and highly multilingual factual Knowledge Graph (KG).
Approach: They propose a dataset for MUltilingual lexico-semantic Classification of Links between Entities using 27K pairs of universal concepts selected from Wikidata.
Outcome: The proposed dataset bridges lexical and conceptual semantics, avoids linguistic memorization, is domain-balanced across entities, and enables enrichment and hierarchical information retrieval.
Cross-Lingual Link Discovery for Under-Resourced Languages (2022.lrec-1)

Copied to clipboard

Challenge: Linked data paradigms can be used to solve under-resourced languages' problem of under-utilization of resources.
Approach: They propose a paradigm for cross-lingual link discovery that can be applied to under-resourced languages . they argue that techniques for cross language linking can be readily applied .
Outcome: The proposed technologies can be applied to under-resourced languages, the authors argue . the authors show that the Linked Data paradigm can be used to solve the problem .
MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations (2024.lrec-main)

Copied to clipboard

Challenge: Prior work has focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs) with some exceptions.
Approach: They propose to use a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian as an experiment on cross-lingual transfer of relational knowledge.
Outcome: The proposed dataset is adapted from a BATS-based dataset in 15 languages including low-resource languages such as Bambara, Lithuanian, and Albanian.
Orchestrating NLP Services for the Legal Domain (2020.lrec-1)

Copied to clipboard

Challenge: a legal technology system under development in the EU is based on semantic services and a multilingual legal knowledge Graph.
Approach: They propose a workflow manager that enables flexible orchestration of workflows . they describe different use cases and propose prototypical solutions .
Outcome: The proposed system is based on a set of natural language processing and document curation services and a multilingual legal knowledge graph that contains semantic information and meaningful references to legal documents.
No clues good clues: out of context Lexical Relation Classification (2023.acl-long)

Copied to clipboard

Challenge: Pre-trained language models (PTLMs) are used to predict lexical relations between words.
Approach: They propose to use pre-trained language models to fine-tune and exploit verbalized text for linguistically motivated tasks.
Outcome: The proposed model outperforms graded Lexical Entailment and lexical relation classification with very simple prompts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations