Papers by Thierry Declerck

9 papers
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)

Copied to clipboard

Challenge: Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets.
Approach: They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies.
Outcome: The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies.
Towards a new Ontology for Sign Languages (2022.lrec-1)

Copied to clipboard

Challenge: Linked Data (LD) compliant datasets for sign languages are not available in the LLOD cloud.
Approach: They propose to create an ontology for representing constitutive elements of Sign Languages (SL) they propose to publish such data in the Linguistic Linked Open Data cloud.
Outcome: The proposed ontology can be used to represent sign languages in the Linguistic Linked Open Data cloud.
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)

Copied to clipboard

Challenge: Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities.
Approach: They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance.
Outcome: The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks.
Using Wiktionary to Create Specialized Lexical Resources and Datasets (2022.lrec-1)

Copied to clipboard

Challenge: Using Wiktionary data to build specialized lexical datasets can be used for evaluating or improving NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Machine Translation (MT).
Approach: They propose to use Wiktionary data to create specialized lexical datasets that can be used for evaluating or improving NLP tasks.
Outcome: The proposed datasets can be used to improve and/or evaluate NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Sense Linking (SL), or machine translation (MT).
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)

Copied to clipboard

Challenge: a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources .
Approach: They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment .
Outcome: The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language.
Language Data Sharing in European Public Services – Overcoming Obstacles and Creating Sustainable Data Sharing Infrastructures (2020.lrec-1)

Copied to clipboard

Challenge: Data is key in training modern language technologies.
Approach: They summarise findings of first pan-European study on barriers to language data sharing . they identify structural challenges, disposition towards CAT tools and lack of digital skills . overcoming language barriers is one of the main challenges european citizens face .
Outcome: The paper summarises the findings of the first pan-European study on barriers to language data sharing . the findings highlight the barriers and recommend solutions to overcome them .
European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management (L18-1)

Copied to clipboard

Challenge: European Language Resource Coordination (ELRC) initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Approach: They propose to initiate actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Outcome: The European Language Resource Coordination (ELRC) consortium initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries.
Comparing Pretrained Multilingual Word Embeddings on an Ontology Alignment Task (L18-1)

Copied to clipboard

Challenge: Existing word embeddings capture a string's semantics and can be trained for multiple languages.
Approach: They propose to compare three different multilingual pretrained word embedding repositories with a string-matching baseline and use it to compute semantic similarities of strings in different languages.
Outcome: The proposed method produces correct alignments on a non-standard dataset on all four languages.
An Integrated Formal Representation for Terminological and Lexical Data included in Classification Schemes (L18-1)

Copied to clipboard

Challenge: e-lexicography is a field of study dealing with the automated creation of specialized multilingual dictionaries from structured data.
Approach: They propose to use a SKOS-XL vocabulary for modelling the multilingual terminological part of comparable taxonomies and OntoLex-Lemon for modelling multilingual lexical entries.
Outcome: The proposed model can be explicitly cross-linked in the context of the Linguistic Linked Open Data (LLOD).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations