Papers by Thierry Declerck
Recent Developments for the Linguistic Linked Open Data Infrastructure (2020.lrec-1)
Copied to clipboard
Thierry Declerck, John Philip McCrae, Matthias Hartung, Jorge Gracia, Christian Chiarcos, Elena Montiel-Ponsoda, Philipp Cimiano, Artem Revenko, Roser Saurí, Deirdre Lee, Stefania Racioppa, Jamal Abdul Nasir, Matthias Orlikowsk, Marta Lanau-Coronas, Christian Fäth, Mariano Rico, Mohammad Fazleh Elahi, Maria Khvalchik, Meritxell Gonzalez, Katharine Cooney
| Challenge: | Language data is rarely 'ready-to-use' and language technology specialists spend over 80% of their time cleaning, organizing and collecting language datasets. |
| Approach: | They propose a methodology for building data value chains based around language resources and language technologies that can be integrated by means of semantic technologies. |
| Outcome: | The proposed methodology is based on language resources and language technologies that can be integrated by means of semantic technologies. |
Towards a new Ontology for Sign Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | Linked Data (LD) compliant datasets for sign languages are not available in the LLOD cloud. |
| Approach: | They propose to create an ontology for representing constitutive elements of Sign Languages (SL) they propose to publish such data in the Linguistic Linked Open Data cloud. |
| Outcome: | The proposed ontology can be used to represent sign languages in the Linguistic Linked Open Data cloud. |
What’s the Meaning of Superhuman Performance in Today’s NLU? (2023.acl-long)
Copied to clipboard
Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli
| Challenge: | Recent research has focused on developing larger pretrained language models and introducing benchmarks such as SuperGLUE and SQuAD to measure their abilities. |
| Approach: | They propose to use benchmarks such as SuperGLUE and SQUAD to evaluate PLMs' abilities in language understanding, reasoning, and reading comprehension to assess their performance. |
| Outcome: | The proposed benchmarks have serious limitations affecting comparison between humans and PLMs and provide recommendations for fairer and more transparent benchmarks. |
Using Wiktionary to Create Specialized Lexical Resources and Datasets (2022.lrec-1)
Copied to clipboard
| Challenge: | Using Wiktionary data to build specialized lexical datasets can be used for evaluating or improving NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Machine Translation (MT). |
| Approach: | They propose to use Wiktionary data to create specialized lexical datasets that can be used for evaluating or improving NLP tasks. |
| Outcome: | The proposed datasets can be used to improve and/or evaluate NLP tasks, like Word Sense Disambiguation (WSD), Word-in-Context challenges (WiC), or Sense Linking (SL), or machine translation (MT). |
A Multilingual Evaluation Dataset for Monolingual Word Sense Alignment (2020.lrec-1)
Copied to clipboard
Sina Ahmadi, John Philip McCrae, Sanni Nimb, Fahad Khan, Monica Monachini, Bolette Pedersen, Thierry Declerck, Tanja Wissik, Andrea Bellandi, Irene Pisani, Thomas Troelsgård, Sussi Olsen, Simon Krek, Veronika Lipp, Tamás Váradi, László Simon, András Gyorffy, Carole Tiberius, Tanneke Schoonheim, Yifat Ben Moshe, Maya Rudich, Raya Abu Ahmad, Dorielle Lonke, Kira Kovalenko, Margit Langemets, Jelena Kallas, Oksana Dereza, Theodorus Fransen, David Cillessen, David Lindemann, Mikel Alonso, Ana Salgado, José Luis Sancho, Rafael-J. Ureña-Ruiz, Jordi Porta Zamorano, Kiril Simov, Petya Osenova, Zara Kancheva, Ivaylo Radev, Ranka Stanković, Andrej Perdih, Dejan Gabrovsek
| Challenge: | a new dataset aims to align monolingual dictionaries with a single sense level for 15 languages . this dataset covers a wide range of languages and resources . |
| Approach: | They propose to manually align monolingual dictionaries with possible semantic relationships . they use 15 languages to create a new baseline for the task of monolingual word sense alignment . |
| Outcome: | The proposed dataset covers 15 languages and covers the more challenging task of linking general-purpose language. |
Language Data Sharing in European Public Services – Overcoming Obstacles and Creating Sustainable Data Sharing Infrastructures (2020.lrec-1)
Copied to clipboard
| Challenge: | Data is key in training modern language technologies. |
| Approach: | They summarise findings of first pan-European study on barriers to language data sharing . they identify structural challenges, disposition towards CAT tools and lack of digital skills . overcoming language barriers is one of the main challenges european citizens face . |
| Outcome: | The paper summarises the findings of the first pan-European study on barriers to language data sharing . the findings highlight the barriers and recommend solutions to overcome them . |
European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management (L18-1)
Copied to clipboard
Andrea Lösch, Valérie Mapelli, Stelios Piperidis, Andrejs Vasiļjevs, Lilli Smal, Thierry Declerck, Eileen Schnur, Khalid Choukri, Josef van Genabith
| Challenge: | European Language Resource Coordination (ELRC) initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
| Approach: | They propose to initiate actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
| Outcome: | The European Language Resource Coordination (ELRC) consortium initiated a number of actions to support the collection of Language Resources (LRs) within the public sector in EU member and CEF-affiliated countries. |
Comparing Pretrained Multilingual Word Embeddings on an Ontology Alignment Task (L18-1)
Copied to clipboard
| Challenge: | Existing word embeddings capture a string's semantics and can be trained for multiple languages. |
| Approach: | They propose to compare three different multilingual pretrained word embedding repositories with a string-matching baseline and use it to compute semantic similarities of strings in different languages. |
| Outcome: | The proposed method produces correct alignments on a non-standard dataset on all four languages. |
An Integrated Formal Representation for Terminological and Lexical Data included in Classification Schemes (L18-1)
Copied to clipboard
| Challenge: | e-lexicography is a field of study dealing with the automated creation of specialized multilingual dictionaries from structured data. |
| Approach: | They propose to use a SKOS-XL vocabulary for modelling the multilingual terminological part of comparable taxonomies and OntoLex-Lemon for modelling multilingual lexical entries. |
| Outcome: | The proposed model can be explicitly cross-linked in the context of the Linguistic Linked Open Data (LLOD). |