Papers by Thomas Demeester

13 papers
BioDEX: Large-Scale Biomedical Adverse Drug Event Extraction for Real-World Pharmacovigilance (2023.findings-emnlp)

Copied to clipboard

Challenge: pharmacovigilance (PV) is a tool for analyzing adverse drug events from biomedical literature . pharmacologists use natural language processing to extract core information from papers .
Approach: They propose a resource for biomedical adverse drug event eXtraction using natural language processing.
Outcome: The proposed model achieves 59.1% F1 (validation) and estimates human performance to be 72.0% F1 . the proposed model could be used to improve drug safety monitoring, also called pharmacovigilance, in the future.
Recipe Instruction Semantics Corpus (RISeC): Resolving Semantic Structure and Zero Anaphora in Recipes (2020.aacl-main)

Copied to clipboard

Challenge: Existing approaches to understanding recipe instructions make assumptions that are domain specific.
Approach: They propose a new dataset for information extraction on recipes . they avoid a priori pre-defining domain-specific predicates to recognize . instead, they focus on basic understanding of the expressed semantics .
Outcome: The proposed dataset avoids a priori pre-defining domain-specific predicates to recognize . instead, it focuses on basic understanding of the expressed semantics rather than reducing them to a simplified state representation.
Injecting Knowledge Base Information into End-to-End Joint Entity and Relation Extraction and Coreference Resolution (2021.findings-acl)

Copied to clipboard

Challenge: Using unsupervised entity linking, we solve named entity recognition, coreference resolution and relation extraction tasks together.
Approach: They propose to use a knowledge base to inject information into a joint IE model by using unsupervised entity linking.
Outcome: The proposed model improves on two datasets with 5% F1 score.
Robustifying Sentiment Classification by Maximally Exploiting Few Counterfactuals (2022.emnlp-main)

Copied to clipboard

Challenge: a recent study found that finetuned language models rely on spurious patterns in training data . this limitation limits their performance on out-of-distribution (OOD) test data.
Approach: They propose a method that only requires annotation of a small fraction of training data . they add 1% manual counterfactuals to training data and generate extra counterfacts in vector space .
Outcome: The proposed approach improves sentiment classification using IMDb data and other sets for OOD tests.
Adversarial training for multi-context joint entity and relation extraction (D18-1)

Copied to clipboard

Challenge: Existing models that use adversarial training (AT) have been used in various tasks such as parsing, POS tagging, relation extraction and translation.
Approach: They propose to use adversarial training (AT) to regularize neural network methods by adding small perturbations to the input data.
Outcome: The proposed model improves state-of-the-art on news, biomedical, and real estate datasets and for different languages.
Jack the Reader – A Machine Reading Framework (P18-4)

Copied to clipboard

Challenge: Many Machine Reading and Natural Language Understanding tasks require reading supporting text in order to answer questions.
Approach: They propose a framework for Machine Reading that allows for quick prototyping by component reuse and evaluation of new models on existing datasets.
Outcome: The proposed framework supports question answering, natural language inference and link prediction tasks.
Diverse Content Selection for Educational Question Generation (2023.eacl-srw)

Copied to clipboard

Challenge: Current automatic Question Generation (QG) systems do not consider content selection as an educational aspect.
Approach: They propose to select content based on relevance and topic diversity for question generation on educational document level.
Outcome: The proposed solution reduces the time and effort required to create questions for students on educational datasets.
A Simple Geometric Method for Cross-Lingual Linguistic Transformations with Pre-trained Autoencoders (2021.emnlp-main)

Copied to clipboard

Challenge: Existing studies have used probing tasks to verify the presence of linguistic properties in vector representations, but it is unclear whether they can be manipulated to indirectly steer them.
Approach: They validate a geometric mapping technique to transform linguistic properties without tuning . they use a pre-trained multilingual autoencoder to transform three linguistic property .
Outcome: The proposed method can be used without tuning of the pre-trained autoencoder . the results are validated in monolingual and cross-lingual settings .
Explaining Character-Aware Neural Networks for Word-Level Prediction: Do They Discover Linguistic Rules? (D18-1)

Copied to clipboard

Challenge: Character-level features are used in many natural language processing algorithms but little is known about the character-level patterns they learn.
Approach: They extend contextual decomposition technique to convolutional neural networks and bidirectional long-term memory networks to evaluate and compare these models for morphological tagging on three morphology-dependent languages.
Outcome: The proposed models implicitly discover understandable linguistic rules for morphological tagging on three morphology-dependent languages.
A Million Tweets Are Worth a Few Points: Tuning Transformers for Customer Service Tasks (2021.naacl-main)

Copied to clipboard

Challenge: In domain-specific customer service applications, many companies struggle to deploy advanced NLP models due to the limited availability of and noise in their datasets.
Approach: They analyze customer service conversations on a multilingual social media corpus and compare different approaches to pretraining and finetuning on different end tasks.
Outcome: The proposed model improves performance on multilingual social media data, especially in non-English settings.
Towards Consistent Document-level Entity Linking: Joint Models for Entity Linking and Coreference Resolution (2022.acl-short)

Copied to clipboard

Challenge: Existing approaches to solve entity linking (EL) jointly with coreference resolution (coref) a coreferenced cluster can only be linked to a single entity or NIL (i.e., a nonlinkable entity)
Approach: They propose to join entity linking and coreference resolution in a single structured prediction task over directed trees and use a globally normalized model to solve it.
Outcome: The proposed model improves on two datasets with a +5% boost in accuracy compared to standalone models . the proposed model is based on current models that predict a single antecedent for each span to resolve .
BioLORD: Learning Ontological Representations from Definitions for Biomedical Concepts and their Textual Descriptions (2022.findings-emnlp)

Copied to clipboard

Challenge: BioLORD is a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts.
Approach: They propose a pre-training strategy for producing meaningful representations for clinical sentences and biomedical concepts using definitions and ontologies.
Outcome: The proposed model produces more semantic representations that match more closely the hierarchical structure of ontologies.
Sub-event detection from twitter streams as a sequence labeling problem (N19-1)

Copied to clipboard

Challenge: Existing methods for sub-event detection do not account for sequential nature of social media streams.
Approach: They propose to use a neural sequence architecture that explicitly accounts for the chronological order of posts to improve sub-event detection.
Outcome: The proposed method outperforms a graph-based state-of-the-art method for binary sub-event detection (2.7% micro-F1 improvement) it also outperformed a recurrent neural network model on the posts sequence level for labeled sub- events (2.4% bin-level improvement).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations