Papers by Paul Buitelaar

16 papers
Evaluation Dataset and Methodology for Extracting Application-Specific Taxonomies from the Wikipedia Knowledge Graph (2020.lrec-1)

Copied to clipboard

Challenge: Recent efforts to extract hierarchical relations from unstructured text have been challenging.
Approach: They propose an iterative method to extract an application-specific gold standard dataset from a Wikipedia knowledge graph and an evaluation framework to assess the quality of noisy automatically extracted taxonomies.
Outcome: The proposed method reduces manual work and provides a first gold standard dataset and evaluation framework.
Linghub2: Language Resource Discovery Tool for Language Technologies (2022.lrec-1)

Copied to clipboard

Challenge: Linghub is a platform for language resources that can be used to find and retrieve data . the platform is based on a popular open source data management system, DSpace .
Approach: This work describes a rejuvenation and modernisation of the 2015 platform into using a popular open source data management system, DSpace, as foundation.
Outcome: Linghub2 1 aims to help language resources and technology users find and retrieve relevant data . the new platform, Ling hub2, contains updated and extended resources and more languages offered .
From Laughter to Inequality: Annotated Dataset for Misogyny Detection in Tamil and Malayalam Memes (2024.lrec-main)

Copied to clipboard

Challenge: a new form of memes has emerged to combat misogyny and harmful stereotypes . authors present a dataset to analyze online misogamy in Tamil and Malayalam communities .
Approach: They propose to create an annotated dataset with detailed annotation guidelines to analyze online misogyny within Tamil and Malayalam-speaking communities.
Outcome: The proposed dataset reveals the world of gender bias and stereotypes in Tamil and Malayalam-speaking communities.
Figure Me Out: A Gold Standard Dataset for Metaphor Interpretation (2020.lrec-1)

Copied to clipboard

Challenge: Metaphor comprehension and understanding is a complex cognitive task that requires interpreting metaphors by grasping the interaction between the meaning of their target and source concepts.
Approach: They propose an automatic retrieval approach to annotate verb-noun metaphors in text . they validated their approach by annotating around 1,500 metaphors from tweets .
Outcome: The proposed method reduces the workload on annotators and maintains consistency . it can be used to interpret verb-noun metaphoric expressions in tweets .
Inference to the Best Explanation in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have found success in real-world applications, but their underlying explanatory process is still poorly understood.
Approach: They propose to use a framework inspired by philosophical accounts on Inference to the Best Explanation (IBE) to advance the interpretation and evaluation of LLMs’ explanations.
Outcome: The proposed framework can identify the best explanation with up to 77% accuracy (27% above random) while being intrinsically more efficient and interpretable.
Teanga: A Linked Data based platform for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: Using linked data, we can use many NLP services from a single interface . integrating components within a development model is endemic to software development .
Approach: They propose a linked data based platform for natural language processing that uses linked data to define the types of services input and output.
Outcome: The proposed platform is easy to install and run, easy to use and able to run multiple NLP tasks from one interface.
Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati (2024.lrec-main)

Copied to clipboard

Challenge: There has been a rise in homophobic and transphobic content targeting LGBT+ individuals on social media platforms.
Approach: They propose to use a dataset to automatically identify homophobic and transphobic content within comments collected from YouTube for three languages.
Outcome: The proposed dataset will identify homophobic and transphobic content within comments collected from YouTube in Telugu, Kannada, and Gujarati.
Contextual Modulation for Relation-Level Metaphor Identification (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to identifying metaphors in text ignore context where metaphor occurs . existing approaches focus on word-level identification without explicitly modelling interaction between metaphor components .
Approach: They propose a method for identifying relation-level metaphoric expressions of certain grammatical relations based on contextual modulation.
Outcome: The proposed architecture achieves state-of-the-art results on benchmark datasets.
LUCE: A Dynamic Framework and Interactive Dashboard for Opinionated Text Analysis (2025.coling-demos)

Copied to clipboard

Challenge: LUCE is an advanced dynamic framework for analysing opinionated text . it features computational modules for different elements of opinions, e.g., sentiment/emotion, suggestion, figurative language, hate/toxic speech, and topics.
Approach: They introduce a dynamic framework with an interactive dashboard for analysing opinionated text . it features computational modules of text classification and extraction for different elements of opinions .
Outcome: The framework is validated in a relevant environment and its capabilities and performance demonstrated . it features trained models, python-based APIs, and a user-friendly dashboard .
A supervised approach to taxonomy extraction using word embeddings (L18-1)

Copied to clipboard

Challenge: a recent evaluation of a method for organizing texts into a hierarchy showed that it did not outperform a baseline.
Approach: They propose a method that uses supervised learning to combine multiple features with a support vector machine classifier including the baseline features.
Outcome: The proposed method outperforms the baseline method and provides stronger method for identifying taxonomic relations than previous methods.
Analysing the Correlation between Lexical Ambiguity and Translation Quality in a Multimodal Setting using WordNet (2022.naacl-srw)

Copied to clipboard

Challenge: Recent studies in machine translation have been focusing on using visual information to improve the translation quality of sentences.
Approach: They propose to use visual information to improve the output quality of a text-based translation model by extracting ambiguity scores from WordNet.
Outcome: The proposed model improves translation quality for all sentences in the English-German dataset.
Automatic Enrichment of Terminological Resources: the IATE RDF Example (L18-1)

Copied to clipboard

Challenge: a recent paper aims to automate the maintenance of terminological resources.
Approach: They propose automatic approaches to maintain and increase lexical coverage of knowledge bases by using machine translation and multilingual word sense disambiguation.
Outcome: The proposed approach outperforms the existing methods with random sentences in most languages .
A Hybrid Approach to Aspect Based Sentiment Analysis Using Transfer Learning (2024.lrec-main)

Copied to clipboard

Challenge: Aspect-Based Sentiment Analysis (ABSA) aims to identify terms or multiword expressions (MWEs) on which sentiments are expressed and the sentiment polarities associated with them.
Approach: They propose a hybrid approach to Aspect-Based Sentiment Analysis using transfer learning . they exploit the strengths of large language models and traditional syntactic dependencies .
Outcome: The proposed method exploits the strengths of large language models and traditional syntactic dependencies.
CALM-Bench: A Multi-task Benchmark for Evaluating Causality-Aware Language Models (2023.findings-eacl)

Copied to clipboard

Challenge: Recent advances in foundation language models have shown the efficacy of pre-trained models across diverse QA tasks.
Approach: They propose a multi-task benchmark for evaluating causality-aware language models to unify causal QA research.
Outcome: The proposed model outperforms single-task fine-tuned models on the CALM-Bench tasks.
A Term Extraction Approach to Survey Analysis in Health Care (2020.lrec-1)

Copied to clipboard

Challenge: a new study examines the impact of customer feedback on health care organizations . the results of the 2017 Irish National Inpatient Survey are compared to a manual framework .
Approach: They propose an approach to patient experience using free text questions from the 2017 Irish National Inpatient Survey campaign.
Outcome: The proposed approach to patient experience is based on the results of the 2017 Irish National Inpatient Survey.
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)

Copied to clipboard

Challenge: a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes .
Approach: They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones .
Outcome: The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations