Papers by Paul Buitelaar
Evaluation Dataset and Methodology for Extracting Application-Specific Taxonomies from the Wikipedia Knowledge Graph (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent efforts to extract hierarchical relations from unstructured text have been challenging. |
| Approach: | They propose an iterative method to extract an application-specific gold standard dataset from a Wikipedia knowledge graph and an evaluation framework to assess the quality of noisy automatically extracted taxonomies. |
| Outcome: | The proposed method reduces manual work and provides a first gold standard dataset and evaluation framework. |
Linghub2: Language Resource Discovery Tool for Language Technologies (2022.lrec-1)
Copied to clipboard
| Challenge: | Linghub is a platform for language resources that can be used to find and retrieve data . the platform is based on a popular open source data management system, DSpace . |
| Approach: | This work describes a rejuvenation and modernisation of the 2015 platform into using a popular open source data management system, DSpace, as foundation. |
| Outcome: | Linghub2 1 aims to help language resources and technology users find and retrieve relevant data . the new platform, Ling hub2, contains updated and extended resources and more languages offered . |
From Laughter to Inequality: Annotated Dataset for Misogyny Detection in Tamil and Malayalam Memes (2024.lrec-main)
Copied to clipboard
Rahul Ponnusamy, Kathiravan Pannerselvam, Saranya R, Prasanna Kumar Kumaresan, Sajeetha Thavareesan, Bhuvaneswari S, Anshid K.a, Susminu S Kumar, Paul Buitelaar, Bharathi Raja Chakravarthi
| Challenge: | a new form of memes has emerged to combat misogyny and harmful stereotypes . authors present a dataset to analyze online misogamy in Tamil and Malayalam communities . |
| Approach: | They propose to create an annotated dataset with detailed annotation guidelines to analyze online misogyny within Tamil and Malayalam-speaking communities. |
| Outcome: | The proposed dataset reveals the world of gender bias and stereotypes in Tamil and Malayalam-speaking communities. |
Figure Me Out: A Gold Standard Dataset for Metaphor Interpretation (2020.lrec-1)
Copied to clipboard
| Challenge: | Metaphor comprehension and understanding is a complex cognitive task that requires interpreting metaphors by grasping the interaction between the meaning of their target and source concepts. |
| Approach: | They propose an automatic retrieval approach to annotate verb-noun metaphors in text . they validated their approach by annotating around 1,500 metaphors from tweets . |
| Outcome: | The proposed method reduces the workload on annotators and maintains consistency . it can be used to interpret verb-noun metaphoric expressions in tweets . |
Inference to the Best Explanation in Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have found success in real-world applications, but their underlying explanatory process is still poorly understood. |
| Approach: | They propose to use a framework inspired by philosophical accounts on Inference to the Best Explanation (IBE) to advance the interpretation and evaluation of LLMs’ explanations. |
| Outcome: | The proposed framework can identify the best explanation with up to 77% accuracy (27% above random) while being intrinsically more efficient and interpretable. |
Teanga: A Linked Data based platform for Natural Language Processing (L18-1)
Copied to clipboard
| Challenge: | Using linked data, we can use many NLP services from a single interface . integrating components within a development model is endemic to software development . |
| Approach: | They propose a linked data based platform for natural language processing that uses linked data to define the types of services input and output. |
| Outcome: | The proposed platform is easy to install and run, easy to use and able to run multiple NLP tasks from one interface. |
Dataset for Identification of Homophobia and Transphobia for Telugu, Kannada, and Gujarati (2024.lrec-main)
Copied to clipboard
| Challenge: | There has been a rise in homophobic and transphobic content targeting LGBT+ individuals on social media platforms. |
| Approach: | They propose to use a dataset to automatically identify homophobic and transphobic content within comments collected from YouTube for three languages. |
| Outcome: | The proposed dataset will identify homophobic and transphobic content within comments collected from YouTube in Telugu, Kannada, and Gujarati. |
Contextual Modulation for Relation-Level Metaphor Identification (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to identifying metaphors in text ignore context where metaphor occurs . existing approaches focus on word-level identification without explicitly modelling interaction between metaphor components . |
| Approach: | They propose a method for identifying relation-level metaphoric expressions of certain grammatical relations based on contextual modulation. |
| Outcome: | The proposed architecture achieves state-of-the-art results on benchmark datasets. |
LUCE: A Dynamic Framework and Interactive Dashboard for Opinionated Text Analysis (2025.coling-demos)
Copied to clipboard
| Challenge: | LUCE is an advanced dynamic framework for analysing opinionated text . it features computational modules for different elements of opinions, e.g., sentiment/emotion, suggestion, figurative language, hate/toxic speech, and topics. |
| Approach: | They introduce a dynamic framework with an interactive dashboard for analysing opinionated text . it features computational modules of text classification and extraction for different elements of opinions . |
| Outcome: | The framework is validated in a relevant environment and its capabilities and performance demonstrated . it features trained models, python-based APIs, and a user-friendly dashboard . |
A supervised approach to taxonomy extraction using word embeddings (L18-1)
Copied to clipboard
| Challenge: | a recent evaluation of a method for organizing texts into a hierarchy showed that it did not outperform a baseline. |
| Approach: | They propose a method that uses supervised learning to combine multiple features with a support vector machine classifier including the baseline features. |
| Outcome: | The proposed method outperforms the baseline method and provides stronger method for identifying taxonomic relations than previous methods. |
Analysing the Correlation between Lexical Ambiguity and Translation Quality in a Multimodal Setting using WordNet (2022.naacl-srw)
Copied to clipboard
| Challenge: | Recent studies in machine translation have been focusing on using visual information to improve the translation quality of sentences. |
| Approach: | They propose to use visual information to improve the output quality of a text-based translation model by extracting ambiguity scores from WordNet. |
| Outcome: | The proposed model improves translation quality for all sentences in the English-German dataset. |
Automatic Enrichment of Terminological Resources: the IATE RDF Example (L18-1)
Copied to clipboard
| Challenge: | a recent paper aims to automate the maintenance of terminological resources. |
| Approach: | They propose automatic approaches to maintain and increase lexical coverage of knowledge bases by using machine translation and multilingual word sense disambiguation. |
| Outcome: | The proposed approach outperforms the existing methods with random sentences in most languages . |
A Hybrid Approach to Aspect Based Sentiment Analysis Using Transfer Learning (2024.lrec-main)
Copied to clipboard
| Challenge: | Aspect-Based Sentiment Analysis (ABSA) aims to identify terms or multiword expressions (MWEs) on which sentiments are expressed and the sentiment polarities associated with them. |
| Approach: | They propose a hybrid approach to Aspect-Based Sentiment Analysis using transfer learning . they exploit the strengths of large language models and traditional syntactic dependencies . |
| Outcome: | The proposed method exploits the strengths of large language models and traditional syntactic dependencies. |
CALM-Bench: A Multi-task Benchmark for Evaluating Causality-Aware Language Models (2023.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances in foundation language models have shown the efficacy of pre-trained models across diverse QA tasks. |
| Approach: | They propose a multi-task benchmark for evaluating causality-aware language models to unify causal QA research. |
| Outcome: | The proposed model outperforms single-task fine-tuned models on the CALM-Bench tasks. |
A Term Extraction Approach to Survey Analysis in Health Care (2020.lrec-1)
Copied to clipboard
| Challenge: | a new study examines the impact of customer feedback on health care organizations . the results of the 2017 Irish National Inpatient Survey are compared to a manual framework . |
| Approach: | They propose an approach to patient experience using free text questions from the 2017 Irish National Inpatient Survey campaign. |
| Outcome: | The proposed approach to patient experience is based on the results of the 2017 Irish National Inpatient Survey. |
A Comparison Of Emotion Annotation Schemes And A New Annotated Data Set (L18-1)
Copied to clipboard
| Challenge: | a series of study on positive/negative sentiments has been conducted on tweets, but recognition of more nuanced affect has received little attention . valence, arousal, dominance and surprise are the most commonly used emotion representation schemes . |
| Approach: | They propose to annotate tweets with scores on four emotion dimensions . they compare annotator agreement with relative annotation schemes over categorical ones . |
| Outcome: | The proposed model improves agreement with relative annotation schemes over categorical ones on Ekman's six basic emotions. |