Papers by Udo Kruschwitz
Scalable Visualisation of Sentiment and Stance (L18-1)
Copied to clipboard
| Challenge: | a novel visualisation approach for sentiment and stance analysis is proposed for large datasets. |
| Approach: | They propose a visualisation approach for scalable visualisation of sentiment and stance from large-scale data. |
| Outcome: | The proposed visualisation approach can be used on a 9,278 user comments with stance explicitly declared by the author. |
Crowdsourcing and Aggregating Nested Markable Annotations (P19-1)
Copied to clipboard
| Challenge: | Existing methods for identifying markables for coreference annotation are task and language-independent and can be used for a variety of other annotation tasks. |
| Approach: | They propose a method for identifying markables for coreference annotation that combines automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model. |
| Outcome: | The proposed method improves mention boundaries on news and other genres by over seven percentage points compared with state-of-the-art, domain-independent automatic mention detectors and almost three points over an in-domain mention detector. |
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations (2024.lrec-main)
Copied to clipboard
| Challenge: | Graph Neural Networks (GNNs) are used to train neural networks to detect fake news based on context-based methods. |
| Approach: | They propose to combine the two by applying pre-training of Graph Neural Networks (GNNs) in the domain of context-based fake news detection. |
| Outcome: | The proposed methods show that transfer learning does not lead to significant improvements over training a model from scratch in the domain of context-based fake news detection. |
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)
Copied to clipboard
| Challenge: | Existing methods to generate annotated corpora for coreference are expensive and limited. |
| Approach: | They propose a model of annotation for aggregating crowdsourced anaphoric annotations. |
| Outcome: | The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation. |
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts (2023.eacl-main)
Copied to clipboard
Juntao Yu, Silviu Paun, Maris Camilleri, Paloma Garcia, Jon Chamberlain, Udo Kruschwitz, Massimo Poesio
| Challenge: | Existing approaches to scale up anaphoric annotation have not overcome these limitations. |
| Approach: | They propose to use a game-with-a-purpose to ‘complete’ markable annotations by using an anaphoric resolver and an aggregation method for anaphorism. |
| Outcome: | The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose. |
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)
Copied to clipboard
| Challenge: | censorship is a potential risk when addressing these issues with automated text classification methods. |
| Approach: | They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset. |
| Outcome: | The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset. |
Tackling Irony Detection using Ensemble Classifiers (2022.lrec-1)
Copied to clipboard
| Challenge: | Automated approaches to irony detection still fall short of what one would consider desirable performance. |
| Approach: | They propose to use transformer-based approaches to automate irony detection in social media . they propose to augmentation training data to address the binary and fine-grained problem . |
| Outcome: | The proposed methods improve performance over baselines and are not decisive for good results. |
Applying Automatic Text Summarization for Fake News Detection (2022.lrec-1)
Copied to clipboard
| Challenge: | Social media has been a driver for the spread of misleading and deliberately wrong information, as there is little to no veracity monitoring. |
| Approach: | They propose a framework that combines transformer-based language models with contextual information to circumvent sequential limits and related loss of information. |
| Outcome: | The proposed framework can circumvent sequential limits and related loss of information on two publicly available datasets and achieve state-of-the-art performance benchmarks. |
A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation (N19-1)
Copied to clipboard
| Challenge: | a corpus of anaphoric information (coreference) is crowdsourced through a game-with-a-purpose . its main feature is the large number of judgments per markable: 20 on average, and over 2.2M in total. |
| Approach: | They propose to crowdsource anaphoric information corpus by a game-with-a-purpose and to use it to train a coreference resolver. |
| Outcome: | The proposed corpus contains annotations for 108,000 markables and 20 judgments per markable, and 2.2M in total. |
A New Dataset for Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in natural language processing have lowered the barriers for people outside the NLP community to tap into the tools and resources applied to a variety of domain-specific applications. |
| Approach: | They propose to annotate court transcripts from genocide-related cases using transformer-based approaches and to establish benchmarks for the task of paragraph identification of violence-related witness statements. |
| Outcome: | The first annotated corpus of genocide-related court transcripts is aimed at providing a first reference corpus for the community and to establish benchmark performances using state-of-the-art transformer-based approaches. |