Papers by Udo Kruschwitz

10 papers
Scalable Visualisation of Sentiment and Stance (L18-1)

Copied to clipboard

Challenge: a novel visualisation approach for sentiment and stance analysis is proposed for large datasets.
Approach: They propose a visualisation approach for scalable visualisation of sentiment and stance from large-scale data.
Outcome: The proposed visualisation approach can be used on a 9,278 user comments with stance explicitly declared by the author.
Crowdsourcing and Aggregating Nested Markable Annotations (P19-1)

Copied to clipboard

Challenge: Existing methods for identifying markables for coreference annotation are task and language-independent and can be used for a variety of other annotation tasks.
Approach: They propose a method for identifying markables for coreference annotation that combines automatic markable detectors with checking with a Game-With-A-Purpose (GWAP) and aggregation using a Bayesian annotation model.
Outcome: The proposed method improves mention boundaries on news and other genres by over seven percentage points compared with state-of-the-art, domain-independent automatic mention detectors and almost three points over an in-domain mention detector.
Challenges in Pre-Training Graph Neural Networks for Context-Based Fake News Detection: An Evaluation of Current Strategies and Resource Limitations (2024.lrec-main)

Copied to clipboard

Challenge: Graph Neural Networks (GNNs) are used to train neural networks to detect fake news based on context-based methods.
Approach: They propose to combine the two by applying pre-training of Graph Neural Networks (GNNs) in the domain of context-based fake news detection.
Outcome: The proposed methods show that transfer learning does not lead to significant improvements over training a model from scratch in the domain of context-based fake news detection.
A Probabilistic Annotation Model for Crowdsourcing Coreference (D18-1)

Copied to clipboard

Challenge: Existing methods to generate annotated corpora for coreference are expensive and limited.
Approach: They propose a model of annotation for aggregating crowdsourced anaphoric annotations.
Outcome: The proposed model can extract from crowdsourced annotations coreference chains comparable to those obtained with expert annotation.
Aggregating Crowdsourced and Automatic Judgments to Scale Up a Corpus of Anaphoric Reference for Fiction and Wikipedia Texts (2023.eacl-main)

Copied to clipboard

Challenge: Existing approaches to scale up anaphoric annotation have not overcome these limitations.
Approach: They propose to use a game-with-a-purpose to ‘complete’ markable annotations by using an anaphoric resolver and an aggregation method for anaphorism.
Outcome: The proposed method could be adopted to greatly speed up annotation time in other projects involving games-with-a-purpose.
Improving Hate Speech Detection with Deep Learning Ensembles (L18-1)

Copied to clipboard

Challenge: censorship is a potential risk when addressing these issues with automated text classification methods.
Approach: They propose to use a neural network-based ensemble method to better classify hate speech using a publicly available embedding model and a popular sentiment dataset.
Outcome: The proposed method improves by 5 points on a hate speech corpus from Twitter and a popular sentiment dataset.
Tackling Irony Detection using Ensemble Classifiers (2022.lrec-1)

Copied to clipboard

Challenge: Automated approaches to irony detection still fall short of what one would consider desirable performance.
Approach: They propose to use transformer-based approaches to automate irony detection in social media . they propose to augmentation training data to address the binary and fine-grained problem .
Outcome: The proposed methods improve performance over baselines and are not decisive for good results.
Applying Automatic Text Summarization for Fake News Detection (2022.lrec-1)

Copied to clipboard

Challenge: Social media has been a driver for the spread of misleading and deliberately wrong information, as there is little to no veracity monitoring.
Approach: They propose a framework that combines transformer-based language models with contextual information to circumvent sequential limits and related loss of information.
Outcome: The proposed framework can circumvent sequential limits and related loss of information on two publicly available datasets and achieve state-of-the-art performance benchmarks.
A Crowdsourced Corpus of Multiple Judgments and Disagreement on Anaphoric Interpretation (N19-1)

Copied to clipboard

Challenge: a corpus of anaphoric information (coreference) is crowdsourced through a game-with-a-purpose . its main feature is the large number of judgments per markable: 20 on average, and over 2.2M in total.
Approach: They propose to crowdsource anaphoric information corpus by a game-with-a-purpose and to use it to train a coreference resolver.
Outcome: The proposed corpus contains annotations for 108,000 markables and 20 judgments per markable, and 2.2M in total.
A New Dataset for Topic-Based Paragraph Classification in Genocide-Related Court Transcripts (2022.lrec-1)

Copied to clipboard

Challenge: Recent advances in natural language processing have lowered the barriers for people outside the NLP community to tap into the tools and resources applied to a variety of domain-specific applications.
Approach: They propose to annotate court transcripts from genocide-related cases using transformer-based approaches and to establish benchmarks for the task of paragraph identification of violence-related witness statements.
Outcome: The first annotated corpus of genocide-related court transcripts is aimed at providing a first reference corpus for the community and to establish benchmark performances using state-of-the-art transformer-based approaches.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations