Papers by Manfred Stede
Adapting Coreference Resolution to Twitter Conversations (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on coreference resolution for Twitter texts show that performance is low. |
| Approach: | They propose to use Twitter conversations to train a system that is originally trained on OntoNotes to improve coreference resolution. |
| Outcome: | The proposed system outperforms existing systems on Twitter by 21.6%. |
Shallow Discourse Parsing for Under-Resourced Languages: Combining Machine Translation and Annotation Projection (2020.lrec-1)
Copied to clipboard
| Challenge: | Shallow Discourse Parsing (SDP) relies on large amounts of training data, which so far exists only for English. |
| Approach: | They propose to translate an existing English Penn Discourse TreeBank into German and use it to create a German corpus annotated for shallow discourse relations in the news domain. |
| Outcome: | The proposed corpus is annotated for shallow discourse relations in the (financial) news domain. |
Semi-Supervised Tri-Training for Explicit Discourse Argument Expansion (2020.lrec-1)
Copied to clipboard
| Challenge: | a novel application of semi-supervision for shallow discourse parsing is described . we focus on explicit discourse arguments, but we leave the sense selection aside . |
| Approach: | They propose a semi-supervised approach for shallow discourse parsing using sequence tagging. |
| Outcome: | The proposed approach improves performance by 2-10% in the first setting and by comparing the results with training relations. |
Variation in Coreference Strategies across Genres and Production Media (2020.coling-main)
Copied to clipboard
| Challenge: | a lack of work on automatic coreference resolution on spoken and written language has led to inconclusive results. |
| Approach: | They propose to use Ontonotes, Switchboard and Twitter to investigate coreference . they find fairly clear patterns of "behavior" for the different genres/medias . |
| Outcome: | The results show that the choice of genre and the medium (spoken versus spoken) relates to the spokenwritten spectrum for coreference strategies. |
Argument Similarity Assessment in German for Intelligent Tutoring: Crowdsourced Dataset and First Experiments (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent years have seen increasing interest in applying natural language processing (NLP) applications to the field of education. |
| Approach: | They propose an NLP-based system that supports german secondary school students in an argumentative writing exercise. |
| Outcome: | The proposed system will support students in a German school exercise . the system will assess similarity between arguments in snippets of argumentative text . |
How Diplomats Dispute: The UN Security Council Conflict Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Until now, there has been little work on how to formalize conflicts in a diplomatic setting. |
| Approach: | They present a corpus of 87 UNSC speeches that are annotated for conflicts and demonstrate the difficulty when dealing with diplomatic language. |
| Outcome: | The proposed method demonstrates that diplomatic language is complex and often implicit along various dimensions. |
Automated Cross-language Intelligibility Analysis of Parkinson’s Disease Patients Using Speech Recognition Technologies (P19-2)
Copied to clipboard
| Challenge: | PD is the second most common neurodegenerative disorder after Alzheimers disease . speech impairments are one of the earliest manifestations in PD patients . |
| Approach: | They propose to analyze the speech signals of PD patients and healthy control subjects in three different languages: German, Spanish, and Czech. |
| Outcome: | The proposed model can discriminate between PD patients and HC subjects even when the language used for train and test is different. |
DiMLex-Bangla: A Lexicon of Bangla Discourse Connectives (2020.lrec-1)
Copied to clipboard
| Challenge: | Discourse connectives are widely believed to be the most explicit, prototypical and most reliable relational signals in discourse processing. |
| Approach: | They present a newly developed lexicon of Bangla discourse connectives . it contains 123 Bangla connective entries, which are primarily compiled from literature . |
| Outcome: | The lexicon contains 123 Bangla connective entries, which are compiled from the linguistic literature and translation of English discourse connectives. |
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)
Copied to clipboard
| Challenge: | Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data. |
| Approach: | They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes. |
| Outcome: | The proposed corpus includes 112 argumentative microtexts and a new annotation tool. |
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)
Copied to clipboard
| Challenge: | Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels. |
| Approach: | They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus. |
| Outcome: | The proposed corpus is more usable for shallow discourse parsing. |
Towards Identifying Alternative-Lexicalization Signals of Discourse Relations (2022.coling-1)
Copied to clipboard
| Challenge: | Existing shallow discourse parsing methods have been limited to identifying relations signaled by a discourse connective and those without a signal. |
| Approach: | They propose to identify relations signalled by a discourse connective and those without . they compare a pattern-based approach and a sequence labeling model . |
| Outcome: | The proposed approach is based on a pattern-based approach and a sequence labeling model. |
GerCCT: An Annotated Corpus for Mining Arguments in German Tweets on Climate Change (2022.lrec-1)
Copied to clipboard
| Challenge: | Recent work on annotated resources focused on single argument components, i.e., claim or evidence. |
| Approach: | They propose to annotate a German climate change argument corpus using sarcasm and toxic language to facilitate filtering out non-argumentative content. |
| Outcome: | The proposed corpus is the first to be annotated for argumentation, sarcasm and toxic language. |
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)
Copied to clipboard
| Challenge: | lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels. |
| Approach: | They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense. |
| Outcome: | The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format. |
Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various Domains (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent work on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and alternative lexicalizations. |
| Approach: | They propose a model for extracting and classifying discourse relation signals from the Penn Discourse Treebank v3 corpus and introduce a new way of modeling rhetorical style by the linear order of coherence relations. |
| Outcome: | The proposed models are based on the Penn Discourse Treebank v3 corpus and employ n-gram patterns to predict genre/domain discrimination. |
Developing the Bangla RST Discourse Treebank (L18-1)
Copied to clipboard
| Challenge: | a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla . |
| Approach: | They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus . |
| Outcome: | The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications . |
Extractive Summarisation for German-language Data: A Text-level Approach with Discourse Features (2022.coling-1)
Copied to clipboard
| Challenge: | Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature. |
| Approach: | They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models. |
| Outcome: | The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models. |
Argumentation Synthesis following Rhetorical Strategies (C18-1)
Copied to clipboard
| Challenge: | Existing argument mining studies focus on logical structure of arguments, identifying their units and relations, and the effects of logical and emotional arguments across audiences. |
| Approach: | They propose to use rhetorical strategies to synthesize argumentative texts with different strategies. |
| Outcome: | The proposed model shows that the experts agree significantly more on selection when following the same strategy. |
Exploiting a lexical resource for discourse connective disambiguation in German (2020.coling-main)
Copied to clipboard
| Challenge: | a connective lexicon can be a valuable resource for languages with limited PDTB-style annotations . connectives are usually understood to be ambiguous in two different ways . |
| Approach: | They propose to augment a purely-empirical approach to connective identification and sense classification in German . they find that a connective lexicon can be a valuable resource for those languages with a large PDTB-style-annotated coprus . |
| Outcome: | The proposed approach improves on published results and achieves an F1 score for German sense classification. |