Papers by Manfred Stede

18 papers
Adapting Coreference Resolution to Twitter Conversations (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on coreference resolution for Twitter texts show that performance is low.
Approach: They propose to use Twitter conversations to train a system that is originally trained on OntoNotes to improve coreference resolution.
Outcome: The proposed system outperforms existing systems on Twitter by 21.6%.
Shallow Discourse Parsing for Under-Resourced Languages: Combining Machine Translation and Annotation Projection (2020.lrec-1)

Copied to clipboard

Challenge: Shallow Discourse Parsing (SDP) relies on large amounts of training data, which so far exists only for English.
Approach: They propose to translate an existing English Penn Discourse TreeBank into German and use it to create a German corpus annotated for shallow discourse relations in the news domain.
Outcome: The proposed corpus is annotated for shallow discourse relations in the (financial) news domain.
Semi-Supervised Tri-Training for Explicit Discourse Argument Expansion (2020.lrec-1)

Copied to clipboard

Challenge: a novel application of semi-supervision for shallow discourse parsing is described . we focus on explicit discourse arguments, but we leave the sense selection aside .
Approach: They propose a semi-supervised approach for shallow discourse parsing using sequence tagging.
Outcome: The proposed approach improves performance by 2-10% in the first setting and by comparing the results with training relations.
Variation in Coreference Strategies across Genres and Production Media (2020.coling-main)

Copied to clipboard

Challenge: a lack of work on automatic coreference resolution on spoken and written language has led to inconclusive results.
Approach: They propose to use Ontonotes, Switchboard and Twitter to investigate coreference . they find fairly clear patterns of "behavior" for the different genres/medias .
Outcome: The results show that the choice of genre and the medium (spoken versus spoken) relates to the spokenwritten spectrum for coreference strategies.
Argument Similarity Assessment in German for Intelligent Tutoring: Crowdsourced Dataset and First Experiments (2022.lrec-1)

Copied to clipboard

Challenge: Recent years have seen increasing interest in applying natural language processing (NLP) applications to the field of education.
Approach: They propose an NLP-based system that supports german secondary school students in an argumentative writing exercise.
Outcome: The proposed system will support students in a German school exercise . the system will assess similarity between arguments in snippets of argumentative text .
How Diplomats Dispute: The UN Security Council Conflict Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Until now, there has been little work on how to formalize conflicts in a diplomatic setting.
Approach: They present a corpus of 87 UNSC speeches that are annotated for conflicts and demonstrate the difficulty when dealing with diplomatic language.
Outcome: The proposed method demonstrates that diplomatic language is complex and often implicit along various dimensions.
Automated Cross-language Intelligibility Analysis of Parkinson’s Disease Patients Using Speech Recognition Technologies (P19-2)

Copied to clipboard

Challenge: PD is the second most common neurodegenerative disorder after Alzheimers disease . speech impairments are one of the earliest manifestations in PD patients .
Approach: They propose to analyze the speech signals of PD patients and healthy control subjects in three different languages: German, Spanish, and Czech.
Outcome: The proposed model can discriminate between PD patients and HC subjects even when the language used for train and test is different.
DiMLex-Bangla: A Lexicon of Bangla Discourse Connectives (2020.lrec-1)

Copied to clipboard

Challenge: Discourse connectives are widely believed to be the most explicit, prototypical and most reliable relational signals in discourse processing.
Approach: They present a newly developed lexicon of Bangla discourse connectives . it contains 123 Bangla connective entries, which are primarily compiled from literature .
Outcome: The lexicon contains 123 Bangla connective entries, which are compiled from the linguistic literature and translation of English discourse connectives.
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)

Copied to clipboard

Challenge: Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data.
Approach: They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes.
Outcome: The proposed corpus includes 112 argumentative microtexts and a new annotation tool.
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels.
Approach: They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus.
Outcome: The proposed corpus is more usable for shallow discourse parsing.
Towards Identifying Alternative-Lexicalization Signals of Discourse Relations (2022.coling-1)

Copied to clipboard

Challenge: Existing shallow discourse parsing methods have been limited to identifying relations signaled by a discourse connective and those without a signal.
Approach: They propose to identify relations signalled by a discourse connective and those without . they compare a pattern-based approach and a sequence labeling model .
Outcome: The proposed approach is based on a pattern-based approach and a sequence labeling model.
GerCCT: An Annotated Corpus for Mining Arguments in German Tweets on Climate Change (2022.lrec-1)

Copied to clipboard

Challenge: Recent work on annotated resources focused on single argument components, i.e., claim or evidence.
Approach: They propose to annotate a German climate change argument corpus using sarcasm and toxic language to facilitate filtering out non-argumentative content.
Outcome: The proposed corpus is the first to be annotated for argumentation, sarcasm and toxic language.
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)

Copied to clipboard

Challenge: lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels.
Approach: They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense.
Outcome: The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format.
Discourse Sense Flows: Modelling the Rhetorical Style of Documents across Various Domains (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent work on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and alternative lexicalizations.
Approach: They propose a model for extracting and classifying discourse relation signals from the Penn Discourse Treebank v3 corpus and introduce a new way of modeling rhetorical style by the linear order of coherence relations.
Outcome: The proposed models are based on the Penn Discourse Treebank v3 corpus and employ n-gram patterns to predict genre/domain discrimination.
Developing the Bangla RST Discourse Treebank (L18-1)

Copied to clipboard

Challenge: a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla .
Approach: They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus .
Outcome: The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications .
Extractive Summarisation for German-language Data: A Text-level Approach with Discourse Features (2022.coling-1)

Copied to clipboard

Challenge: Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature.
Approach: They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models.
Outcome: The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models.
Argumentation Synthesis following Rhetorical Strategies (C18-1)

Copied to clipboard

Challenge: Existing argument mining studies focus on logical structure of arguments, identifying their units and relations, and the effects of logical and emotional arguments across audiences.
Approach: They propose to use rhetorical strategies to synthesize argumentative texts with different strategies.
Outcome: The proposed model shows that the experts agree significantly more on selection when following the same strategy.
Exploiting a lexical resource for discourse connective disambiguation in German (2020.coling-main)

Copied to clipboard

Challenge: a connective lexicon can be a valuable resource for languages with limited PDTB-style annotations . connectives are usually understood to be ambiguous in two different ways .
Approach: They propose to augment a purely-empirical approach to connective identification and sense classification in German . they find that a connective lexicon can be a valuable resource for those languages with a large PDTB-style-annotated coprus .
Outcome: The proposed approach improves on published results and achieves an F1 score for German sense classification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations