Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.

Similar Papers

TED-CDB: A Large-Scale Chinese Discourse Relation Dataset on TED Talks (2020.emnlp-main)

Copied to clipboard

Challenge: TED-CDB dataset is a unique corpus of spoken discourse in Chinese . TED is based on the concept that discourse relations are grounded in an identifiable set of discourse connectives or Altlex expressions.
Approach: They have created a dataset that annotates TED talks in Chinese . they propose to adapt the dataset to Chinese news text to improve its performance .
Outcome: The TED-CDB dataset can improve the performance of systems for languages other than Chinese . it is adapted to features that are not present in English and can extract discourse semantic features .
Shallow Discourse Annotation for Chinese TED Talks (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to annotate text with discourse properties are limited to newspaper articles and are not available in Chinese.
Approach: They propose to annotate TED talks with Chinese-related properties using the Penn Discourse TreeBank annotation style . they propose to use planned monologues instead of written text to annnotate Chinese-specific properties.
Outcome: The proposed method is able to achieve reliable results in Chinese spoken monologues, and is based on the Penn Discourse TreeBank annotation style.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
ISO-based Annotated Multilingual Parallel Corpus for Discourse Markers (2022.lrec-1)

Copied to clipboard

Challenge: Discourse markers carry information about the discourse structure and organization, and also signal local dependencies or epistemic stance of speaker.
Approach: They propose an ISO-based annotated multilingual parallel corpus for discourse markers . they propose an annotation scheme for discourse relations with a plug-in to ISO 24617-2 .
Outcome: The proposed language resource is based on an ISO-based annotated multilingual parallel corpus of discourse markers.
WikiBank: Using Wikidata to Improve Multilingual Frame-Semantic Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Frame-semantic annotations exist for a tiny fraction of the world’s languages, however, Wikidata provides a common, distant supervision signal for semantic parsers.
Approach: They propose a multilingual resource with partial semantic dependency structures that can be used to extend pre-existing resources rather than creating new man-made resources from scratch.
Outcome: The proposed resource can be used to augment pre-existing resources or reduce the annotation effort for low-resource languages.
An Empirical Evaluation of Annotation Practices in Corpora from Language Documentation (2020.lrec-1)

Copied to clipboard

Challenge: Language documentation projects have produced substantial amounts of primary data from a wide variety of endangered languages.
Approach: They propose to use common annotation conventions in existing corpora to facilitate their future processing.
Outcome: The proposed formats are based on the common ELAN and Toolbox formats and are used to facilitate their future processing.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)

Copied to clipboard

Challenge: Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data.
Approach: They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes.
Outcome: The proposed corpus includes 112 argumentative microtexts and a new annotation tool.
From Unaligned to Aligned: Scaling Multilingual LLMs with Multi-Way Parallel Corpora (2025.emnlp-main)

Copied to clipboard

Challenge: Experiments show that models trained on multi-way parallel data outperform those trained on unaligned data.
Approach: They propose a large-scale, high-quality multi-way parallel corpus based on TED Talks that spans 113 languages with up to 50 languages aligned in parallel.
Outcome: The proposed model outperforms models trained on unaligned multilingual data on six multilingual benchmarks.
Towards Unification of Discourse Annotation Frameworks (2022.acl-srw)

Copied to clipboard

Challenge: Discourse information is difficult to represent and annotate, and corpora annotated under different frameworks vary considerably.
Approach: They propose to use automatic means to unify discourse structures and relations . they will also explore the application of the unified framework in multi-task learning and graphical models .
Outcome: The proposed method can be used in multi-task learning and graphical models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations