Developing the Bangla RST Discourse Treebank (L18-1)

Copied to clipboard

Challenge: a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla .
Approach: They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus .
Outcome: The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications .

Similar Papers

Developing a Rhetorical Structure Theory Treebank for Czech (2024.lrec-main)

Copied to clipboard

Challenge: a paper on the Czech RST Discourse Treebank is the first version of a textual annotation system based on the Rhetorical Structure Theory . document is annotated using the RST, a global coherence model proposed by Mann and Thompson .
Approach: They introduce the first version of the Czech RST Discourse Treebank . paper presents an annotation process and provides corpus statistics and evaluation .
Outcome: The paper presents the first version of the Czech RST Discourse Treebank . the treebank includes two gold annotations representing divergent interpretations .
DiMLex-Bangla: A Lexicon of Bangla Discourse Connectives (2020.lrec-1)

Copied to clipboard

Challenge: Discourse connectives are widely believed to be the most explicit, prototypical and most reliable relational signals in discourse processing.
Approach: They present a newly developed lexicon of Bangla discourse connectives . it contains 123 Bangla connective entries, which are primarily compiled from literature .
Outcome: The lexicon contains 123 Bangla connective entries, which are compiled from the linguistic literature and translation of English discourse connectives.
MEGA RST Discourse Treebanks with Structure and Nuclearity from Scalable Distant Sentiment Supervision (2020.emnlp-main)

Copied to clipboard

Challenge: Existing discourse treebanks are limited in the application of data-driven approaches to discourse parsing.
Approach: They propose a method to automatically generate discourse treebanks using distant supervision from sentiment annotated datasets by heuristic beam-search strategy extended with a stochastic component.
Outcome: The proposed method generates discourse trees incorporating structure and nuclearity for documents of arbitrary length using an efficient beam-search strategy, extended with a stochastic component.
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)

Copied to clipboard

Challenge: Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels.
Approach: They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus.
Outcome: The proposed corpus is more usable for shallow discourse parsing.
Gold Standard Bangla OCR Dataset: An In-Depth Look at Data Preprocessing and Annotation Processes (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing datasets designed specifically for the Bengali language have been limited.
Approach: They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR.
Outcome: The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images.
Neural RST-based Evaluation of Discourse Coherence (2020.aacl-main)

Copied to clipboard

Challenge: Existing discourse parsers cannot predict coherent texts without using silver-standard features.
Approach: They propose a tree-recursive neural model which takes advantage of the text’s RST features produced by a state of the art RST parser and compares it to the current state of art.
Outcome: The proposed model achieves state-of-the-art accuracy on the Grammarly Corpus for Discourse Coherence (GCDC) and has 62% fewer parameters than existing models.
Persian Discourse Treebank and coreference corpus (L18-1)

Copied to clipboard

Challenge: Currently, we are adding a new document-level discourse annotation to our new corpus.
Approach: They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution.
Outcome: The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens.
SHONGLAP: A Large Bengali Open-Domain Dialogue Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing open-domain dialogue systems suffer from data scarcity due to unavailability of high-quality datasets for low-resource languages like Bengali.
Approach: They propose to prepare large-scale open-domain dialogue datasets from podcasts and talk-shows and label them based on weak-supervision techniques.
Outcome: The proposed corpus improves performance of large language models in case of downstream classification tasks during fine-tuning.
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)

Copied to clipboard

Challenge: CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses)
Approach: They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic.
Outcome: The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses)
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task.
Approach: They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank.
Outcome: The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations