| Challenge: | a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla . |
| Approach: | They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus . |
| Outcome: | The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications . |
Similar Papers
Developing a Rhetorical Structure Theory Treebank for Czech (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper on the Czech RST Discourse Treebank is the first version of a textual annotation system based on the Rhetorical Structure Theory . document is annotated using the RST, a global coherence model proposed by Mann and Thompson . |
| Approach: | They introduce the first version of the Czech RST Discourse Treebank . paper presents an annotation process and provides corpus statistics and evaluation . |
| Outcome: | The paper presents the first version of the Czech RST Discourse Treebank . the treebank includes two gold annotations representing divergent interpretations . |
DiMLex-Bangla: A Lexicon of Bangla Discourse Connectives (2020.lrec-1)
Copied to clipboard
| Challenge: | Discourse connectives are widely believed to be the most explicit, prototypical and most reliable relational signals in discourse processing. |
| Approach: | They present a newly developed lexicon of Bangla discourse connectives . it contains 123 Bangla connective entries, which are primarily compiled from literature . |
| Outcome: | The lexicon contains 123 Bangla connective entries, which are compiled from the linguistic literature and translation of English discourse connectives. |
MEGA RST Discourse Treebanks with Structure and Nuclearity from Scalable Distant Sentiment Supervision (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing discourse treebanks are limited in the application of data-driven approaches to discourse parsing. |
| Approach: | They propose a method to automatically generate discourse treebanks using distant supervision from sentiment annotated datasets by heuristic beam-search strategy extended with a stochastic component. |
| Outcome: | The proposed method generates discourse trees incorporating structure and nuclearity for documents of arbitrary length using an efficient beam-search strategy, extended with a stochastic component. |
The Potsdam Commentary Corpus 2.2: Extending Annotations for Shallow Discourse Parsing (2020.lrec-1)
Copied to clipboard
| Challenge: | Potsdam Commentary Corpus 2.2 is a german corpus of news editorials annotated on several levels. |
| Approach: | They propose to add relation senses to an already existing layer of discourse connectives and their arguments and a new layer with additional coherence relation types to the potsdam commentary corpus. |
| Outcome: | The proposed corpus is more usable for shallow discourse parsing. |
Gold Standard Bangla OCR Dataset: An In-Depth Look at Data Preprocessing and Annotation Processes (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Existing datasets designed specifically for the Bengali language have been limited. |
| Approach: | They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR. |
| Outcome: | The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images. |
Neural RST-based Evaluation of Discourse Coherence (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing discourse parsers cannot predict coherent texts without using silver-standard features. |
| Approach: | They propose a tree-recursive neural model which takes advantage of the text’s RST features produced by a state of the art RST parser and compares it to the current state of art. |
| Outcome: | The proposed model achieves state-of-the-art accuracy on the Grammarly Corpus for Discourse Coherence (GCDC) and has 62% fewer parameters than existing models. |
Persian Discourse Treebank and coreference corpus (L18-1)
Copied to clipboard
| Challenge: | Currently, we are adding a new document-level discourse annotation to our new corpus. |
| Approach: | They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution. |
| Outcome: | The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens. |
SHONGLAP: A Large Bengali Open-Domain Dialogue Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing open-domain dialogue systems suffer from data scarcity due to unavailability of high-quality datasets for low-resource languages like Bengali. |
| Approach: | They propose to prepare large-scale open-domain dialogue datasets from podcasts and talk-shows and label them based on weak-supervision techniques. |
| Outcome: | The proposed corpus improves performance of large language models in case of downstream classification tasks during fine-tuning. |
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)
Copied to clipboard
| Challenge: | CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses) |
| Approach: | They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic. |
| Outcome: | The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses) |
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)
Copied to clipboard
| Challenge: | a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task. |
| Approach: | They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank. |
| Outcome: | The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank. |