TED-CDB: A Large-Scale Chinese Discourse Relation Dataset on TED Talks (2020.emnlp-main)
Copied to clipboard
| Challenge: | TED-CDB dataset is a unique corpus of spoken discourse in Chinese . TED is based on the concept that discourse relations are grounded in an identifiable set of discourse connectives or Altlex expressions. |
| Approach: | They have created a dataset that annotates TED talks in Chinese . they propose to adapt the dataset to Chinese news text to improve its performance . |
| Outcome: | The TED-CDB dataset can improve the performance of systems for languages other than Chinese . it is adapted to features that are not present in English and can extract discourse semantic features . |
Similar Papers
Shallow Discourse Annotation for Chinese TED Talks (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to annotate text with discourse properties are limited to newspaper articles and are not available in Chinese. |
| Approach: | They propose to annotate TED talks with Chinese-related properties using the Penn Discourse TreeBank annotation style . they propose to use planned monologues instead of written text to annnotate Chinese-specific properties. |
| Outcome: | The proposed method is able to achieve reliable results in Chinese spoken monologues, and is based on the Penn Discourse TreeBank annotation style. |
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)
Copied to clipboard
| Challenge: | Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications. |
| Approach: | They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages. |
| Outcome: | The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages. |
GCDT: A Chinese RST Treebank for Multigenre and Multilingual Discourse Parsing (2022.aacl-short)
Copied to clipboard
| Challenge: | GCDT is the largest hierarchical discourse treebank for Mandarin Chinese in the framework of Rhetorical Structure Theory (RST). |
| Approach: | They propose to use a Chinese hierarchical discourse treebank to parse Mandarin Chinese using relation inventory and a multilingual training program. |
| Outcome: | The proposed dataset includes state-of-the-art scores for Chinese RST parsing and RST Parsing on the English GUM dataset, using cross-lingual training in Chinese and English with multilingual embeddings. |
MCDTB: A Macro-level Chinese Discourse TreeBank (C18-1)
Copied to clipboard
| Challenge: | Discourse analysis is becoming increasingly important in the field of natural language processing. |
| Approach: | They propose to annotate macro discourse information and additional discourse information to make annotation more objective and accurate. |
| Outcome: | The results show that the annotations are more objective and accurate than the previous ones. |
A Unified RvNN Framework for End-to-End Chinese Discourse Parsing (C18-2)
Copied to clipboard
| Challenge: | Existing work on Chinese discourse parser relies on external packages to extract linguistic features from free text. |
| Approach: | They propose an end-to-end Chinese discourse parser based on recursive neural network to jointly model the subtasks including elementary discourse unit segmentation, tree structure construction, center labeling, and sense labeling. |
| Outcome: | The proposed framework achieves state-of-the-art in the Chinese Discourse Treebank dataset. |
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)
Copied to clipboard
| Challenge: | Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature. |
| Approach: | They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders . |
| Outcome: | The proposed evaluation protocol improves the existing framework and provides strong baseline results. |
Multi-Label Classification for Implicit Discourse Relation Recognition (2024.findings-acl)
Copied to clipboard
| Challenge: | Prior research in discourse relation recognition has treated these instances as separate examples during training, with a gold-standard prediction matching one of the labels considered correct at test time. |
| Approach: | They propose to use multiple labels to annotate an example when multiple relations are believed to hold simultaneously. |
| Outcome: | The proposed frameworks don't depress performance for single-label prediction. |
CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | a dataset of Chinese large language models is used to measure societal biases . many studies have shown that LLMs exhibit harmful societal biased outputs despite human data . |
| Approach: | They present a Chinese Bias Benchmark dataset that includes over 100K questions constructed by human experts and generative language models. |
| Outcome: | The proposed dataset covers stereotypes and societal biases in 14 social dimensions related to Chinese culture and values. |
TED-Q: TED Talks and the Questions they Evoke (2020.lrec-1)
Copied to clipboard
| Challenge: | Evoked questions represent a hitherto unexplored type of linguistic data, promising to open up important new lines of research. |
| Approach: | They propose a method to annotate TED-talks with the questions they evoke and, where available, the answers to these questions. |
| Outcome: | The proposed method is designed to scale up, relying on crowdsourcing by non-expert annotators, with its utility for Natural Language Processing in mind. |
GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and Domains (2024.emnlp-main)
Copied to clipboard
Yang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu, Shabnam Behzad, Lauren Levine, Jessica Lin, Devika Tiwari, Amir Zeldes
| Challenge: | Existing shallow discourse parsing systems focus on the Wall Street Journal corpus, but the data is limited to the news domain and is 35 years old. |
| Approach: | They propose to use the Wall Street Journal corpus as a benchmark for PDTB-style shallow discourse parsing. |
| Outcome: | The proposed dataset is compatible with PDTB, but suffers from degradation out-of-domain. |