Persian Discourse Treebank and coreference corpus (L18-1)

Copied to clipboard

Challenge: Currently, we are adding a new document-level discourse annotation to our new corpus.
Approach: They propose to build a Persian discourse treebank and a comprehensive Persian coreference corpus based on discourse analysis and coreference resolution.
Outcome: The proposed corpus includes 30000 individual sentences with morphological, syntactic and semantic labels and nearly half a million tokens.

Similar Papers

SzegedKoref: A Hungarian Coreference Corpus (L18-1)

Copied to clipboard

Challenge: SzegedKoref is a treebank of Hungarian that contains manual annotation at several linguistic layers.
Approach: They introduce a Hungarian corpus in which coreference relations are manually annotated.
Outcome: The proposed corpus can be used in training and testing machine learning based coreference resolution systems.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
Cost-Effective Discourse Annotation in the Prague Czech–English Dependency Treebank (2024.lrec-main)

Copied to clipboard

Challenge: a method for obtaining a high-quality annotation of explicit discourse relations is a resource-demanding task.
Approach: They propose a method for obtaining a high-quality annotation of explicit discourse relations in the Czech part of the Prague Czech–English Dependency Treebank.
Outcome: The proposed method solves the problem of identifying discrepancies between the annotations in the Czech part of the Penn Treebank.
The Persian Dependency Treebank Made Universal (2022.lrec-1)

Copied to clipboard

Challenge: Existing universal dependency treebanks are lacking sufficient annotated data.
Approach: They propose a method for converting Persian Dependency Treebank to Universal Dependencies using an automatic method.
Outcome: The proposed method is more compatible with Universal Dependencies than the Uppsala Persian Universal Dependency Treebank.
Camel Treebank: An Open Multi-genre Arabic Dependency Treebank (2022.lrec-1)

Copied to clipboard

Challenge: CAMELTB is an open-source dependency treebank of Arabic with 13 sub-corpora . texts are publicly available (out of copyright, creative commons, or under open licenses)
Approach: They present the Camel Treebank, a 188K word open-source dependency treebank of Arabic.
Outcome: The CAMELTB is a 188K word open-source dependency treebank of Arabic . the texts are publicly available (out of copyright, creative commons, or under open licenses)
ezCoref: Towards Unifying Annotation Guidelines for Coreference Resolution (2023.findings-eacl)

Copied to clipboard

Challenge: Existing datasets vary in definition of coreferences and are curated for linguistic experts.
Approach: They propose to use ezCoref to create a crowdsourcing-friendly coreference annotation methodology that teaches annotators only cases that are treated similarly across existing datasets.
Outcome: The proposed method reannotates 240 passages from seven existing english coreference datasets while teaching annotators only cases that are treated similarly across them.
Implicit Discourse Relation Classification: We Need to Talk about Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in literature.
Approach: They propose an improved evaluation protocol for implicit relation classification on PDTB 2.0 . they report strong baseline results from pretrained sentence encoders .
Outcome: The proposed evaluation protocol improves the existing framework and provides strong baseline results.
ParCorFull: a Parallel Corpus Annotated with Full Coreference (L18-1)

Copied to clipboard

Challenge: Recent research in multilingual coreference and automatic pronoun translation has led to important insights into the problem and some promising results.
Approach: They propose a corpus annotated with full coreference chains that addresses a problem that machine translation and other multilingual natural language processing (NLP) technologies face: translation of coreference across languages.
Outcome: The proposed corpus contains parallel texts for the language pair English-German, two major European languages.
MCDTB: A Macro-level Chinese Discourse TreeBank (C18-1)

Copied to clipboard

Challenge: Discourse analysis is becoming increasingly important in the field of natural language processing.
Approach: They propose to annotate macro discourse information and additional discourse information to make annotation more objective and accurate.
Outcome: The results show that the annotations are more objective and accurate than the previous ones.
Informal Persian Universal Dependency Treebank (2022.lrec-1)

Copied to clipboard

Challenge: phonological, morphological, and syntactic distinctions between formal and informal Persian are important . formal Persian is not a universally recognized form of language, but is a dialect of informal Persian .
Approach: They develop an open-source treebank for informal Persian to train dependency parsers . they then train dependency lexicographers on existing treebanks and evaluate them on out-of-domain data .
Outcome: The proposed treebanks show that they perform poorly when training on formal and informal Persians.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations