Challenge: Using BiPaR, we build monolingual, multilingual and cross-lingual MRC on novels.
Approach: They propose a bilingual parallel novel-style machine reading comprehension dataset BiPaR . they collect 3,667 bilingual parallel paragraphs from Chinese and English novels .
Outcome: The proposed dataset supports multilingual and cross-lingual reading comprehension.

Similar Papers

The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for text comprehension only cover 30 languages, but lack of labeled data is a major obstacle to building functional systems in most languages.
Approach: They present a multiple-choice machine reading comprehension dataset spanning 122 languages . they use it to evaluate the capabilities of multilingual masked language models and large language models .
Outcome: The proposed dataset enables the evaluation of text models in high-, medium- and low-resource languages.
Cross-Lingual Machine Reading Comprehension (D19-1)

Copied to clipboard

Challenge: Existing work on machine reading comprehension task is focused on English, but there are few efforts on other languages due to the lack of large-scale training data.
Approach: They propose a cross-lingual machine reading comprehension task for other languages . they propose cloze-style reading comprehension and various neural network approaches .
Outcome: The proposed model improves reading comprehension performance of Chinese datasets over state-of-the-art systems by a large margin over existing systems.
DuoRC: Towards Complex Language Understanding with Paraphrased Reading Comprehension (P18-1)

Copied to clipboard

Challenge: DuoRC contains 186,089 unique question-answer pairs created from 7680 movie plots .
Approach: They propose a novel dataset for Reading Comprehension that motivates new challenges for neural approaches in language understanding beyond those offered by existing RC datasets.
Outcome: The proposed dataset motivates several new challenges for neural approaches in language understanding beyond those offered by existing RC datasets.
A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora (L18-1)

Copied to clipboard

Challenge: BUCC Shared Task aims to extract parallel sentences from comparable corporad . resulting corpus contains about 3.5 million distinct sentences in english, french, german, Russian, and Chinese .
Approach: They present challenges faced to build a parallel sentences dataset from comparable corporad . they emphasize issues faced to include Chinese as one of the languages .
Outcome: The 2017 BUCC Shared Task was a first for this task . the dataset contains 3.5 million sentences in English, French, German, Russian, and Chinese .
A Span-Extraction Dataset for Chinese Machine Reading Comprehension (D19-1)

Copied to clipboard

Challenge: Existing reading comprehension datasets are mostly in English . MRC is a new field of research that aims to comprehend the context of articles and answer the questions based on them.
Approach: They propose a Span-Extraction dataset for Chinese machine reading comprehension to add language diversities to existing reading comprehension datasets.
Outcome: The proposed dataset is composed of 20,000 real questions annotated on Wikipedia paragraphs by human experts.
Dataset for the First Evaluation on Chinese Machine Reading Comprehension (L18-1)

Copied to clipboard

Challenge: Existing reading comprehension datasets are mostly in English .
Approach: They propose a Chinese reading comprehension dataset to add diversity to existing reading comprehension data . proposed dataset contains cloze-style reading comprehension and user query reading comprehension .
Outcome: The proposed dataset is based on a Chinese reading comprehension dataset . it includes two types of cloze-style and user query reading comprehension . the proposed dataset hosted the 1st Evaluation on Chinese Machine Reading Comprehension (CMRC-2017)
CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1,500+ Language Pairs (2023.acl-long)

Copied to clipboard

Challenge: a large-scale cross-lingual summarization dataset is available for free . a cross-linguistic summarizing model can be trained in any target language .
Approach: They propose a multistage data sampling algorithm to train a cross-lingual summarization model capable of summarizing an article in any target language.
Outcome: The proposed model outperforms baseline models on ROUGE and LaSE.
A Survey on Cross-Lingual Summarization (2022.tacl-1)

Copied to clipboard

Challenge: Cross-lingual summarization is a task of generating a summary in one language for a given document in a different language.
Approach: They present a systematic review of the literature on cross-lingual summarization . they summarize previous efforts and compare them with each other .
Outcome: The proposed approach is compared with previous approaches and summarizes them to provide a deeper analysis.
MLSUM: The Multilingual Summarization Corpus (2020.emnlp-main)

Copied to clipboard

Challenge: Existing biases in multi-lingual datasets are limiting the use of multilingual data in document summarization tasks.
Approach: They present MLSUM, the first large-scale MultiLingual SUMmarization dataset.
Outcome: The proposed dataset contains 1.5M+ article/summary pairs in five different languages.
English Machine Reading Comprehension Datasets: A Survey (2021.emnlp-main)

Copied to clipboard

Challenge: a survey of English Machine Reading Comprehension datasets is carried out . the aim is to provide a concise yet informative overview of the landscape .
Approach: They survey 60 English Machine Reading Comprehension datasets to provide a resource for other researchers interested in this problem.
Outcome: The proposed survey covers 60 English MRC datasets with a view to providing a resource for other researchers interested in the problem.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations