Challenge: Existing studies have highlighted the importance of and need to create learner corpora.
Approach: They propose to create a Qatari corpus of argumentative writing (QCAW) the corpus contains 200,000 tokens of argumentation written by Qatari university students .
Outcome: The QCAW contains 195 essays written by 195 students, 159 females and 36 males.

Similar Papers

The Bahrain Corpus: A Multi-genre Corpus of Bahraini Arabic (2022.lrec-1)

Copied to clipboard

Challenge: Various corpora of various sizes and representing different genres, have been created for various Arabic dialects.
Approach: They propose to create a specialized corpus of Bahraini Arabic dialect, which includes written texts as well as transcripts of audio files.
Outcome: The proposed corpus includes 620K words representing the Bahraini Arabic dialect . the annotated corpus is available to support researchers interested in Arabic NLP .
ZAEBUC: An Annotated Arabic-English Bilingual Writer Corpus (2022.lrec-1)

Copied to clipboard

Challenge: ZAEBUC is an annotated Arabic-English bilingual writer corpus . it is a corpus of short essays written by first-year university students .
Approach: They propose to use a standard Arabic-English bilingual writer corpus to match comparable texts written by the same writer on different occasions.
Outcome: The ZAEBUC corpus is an annotated Arabic-English bilingual writer corpus by first-year university students at Zayed University in the United Arab Emirates.
A Morphologically Annotated Corpus of Emirati Arabic (L18-1)

Copied to clipboard

Challenge: Emirati Arabic corpus is first large-scale morphologically manually annotated corpus . resources for dialectal Arabic NLP tasks are still lacking compared to those for modern standard Arabic (MSA).
Approach: They propose to annotate a large-scale corpus of Emirati Arabic using a morphologically manually annotated corpus from eight Gumar novels . they discuss the guidelines for each part of the annotation components, and the annotation interface they use.
Outcome: The annotated corpus includes about 200,000 words from eight Gumar novels in the Emirati Arabic variety.
A School Student Essay Corpus for Analyzing Interactions of Argumentative Structure and Quality (2024.naacl-long)

Copied to clipboard

Challenge: Existing arguments mining corpus with ground-truth quality annotations is lacking . authors propose baseline approaches to argument mining and essay scoring .
Approach: They propose to use argumentative structure to support argumentative writing . they use an annotated german corpus to analyze interactions between the two tasks .
Outcome: The proposed methods can be used to support argumentative writing . they analyze interactions between argumentative structure and quality annotations .
A Leveled Reading Corpus of Modern Standard Arabic (L18-1)

Copied to clipboard

Challenge: Using a reading corpus in Modern Standard Arabic, we explore the lexical coverage of textbooks and unabridged works of fiction.
Approach: They propose to use textbooks from the United Arab Emirates curriculum and a reading corpus in Modern Standard Arabic to enrich the sparse collection of resources available for educational applications.
Outcome: The corpus spans all 12 grades and contains 129 unabridged works of fiction spanning grades 1-12 . lexical coverage is compared to other genres, and the results show that the two sub-corpora are similar to each other to measure their genres.
A Corpus of Non-Native Written English Annotated for Metaphor (N18-2)

Copied to clipboard

Challenge: Using argumentation-relevant metaphor predicts a holistic score of essay quality, we show .
Approach: They present a corpus of argumentative essays annotated for metaphor by non-native speakers of English . they also examine the relationship between writing proficiency and metaphor use .
Outcome: The proposed corpus is made publicly available and evaluated . it shows that metaphor is a significant predictor of a holistic score of essay quality .
DARIUS: A Comprehensive Learner Corpus for Argument Mining in German-Language Essays (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora focus on specific out-of-school domains, such as legal documents.
Approach: They present a digital argumentation instruction for science corpus on 4589 essays written by 1839 german secondary school students.
Outcome: The proposed corpus is annotated according to a fine-grained annotation scheme on 4589 essays written by 1839 german secondary school students.
Constructing a Bilingual Hadith Corpus Using a Segmentation Tool (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on Hadith have focused on the Quran, leaving it relatively unexplored.
Approach: They propose to gather and construct a bilingual parallel corpus of Islamic Hadith using a custom segmentation tool that annotates the two Hadithe components with 92% accuracy.
Outcome: The proposed method minimises the costs of language resource creation and produces consistent results independently from previous knowledge and experiences that usually influence human annotators.
A Corpus of Encyclopedia Articles with Logical Forms (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of annotated typed lambda calculus translations is described in this paper . typed Lambda Calculus expressions are intended to serve as a theory-neutral formal representation .
Approach: They describe an annotated corpus of typed lambda calculus translations for 2,000 sentences in Simple English Wikipedia.
Outcome: The annotated typed lambda calculus translations are used in a corpus of 2,000 sentences in Simple English Wikipedia.
The Hebrew Essay Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers .
Approach: They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use.
Outcome: The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations