Multi-lingual Argumentative Corpora in English, Turkish, Greek, Albanian, Croatian, Serbian, Macedonian, Bulgarian, Romanian and Arabic (L18-1)
Copied to clipboard
Alfred Sliwa, Yuan Ma, Ruishen Liu, Niravkumar Borad, Seyedeh Ziyaei, Mina Ghobadi, Firas Sabbah, Ahmet Aker
| Challenge: | Argumentative corpora are costly to create and available only in few languages with English dominating the area. |
| Approach: | They use 8 different argument mining classifiers trained for English to build a parallel corpora in which the source language is English and the target language is either a Balkan language or Arabic. |
| Outcome: | The proposed method is based on 8 different argument mining classifiers trained for English and project the decision to the target language. |
Similar Papers
End-to-end Argument Mining with Cross-corpora Multi-task Learning (2022.tacl-1)
Copied to clipboard
| Challenge: | Argument(ation) mining is a task of identifying argument structure from text . lack of training data makes it difficult to train models based on limited data sets. |
| Approach: | They propose an end-to-end cross-corpus argument mining method that uses auxiliary argument mining corpora to train models. |
| Outcome: | The proposed method outperforms models trained on a single corpus on arguments on arguments in argument mining tasks. |
Annotating Arguments in a Corpus of Opinion Articles (2022.lrec-1)
Copied to clipboard
Gil Rocha, Luís Trigo, Henrique Lopes Cardoso, Rui Sousa-Silva, Paula Carvalho, Bruno Martins, Miguel Won
| Challenge: | Argument annotation is the process of exposing and justifying one's points of view, with the aim of conveying a logical reasoning through a set of semantically related propositions. |
| Approach: | They propose to use argumentative discourse units to annotate arguments in Portuguese using a multi-layered process to analyze the annotations produced. |
| Outcome: | The proposed model exploits the best practices identified in previous studies while fostering the potential use of the resulting annotated corpus for new purposes. |
Cross-lingual Argumentation Mining: Machine Translation (and a bit of Projection) is All You Need! (C18-1)
Copied to clipboard
| Challenge: | Argumentation mining (AM) requires the identification of complex discourse structures . existing resources are not adequate for assessing cross-lingual AM due to their heterogeneity or lack of complexity. |
| Approach: | They propose to use a dataset to translate persuasive student essays into German, French, Spanish, and Chinese to compare arguments mining and annotation projection. |
| Outcome: | The proposed methods perform better when using expensive human or cheap machine translations and almost eliminate loss from cross-lingual transfer. |
A Multilingual Parallel Corpora Collection Effort for Indian Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Currently, neural network based approaches for machine translation are data hungry and sentence-level aligned parallel pairs are the currency. |
| Approach: | They propose to build sentence aligned parallel corpora across 10 Indian languages using online sources which have content shared across languages. |
| Outcome: | The proposed corpora significantly extends existing resources that are either not large enough or are restricted to a specific domain (such as health). |
A Multilingual Parallel Corpus for Aromanian (2024.lrec-main)
Copied to clipboard
| Challenge: | Aromanian is an endangered 1 language that currently lacks corpora and electronic resources that can potentially contribute to the preservation of its cultural heritage. |
| Approach: | They propose to create a corpus of Aromanian and equivalent sentence-aligned translations into Romanian, English, and French using orthographic standards. |
| Outcome: | The authors report that the first high-quality corpus of Aromanian is available in the Balkans and is available for download in Romanian, English, and French. |
DARIUS: A Comprehensive Learner Corpus for Argument Mining in German-Language Essays (2024.lrec-main)
Copied to clipboard
Nils-Jonathan Schaller, Andrea Horbach, Lars Ingver Höft, Yuning Ding, Jan Luca Bahr, Jennifer Meyer, Thorben Jansen
| Challenge: | Existing corpora focus on specific out-of-school domains, such as legal documents. |
| Approach: | They present a digital argumentation instruction for science corpus on 4589 essays written by 1839 german secondary school students. |
| Outcome: | The proposed corpus is annotated according to a fine-grained annotation scheme on 4589 essays written by 1839 german secondary school students. |
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Recent studies have highlighted the potential of exploiting parallel corpora to enhance multilingual large language models. |
| Approach: | They investigate the impact of parallel corpora quality and quantity, training objectives, and model size on performance of multilingual large language models enhanced with parallel corporeal. |
| Outcome: | The proposed approach improves performance in bilingual and general-purpose tasks. |
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)
Copied to clipboard
| Challenge: | Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data. |
| Approach: | They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes. |
| Outcome: | The proposed corpus includes 112 argumentative microtexts and a new annotation tool. |
The Discussion Tracker Corpus of Collaborative Argumentation (2020.lrec-1)
Copied to clipboard
| Challenge: | The Discussion Tracker corpus is an annotated dataset of transcripts of spoken, multi-party argumentation transcribed from 985 minutes of audio . |
| Approach: | They analyze 29 multi-party arguments transcribed from 985 minutes of audio . they provide descriptive statistics and code for predicting each dimension separately. |
| Outcome: | The Discussion Tracker corpus was collected in high school English classes and annotated for argument moves, specificity, specificities and collaboration dimensions. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |