BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)
Copied to clipboard
Nayla Escribano, Jon Ander Gonzalez, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Simón Peña-Fernández, Olatz Perez-de-Viñaspre, Rodrigo Agerri
| Challenge: | a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research . |
| Approach: | They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches. |
| Outcome: | The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender . |
Similar Papers
The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)
Copied to clipboard
Väinö Aleksi Yrjänäinen, Fredrik Mohammadi Norén, Robert Borges, Johan Jarlbrink, Lotta Åberg Brorsson, Anders P. Olsson, Pelle Snickars, Måns Magnusson
| Challenge: | The Swedish Parliament Corpus is a new research corpus for the Swedish parliament. |
| Approach: | They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years. |
| Outcome: | The new corpus facilitates detailed analysis of parliamentary speeches in several research fields. |
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve . |
| Approach: | They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus . |
| Outcome: | The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies. |
FREDSum: A Dialogue Summarization Corpus for French Political Debates (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in deep learning have improved the performance of abstractive summarization systems. |
| Approach: | They present a dataset of french political debates to enhance resources for multi-lingual dialogue summarization. |
| Outcome: | The proposed dataset will be made publicly available for use by the research community. |
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)
Copied to clipboard
| Challenge: | specialized corpora on a variety of topics are available online, but online corporates are scarce. |
| Approach: | They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists . |
| Outcome: | The proposed corpus contains more than six million speeches in English and Chinese and is available for free online. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)
Copied to clipboard
| Challenge: | Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country. |
| Approach: | They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching. |
| Outcome: | The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching. |
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-1)
Copied to clipboard
| Challenge: | a dataset of written code-switched productions is curated from topical threads of multiple bilingual communities on the Reddit discussion platform. |
| Approach: | They analyze a dataset of written code-switched productions curated from multiple bilingual communities on the reddit discussion platform and examine whether findings are carried over to written codeswitching in discussion forums. |
| Outcome: | The proposed dataset can facilitate a range of research and practical activities. |
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-55)
Copied to clipboard
| Challenge: | a dataset of written multilingual productions is released to explore the sociolinguistic underpinnings of written code-switching . |
| Approach: | They use a reddit discussion platform to collect written code-switched productions . they examine whether oral code-witching findings are carried over to written code . |
| Outcome: | The proposed dataset can facilitate a range of research and practical activities. |
Computational Analysis of Political Texts: Bridging Research Efforts Across Communities (P19-4)
Copied to clipboard
| Challenge: | Political scientists have developed and adopted natural language processing (NLP) methods to exploit text as an additional source of data in their analyses. |
| Approach: | This tutorial aims to provide a gentle introduction to methods and tasks related to computational analysis of political texts from both communities. |
| Outcome: | The main goal of this tutorial is to bring the two research communities closer to each other and contribute to faster and more significant developments in this interdisciplinary area. |