Challenge: a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research .
Approach: They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches.
Outcome: The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender .

Similar Papers

The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)

Copied to clipboard

Challenge: The Swedish Parliament Corpus is a new research corpus for the Swedish parliament.
Approach: They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years.
Outcome: The new corpus facilitates detailed analysis of parliamentary speeches in several research fields.
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve .
Approach: They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus .
Outcome: The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies.
FREDSum: A Dialogue Summarization Corpus for French Political Debates (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in deep learning have improved the performance of abstractive summarization systems.
Approach: They present a dataset of french political debates to enhance resources for multi-lingual dialogue summarization.
Outcome: The proposed dataset will be made publicly available for use by the research community.
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)

Copied to clipboard

Challenge: specialized corpora on a variety of topics are available online, but online corporates are scarce.
Approach: They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists .
Outcome: The proposed corpus contains more than six million speeches in English and Chinese and is available for free online.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country.
Approach: They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching.
Outcome: The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching.
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-1)

Copied to clipboard

Challenge: a dataset of written code-switched productions is curated from topical threads of multiple bilingual communities on the Reddit discussion platform.
Approach: They analyze a dataset of written code-switched productions curated from multiple bilingual communities on the reddit discussion platform and examine whether findings are carried over to written codeswitching in discussion forums.
Outcome: The proposed dataset can facilitate a range of research and practical activities.
CodeSwitch-Reddit: Exploration of Written Multilingual Discourse in Online Discussion Forums (D19-55)

Copied to clipboard

Challenge: a dataset of written multilingual productions is released to explore the sociolinguistic underpinnings of written code-switching .
Approach: They use a reddit discussion platform to collect written code-switched productions . they examine whether oral code-witching findings are carried over to written code .
Outcome: The proposed dataset can facilitate a range of research and practical activities.
Computational Analysis of Political Texts: Bridging Research Efforts Across Communities (P19-4)

Copied to clipboard

Challenge: Political scientists have developed and adopted natural language processing (NLP) methods to exploit text as an additional source of data in their analyses.
Approach: This tutorial aims to provide a gentle introduction to methods and tasks related to computational analysis of political texts from both communities.
Outcome: The main goal of this tutorial is to bring the two research communities closer to each other and contribute to faster and more significant developments in this interdisciplinary area.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations