The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)

Copied to clipboard

Challenge: The Swedish Parliament Corpus is a new research corpus for the Swedish parliament.
Approach: They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years.
Outcome: The new corpus facilitates detailed analysis of parliamentary speeches in several research fields.

Similar Papers

BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions (2022.lrec-1)

Copied to clipboard

Challenge: a new corpus of Basque parliamentary transcripts is released to study political discourse in contrasting languages . a corpus containing political discourses from public institutions can be used for computational social science research .
Approach: They present a corpus from Basque parliamentary transcripts and enrich it with metadata related to relevant attributes of speakers and speeches.
Outcome: The proposed corpus is characterized by heavy Basque-Spanish code-switching . it provides interesting insights about language use of political representatives across time, parties and gender .
German Parliamentary Corpus (GerParCor) Reloaded (2024.lrec-main)

Copied to clipboard

Challenge: In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor.
Approach: They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format.
Outcome: The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version.
A corpus of German political speeches from the 21st century (L18-1)

Copied to clipboard

Challenge: a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level .
Approach: a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level .
Outcome: The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts .
SLäNDa: An Annotated Corpus of Narrative and Dialogue in Swedish Literary Fiction (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary fiction has been annotated for cited materials with a focus on dialogue.
Approach: They propose to annotate a new corpus of Swedish literary fiction for cited materials with a focus on dialogue.
Outcome: The proposed corpus can be used to train and analyze models for different types of analysis of literary narrative and speech.
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
German Parliamentary Corpus (GerParCor) (2022.lrec-1)

Copied to clipboard

Challenge: German parliaments have a large and partly unexploited treasure trove of publicly accessible texts.
Approach: a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus .
Outcome: the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries.
SLäNDa version 2.0: Improved and Extended Annotation of Narrative and Dialogue in Swedish Literature (2022.lrec-1)

Copied to clipboard

Challenge: SLäNDa is the only large-scale corpus of literature annotated for narrative, speech and speakers.
Approach: They propose to annotate a version 2.0 of the SLäNDa corpus which includes 19 novels . they specifically examine different ways of marking speech segments such as quotation marks, dashes, or no marking at all.
Outcome: The proposed corpus contains excerpts from 19 novels written between 1809 and 1940.
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve .
Approach: They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus .
Outcome: The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies.
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results .
Approach: They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach .
Outcome: The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts.
Open Political Corpora: Structuring, Searching, and Analyzing Political Text Collections with PoliCorp (2025.emnlp-demos)

Copied to clipboard

Challenge: PoliCorp provides researchers with access to rich textual data, enabling in-depth analysis of parliamentary discourse over time.
Approach: They present a web portal that allows researchers to search political text corpora . the platform currently contains a collection of transcripts from the german parliament .
Outcome: The proposed platform provides researchers with access to rich textual data, enabling in-depth analysis of parliamentary discourse over time.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations