| Challenge: | a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level . |
| Approach: | a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level . |
| Outcome: | The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts . |
Similar Papers
German Parliamentary Corpus (GerParCor) Reloaded (2024.lrec-main)
Copied to clipboard
| Challenge: | In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor. |
| Approach: | They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format. |
| Outcome: | The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version. |
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)
Copied to clipboard
| Challenge: | specialized corpora on a variety of topics are available online, but online corporates are scarce. |
| Approach: | They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists . |
| Outcome: | The proposed corpus contains more than six million speeches in English and Chinese and is available for free online. |
How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates (2024.lrec-main)
Copied to clipboard
| Challenge: | a new perspective on framing through the lens of speech acts investigates how politicians make use of different pragmatic speech act functions in political debates. |
| Approach: | They propose a new framework for framing through the lens of speech acts and an annotation scheme for political debates. |
| Outcome: | The proposed framework can predict speech acts with an avg. F1 of around 82.0% . the proposed framework is based on a dataset of German parliamentary debates . |
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve . |
| Approach: | They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus . |
| Outcome: | The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies. |
German Parliamentary Corpus (GerParCor) (2022.lrec-1)
Copied to clipboard
| Challenge: | German parliaments have a large and partly unexploited treasure trove of publicly accessible texts. |
| Approach: | a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus . |
| Outcome: | the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries. |
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)
Copied to clipboard
| Challenge: | DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text. |
| Approach: | They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports. |
| Outcome: | The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year. |
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)
Copied to clipboard
| Challenge: | Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications. |
| Approach: | They analyze available dialogue corpora and propose possible approaches to create new ones. |
| Outcome: | The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed . |
The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)
Copied to clipboard
Väinö Aleksi Yrjänäinen, Fredrik Mohammadi Norén, Robert Borges, Johan Jarlbrink, Lotta Åberg Brorsson, Anders P. Olsson, Pelle Snickars, Måns Magnusson
| Challenge: | The Swedish Parliament Corpus is a new research corpus for the Swedish parliament. |
| Approach: | They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years. |
| Outcome: | The new corpus facilitates detailed analysis of parliamentary speeches in several research fields. |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)
Copied to clipboard
| Challenge: | a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |
| Approach: | They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems. |
| Outcome: | The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged. |