Challenge: a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level .
Approach: a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level .
Outcome: The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts .

Similar Papers

German Parliamentary Corpus (GerParCor) Reloaded (2024.lrec-main)

Copied to clipboard

Challenge: In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor.
Approach: They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format.
Outcome: The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version.
Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)

Copied to clipboard

Challenge: specialized corpora on a variety of topics are available online, but online corporates are scarce.
Approach: They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists .
Outcome: The proposed corpus contains more than six million speeches in English and Chinese and is available for free online.
How to Do Politics with Words: Investigating Speech Acts in Parliamentary Debates (2024.lrec-main)

Copied to clipboard

Challenge: a new perspective on framing through the lens of speech acts investigates how politicians make use of different pragmatic speech act functions in political debates.
Approach: They propose a new framework for framing through the lens of speech acts and an annotation scheme for political debates.
Outcome: The proposed framework can predict speech acts with an avg. F1 of around 82.0% . the proposed framework is based on a dataset of German parliamentary debates .
A Corpus of Spanish Political Speeches from 1937 to 2019 (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of political speeches in Spanish is documented from 1937 to 2019 . the corpus contains the speeches delivered by the head of state of Spain on Christmas Eve .
Approach: They propose to collect political speeches from the Christmas Eve national speeches from 1937 to 2019 . they propose a Python interface that allows querying and analyzing the corpus .
Outcome: The proposed corpus contains speeches delivered by the king of Spain from 1937 to 2019 . the documents reflect some of the most significant events and political changes in recent history . a set of HTML visualizations is provided to navigate the corpus and explore differences between TF-IDF frequencies.
German Parliamentary Corpus (GerParCor) (2022.lrec-1)

Copied to clipboard

Challenge: German parliaments have a large and partly unexploited treasure trove of publicly accessible texts.
Approach: a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus .
Outcome: the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries.
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)

Copied to clipboard

Challenge: DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text.
Approach: They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports.
Outcome: The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year.
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)

Copied to clipboard

Challenge: The Swedish Parliament Corpus is a new research corpus for the Swedish parliament.
Approach: They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years.
Outcome: The new corpus facilitates detailed analysis of parliamentary speeches in several research fields.
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
A Large Annotated Reference Corpus of New High German Poetry (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of public domain German poetry covering the time period 1600 to the 1920s contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.
Approach: They present a large annotated corpus of public domain German poetry covering the time period 1600 to the 1920s with 65k poems.
Outcome: The corpus contains 65k unique poems and over 1.6M lines, each tokenized, syllabified, pos-tagged, and meter-tagged.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations