| Challenge: | In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor. |
| Approach: | They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format. |
| Outcome: | The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version. |
Similar Papers
German Parliamentary Corpus (GerParCor) (2022.lrec-1)
Copied to clipboard
| Challenge: | German parliaments have a large and partly unexploited treasure trove of publicly accessible texts. |
| Approach: | a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus . |
| Outcome: | the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries. |
A corpus of German political speeches from the 21st century (L18-1)
Copied to clipboard
| Challenge: | a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level . |
| Approach: | a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level . |
| Outcome: | The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts . |
The GermaParl Corpus of Parliamentary Protocols (L18-1)
Copied to clipboard
| Challenge: | Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making. |
| Approach: | They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations. |
| Outcome: | The proposed corpus provides a valuable resource for research and teaching purposes. |
The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)
Copied to clipboard
Väinö Aleksi Yrjänäinen, Fredrik Mohammadi Norén, Robert Borges, Johan Jarlbrink, Lotta Åberg Brorsson, Anders P. Olsson, Pelle Snickars, Måns Magnusson
| Challenge: | The Swedish Parliament Corpus is a new research corpus for the Swedish parliament. |
| Approach: | They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years. |
| Outcome: | The new corpus facilitates detailed analysis of parliamentary speeches in several research fields. |
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results . |
| Approach: | They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach . |
| Outcome: | The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts. |
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)
Copied to clipboard
| Challenge: | DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text. |
| Approach: | They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports. |
| Outcome: | The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year. |
The MARCELL Legislative Corpus (2020.lrec-1)
Copied to clipboard
Tamás Váradi, Svetla Koeva, Martin Yamalov, Marko Tadić, Bálint Sass, Bartłomiej Nitoń, Maciej Ogrodniczuk, Piotr Pęzik, Verginica Barbu Mititelu, Radu Ion, Elena Irimia, Maria Mitrofan, Vasile Păiș, Dan Tufiș, Radovan Garabík, Simon Krek, Andraz Repar, Matjaž Rihtar, Janez Brank
| Challenge: | MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification. |
| Approach: | They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents . |
| Outcome: | The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents. |
Evolving Large Text Corpora: Four Versions of the Icelandic Gigaword Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | The Icelandic Gigaword Corpus was first published in 2018 and has since grown to include more than 50 million words. |
| Approach: | They describe the evolution of the Icelandic Gigaword Corpus in its first four years . they show how the corpus has grown almost 50% in size from the first version to the fourth . |
| Outcome: | The Gigaword corpus has grown 50% from its first version to its fourth version and is now available under permissive licenses. |
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)
Copied to clipboard
| Challenge: | outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels . |
| Approach: | They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words . |
| Outcome: | The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language. |
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)
Copied to clipboard
Florian Borchert, Christina Lohr, Luise Modersohn, Jonas Witt, Thomas Langer, Markus Follmann, Matthias Gietzelt, Bert Arnrich, Udo Hahn, Matthieu-P. Schapranow
| Challenge: | despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language. |
| Approach: | They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set . |
| Outcome: | The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology. |