German Parliamentary Corpus (GerParCor) Reloaded (2024.lrec-main)

Copied to clipboard

Challenge: In 2022, the largest German-speaking corpus of parliamentary protocols from three different centuries has been published - GerParCor.
Approach: They propose to update the largest German-speaking corpus of parliamentary protocols from three different centuries, on a national and federal level, from Germany, Austria, Switzerland and Liechtenstein, and to make them available in XMI format.
Outcome: The updated corpus includes all new parliamentary protocols and adds and preprocesses further parliamentary protocol not covered in the previous version.

Similar Papers

German Parliamentary Corpus (GerParCor) (2022.lrec-1)

Copied to clipboard

Challenge: German parliaments have a large and partly unexploited treasure trove of publicly accessible texts.
Approach: a new corpus of German-language parliamentary protocols is made available in XMI format . the corpus is genre-specific and contains conversions of scanned protocols . a researcher at the university of berlin and a professor at the berlin university created the corpuus .
Outcome: the German Parliamentary Corpus is a genre-specific corpus of German-language parliamentary protocols from three centuries and four countries.
A corpus of German political speeches from the 21st century (L18-1)

Copied to clipboard

Challenge: a german political speeches corpus was released in 2017 . the corpus includes the four highest ranked functions on federal state level .
Approach: a new german political speeches corpus is presented . the corpus includes the four highest ranked functions on federal state level .
Outcome: The present German political speeches corpus is updated and extended . it includes the four highest ranked functions on federal state level . the main contributions are an extensive description of the corpus and an interface to navigate through the texts .
The GermaParl Corpus of Parliamentary Protocols (L18-1)

Copied to clipboard

Challenge: Parliamentary debates convey the arguments, interpretations and disputes that shape political decision-making.
Approach: They outline available data, the data preparation process for preparing corpora of parliamentary debates and tools to obtain hand-coded annotations.
Outcome: The proposed corpus provides a valuable resource for research and teaching purposes.
The Swedish Parliament Corpus 1867 – 2022 (2024.lrec-main)

Copied to clipboard

Challenge: The Swedish Parliament Corpus is a new research corpus for the Swedish parliament.
Approach: They propose to expand the Swedish Parliament corpus by providing a database of all members of parliament over 150 years.
Outcome: The new corpus facilitates detailed analysis of parliamentary speeches in several research fields.
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results .
Approach: They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach .
Outcome: The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts.
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)

Copied to clipboard

Challenge: DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text.
Approach: They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports.
Outcome: The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year.
The MARCELL Legislative Corpus (2020.lrec-1)

Copied to clipboard

Challenge: MARCELL corpus provides a rich and valuable source for further studies and developments in machine learning, cross-lingual terminological data extraction and classification.
Approach: They present the results of the project MARCELL CEF Telecom . they aim to collect and deeply annotate a large comparable corpus of legal documents .
Outcome: The MARCELL corpus includes 7 monolingual sub-corpora containing the body of respective national legislative documents.
Evolving Large Text Corpora: Four Versions of the Icelandic Gigaword Corpus (2022.lrec-1)

Copied to clipboard

Challenge: The Icelandic Gigaword Corpus was first published in 2018 and has since grown to include more than 50 million words.
Approach: They describe the evolution of the Icelandic Gigaword Corpus in its first four years . they show how the corpus has grown almost 50% in size from the first version to the fourth .
Outcome: The Gigaword corpus has grown 50% from its first version to its fourth version and is now available under permissive licenses.
Introducing a Parsed Corpus of Historical High German (2024.lrec-main)

Copied to clipboard

Challenge: outlines the development of the Indiana Parsed Corpus of (Historical) High German . outlines selection of texts, decisions on part-of-speech tags and other labels .
Approach: They propose to build a parsed German corpus that spans Germanic from 1050 to 1950 . they propose to use Penn-style treebanks to capture syntactic relationships between words .
Outcome: The proposed corpus spans Germanic languages from 1050 to 1950 and illustrative annotation issues unique to the language.
GGPONC 2.0 - The German Clinical Guideline Corpus for Oncology: Curation Workflow, Annotation Policy, Baseline NER Taggers (2022.lrec-1)

Copied to clipboard

Challenge: despite advances in language resources, there is still a shortage of annotated corpora covering (German) medical language.
Approach: They propose to build on clinical guidelines with an annotation scheme based on SNOMED CT . they also train named entity recognition models on the new data set .
Outcome: The new corpus can be built upon clinical guidelines with reasonable coverage of medical terminology.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations