FinCorpus-DE10k: A Corpus for the German Financial Domain (2024.lrec-main)

Copied to clipboard

Challenge: a predominantly German corpus of financial documents is available for the first time . financial text is characterized by a unique vocabulary with implications including sentiment analysis .
Approach: They propose a predominantly German financial corpus comprising 12.5k PDF documents . they hope it will fill this gap and foster further research in the financial domain .
Outcome: The proposed corpus is the first non-email German financial corpus available . it aims to provide insights into financial discourse in the German language and multilingually.

Similar Papers

MultiFin: A Dataset for Multilingual Financial NLP (2023.findings-eacl)

Copied to clipboard

Challenge: Multilingual models are needed to process financial text, which is produced across the world and requires a large dataset.
Approach: They propose to annotate a publicly available financial dataset using a hierarchical label structure and an annotation schema based on a real-world application.
Outcome: The proposed model can be used in high-resource languages, but there is room for improvement in low-resourced languages.
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language.
Approach: They propose to use natural language processing to analyse financial documents to find the best summarisation methods.
Outcome: The proposed dataset is the first to provide a comprehensive set of financial text written in French.
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)

Copied to clipboard

Challenge: Using monolingual-only data, we can automate readability assessment and text simplification of simplified language.
Approach: They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data.
Outcome: The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images.
NLP Analytics in Finance with DoRe: A French 250M Tokens Corpus of Corporate Annual Reports (2020.lrec-1)

Copied to clipboard

Challenge: Recent advances in neural computing and word embeddings for semantic processing open many new applications areas which had been left unaddressed due to inadequate language understanding capacity.
Approach: They propose a French and dialectal French corpus for NLP analytics in finance, regulation and investment.
Outcome: The proposed corpus is designed to be as modular as possible to allow for maximum reuse in different tasks pertaining to Economics, Finance and Investment.
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data .
Approach: They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books.
Outcome: The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation.
SEDAR: a Large Scale French-English Financial Domain Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches for neural machine translation use small amount of data or monolingual data.
Approach: They describe acquisition, preprocessing and characteristics of a large English-French parallel corpus for the financial domain.
Outcome: The proposed corpus contains 8.6 million high quality sentence pairs . the first release of the corpus is available on github.
Corpus REDEWIEDERGABE (2020.lrec-1)

Copied to clipboard

Challenge: The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind.
Approach: This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR).
Outcome: The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
GIL-GALaD: Gender Inclusive Language - German Auto-Assembled Large Database (2024.lrec-main)

Copied to clipboard

Challenge: grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning.
Approach: a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german .
Outcome: a corpus of German gender-inclusive language is assembled and will be included in the release.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations