| Challenge: | a predominantly German corpus of financial documents is available for the first time . financial text is characterized by a unique vocabulary with implications including sentiment analysis . |
| Approach: | They propose a predominantly German financial corpus comprising 12.5k PDF documents . they hope it will fill this gap and foster further research in the financial domain . |
| Outcome: | The proposed corpus is the first non-email German financial corpus available . it aims to provide insights into financial discourse in the German language and multilingually. |
Similar Papers
MultiFin: A Dataset for Multilingual Financial NLP (2023.findings-eacl)
Copied to clipboard
| Challenge: | Multilingual models are needed to process financial text, which is produced across the world and requires a large dataset. |
| Approach: | They propose to annotate a publicly available financial dataset using a hierarchical label structure and an annotation schema based on a real-world application. |
| Outcome: | The proposed model can be used in high-resource languages, but there is room for improvement in low-resourced languages. |
CoFiF Plus: A French Financial Narrative Summarisation Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language. |
| Approach: | They propose to use natural language processing to analyse financial documents to find the best summarisation methods. |
| Outcome: | The proposed dataset is the first to provide a comprehensive set of financial text written in French. |
A Corpus for Automatic Readability Assessment and Text Simplification of German (2020.lrec-1)
Copied to clipboard
| Challenge: | Using monolingual-only data, we can automate readability assessment and text simplification of simplified language. |
| Approach: | They present a corpus for automatic readability assessment and automatic text simplification for German using parallel and monolingual data. |
| Outcome: | The proposed corpus is compiled from web sources and contains information on text structure, typography, font style, and images. |
NLP Analytics in Finance with DoRe: A French 250M Tokens Corpus of Corporate Annual Reports (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent advances in neural computing and word embeddings for semantic processing open many new applications areas which had been left unaddressed due to inadequate language understanding capacity. |
| Approach: | They propose a French and dialectal French corpus for NLP analytics in finance, regulation and investment. |
| Outcome: | The proposed corpus is designed to be as modular as possible to allow for maximum reuse in different tasks pertaining to Economics, Finance and Investment. |
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application (2026.acl-long)
Copied to clipboard
Xueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie
| Challenge: | Existing evaluations of LLMs in finance are text-only, monolingual, and largely saturated by current models. |
| Approach: | They propose a multilingual and multimodal benchmark for evaluating LLMs in real financial contexts. |
| Outcome: | The first expert-annotated multilingual and multimodal benchmark is released . it evaluates 21 leading LLMs and shows they perform better in multilingual settings . |
LibriVoxDeEn: A Corpus for German-to-English Speech Translation and German Speech Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of sentence-aligned triples of German audio, German text, and English translation is available for speech recognition . a large corpus is available to date for end-to-end speech translation based on parallel data . |
| Approach: | They present a corpus of sentence-aligned triples of German audio, German text, and English translation based on German audio books. |
| Outcome: | The proposed corpus is the largest resource for German speech recognition and for end-to-end German-to English speech translation. |
SEDAR: a Large Scale French-English Financial Domain Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing approaches for neural machine translation use small amount of data or monolingual data. |
| Approach: | They describe acquisition, preprocessing and characteristics of a large English-French parallel corpus for the financial domain. |
| Outcome: | The proposed corpus contains 8.6 million high quality sentence pairs . the first release of the corpus is available on github. |
Corpus REDEWIEDERGABE (2020.lrec-1)
Copied to clipboard
| Challenge: | The corpus REDEWIEDERGABE contains detailed annotations for speech, thought and writing representation (ST&WR) with approximately 490,000 tokens, it is the largest resource of its kind. |
| Approach: | This paper presents corpus REDEWIEDERGABE, a German-language historical corpus with detailed annotations for speech, thought and writing representation (ST&WR). |
| Outcome: | The corpus REDEWIEDERGABE contains 490,000 tokens and is the largest resource of its kind. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
GIL-GALaD: Gender Inclusive Language - German Auto-Assembled Large Database (2024.lrec-main)
Copied to clipboard
| Challenge: | grammatically gendered languages such as German pose unique challenges in generating gender-inclusive language for corrective model training or fine-tuning. |
| Approach: | a corpus of German gender-inclusive language is assembled to help improve model training . grammatically gendered languages such as german pose unique challenges . authors describe most common strategies for gender- inclusive language in german . |
| Outcome: | a corpus of German gender-inclusive language is assembled and will be included in the release. |