Challenge: Existing corpora for financial narrative summarisation exists in English but there is a significant lack of financial text resources in the French language.
Approach: They propose to use natural language processing to analyse financial documents to find the best summarisation methods.
Outcome: The proposed dataset is the first to provide a comprehensive set of financial text written in French.

Similar Papers

NLP Analytics in Finance with DoRe: A French 250M Tokens Corpus of Corporate Annual Reports (2020.lrec-1)

Copied to clipboard

Challenge: Recent advances in neural computing and word embeddings for semantic processing open many new applications areas which had been left unaddressed due to inadequate language understanding capacity.
Approach: They propose a French and dialectal French corpus for NLP analytics in finance, regulation and investment.
Outcome: The proposed corpus is designed to be as modular as possible to allow for maximum reuse in different tasks pertaining to Economics, Finance and Investment.
MultiFin: A Dataset for Multilingual Financial NLP (2023.findings-eacl)

Copied to clipboard

Challenge: Multilingual models are needed to process financial text, which is produced across the world and requires a large dataset.
Approach: They propose to annotate a publicly available financial dataset using a hierarchical label structure and an annotation schema based on a real-world application.
Outcome: The proposed model can be used in high-resource languages, but there is room for improvement in low-resourced languages.
FinCorpus-DE10k: A Corpus for the German Financial Domain (2024.lrec-main)

Copied to clipboard

Challenge: a predominantly German corpus of financial documents is available for the first time . financial text is characterized by a unique vocabulary with implications including sentiment analysis .
Approach: They propose a predominantly German financial corpus comprising 12.5k PDF documents . they hope it will fill this gap and foster further research in the financial domain .
Outcome: The proposed corpus is the first non-email German financial corpus available . it aims to provide insights into financial discourse in the German language and multilingually.
A French Corpus and Annotation Schema for Named Entity Recognition and Relation Extraction of Financial News (2020.lrec-1)

Copied to clipboard

Challenge: Strict regulatory regimes mandate financial institutions to rigorously monitor their customers' financial activities.
Approach: They propose to use an ontology of compliance-related concepts and relationships along with a corpus annotated according to it to train and evaluate named entity recognition algorithms.
Outcome: The proposed ontology allows for training and evaluating domain-specific named entity recognition and relation extraction algorithms.
Proceedings of the Second Workshop on Economics and Natural Language Processing (D19-51)

Copied to clipboard

Challenge: ECONLP 2019 will focus on the many ways natural language processing influences business relations and procedures .
Approach: a talk will discuss use-cases of natural language processing to aid in regulatory workflows . a workshop will focus on the many ways how NLP influences business relations and procedures .
Outcome: This talk covers use-cases of natural language processing to aid in regulatory workflows . it also discusses shortcomings of current NLP technologies for financial regulation .
A Graph-Based Method for Unsupervised Knowledge Discovery from Financial Texts (2022.lrec-1)

Copied to clipboard

Challenge: A financial analyst's work involves manually reviewing lengthy filings and financial news articles in order to extract relevant pieces of information.
Approach: They propose an end-to-end, fully unsupervised method for knowledge discovery from financial texts that integrates existing resources to construct a knowledge graph of companies and related entities.
Outcome: The proposed method calculates the environmental rating for companies in the S&P 500 based on company filings with the SEC and provides an independent assessment of its outputs with an independent MSCI source.
SEDAR: a Large Scale French-English Financial Domain Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches for neural machine translation use small amount of data or monolingual data.
Approach: They describe acquisition, preprocessing and characteristics of a large English-French parallel corpus for the financial domain.
Outcome: The proposed corpus contains 8.6 million high quality sentence pairs . the first release of the corpus is available on github.
FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosure (2026.acl-demo)

Copied to clipboard

Challenge: FinReporting is an agentic workflow for localized cross-jurisdiction financial reporting . existing approaches assume a single-market setting and overlook structural differences across jurisdictions .
Approach: They propose a workflow that decomposes financial reporting into auditable stages . they use Large Language Models to extract and summarize corporate disclosures .
Outcome: The proposed system decomposes reporting into auditable stages . it improves consistency and reliability under heterogeneous reporting regimes.
FinQA: A Dataset of Numerical Reasoning over Financial Data (2021.emnlp-main)

Copied to clipboard

Challenge: Popular, large, pre-trained models fall far short of expert humans in acquiring finance knowledge and in complex multi-step numerical reasoning on that knowledge.
Approach: They propose a large-scale dataset with Question-Answering pairs over financial reports written by financial experts to facilitate analytical progress.
Outcome: The proposed dataset is the first of its kind and is available on github.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations