Challenge: Recent advances in the Financial AI realm have expanded the scope of data and methods they use, such as textual and audio cues from financial earnings calls, but limitations exist.
Approach: They propose a Saliency-guided Hierarchical Mixup augmentation technique for multimodal financial prediction tasks.
Outcome: The proposed technique outperforms state-of-the-art methods by 3-7% on financial earnings and conference call datasets.

Similar Papers

DocFin: Multimodal Financial Prediction and Bias Mitigation using Semi-structured Documents (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing research focuses on textual and audio modalities of financial disclosures but ignores the rich tabular data available in financial reports.
Approach: They propose to combine tabular financial data with text transcripts and audio recordings to improve stock volatility and price movement prediction by 5-12% and reduce gender bias by over 30%.
Outcome: The combined data improves stock volatility and price movement prediction by 5-12% and reduces gender bias caused due to audio-based neural networks by over 30%.
Financial Forecasting from Textual and Tabular Time Series (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models that combine multiple data sources and combine them to form accurate financial predictions are challenging to model without inductive biases.
Approach: They propose to use numerical financial results, macroeconomic states, and long financial documents to model company earnings relative to analyst expectations.
Outcome: The proposed model outperforms existing models in a simulated trading environment and demonstrates that each modality contains unique information.
Bridging the Data Gap in Financial Sentiment: LLM-Driven Augmentation (2025.acl-srw)

Copied to clipboard

Challenge: Existing datasets that are outdated and inaccurate hinder accuracy of Financial Sentiment Analysis (FSA) .
Approach: They propose a data augmentation technique using Retrieval Augmented Generation (RAG) to infuse established benchmarks with up-to-date contextual information from contemporary financial news.
Outcome: The proposed method modernizes established benchmarks with up-to-date contextual information while addressing class imbalances.
SSMix: Saliency-Based Span Mixup for Text Classification (2021.findings-acl)

Copied to clipboard

Challenge: SSMix synthesizes a sentence while preserving the locality of two original texts by span-based mixing and keeping more tokens related to the prediction relying on saliency information.
Approach: They propose a new method where the operation is performed on input text rather than on hidden vectors like previous approaches.
Outcome: The proposed method outperforms hidden-level mixup methods on a wide range of text classification benchmarks including textual entailment, sentiment classification, and questiontype classification.
SAE-FiRE: Enhancing Earnings Surprise Predictions Through Sparse Autoencoder Feature Selection (2026.findings-acl)

Copied to clipboard

Challenge: Conference call transcripts contain significant redundancy and industry-specific terminology that creates obstacles for language models.
Approach: They propose a Sparse Autoencoder for Financial Representation Enhancement framework to extract key information from earnings conference call transcripts and eliminate redundancy.
Outcome: The proposed method outperforms baselines in analyzing earnings conference call transcripts.
HypMix: Hyperbolic Interpolative Data Augmentation (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for data augmentation involve performing mathematical operations over the raw input samples or their latent states representations, but these operations are performed in the Euclidean space, simplifying these representations and resulting in noisy interpolations.
Approach: They propose a model-, data-, and modality-agnostic interpolative data augmentation technique operating in the hyperbolic space that captures the complex geometry of input and hidden state hierarchies better than its contemporaries.
Outcome: The proposed technique outperforms state-of-the-art methods on benchmark and low resource datasets across speech, text, and vision modalities.
VolTAGE: Volatility Forecasting via Text Audio Fusion with Graph Convolution Networks for Earnings Calls (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to stock volatility forecasting ignore correlations between stocks.
Approach: They propose to combine vocal cues with verbal and financial cue data to create a multimodal stock volatility prediction model that accounts for stock interdependence via graph convolutions.
Outcome: The proposed model outperforms existing methods showing that it can predict volatility using multimodal learning.
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction (2026.acl-long)

Copied to clipboard

Challenge: Existing models for earnings surprise prediction rely on expensive, proprietary data.
Approach: They propose to use textual transcripts and audio recordings to build a dataset for earnings surprise prediction.
Outcome: The proposed dataset includes 2,688 unique conference calls from 2019 to 2021.
An Empirical Investigation of Bias in the Multimodal Analysis of Financial Earnings Calls (2021.naacl-main)

Copied to clipboard

Challenge: Existing research focuses on textual elements of financial disclosures but ignores the rich acoustic features in the executives’ speech.
Approach: They propose to use a multimodal approach that leverages the verbal and vocal cues of speakers in financial disclosures to predict volatility and risk.
Outcome: The proposed models outperform existing models in the financial realm but still underrepresent the diverse communities spanning demographics, gender, and native speech.
A Unified Framework for Modeling Heterogeneous Financial Data via Dual-Granularity Prompting (2026.acl-industry)

Copied to clipboard

Challenge: Recent industrial credit scoring models rely heavily on manually tuned statistical learning methods due to the complexity of heterogeneous financial data and the challenge of modeling evolving creditworthiness.
Approach: They propose a framework that reformulates credit scoring as a multi-scale sequential learning problem.
Outcome: FinLangNet improves KS and bad debt rate by 6.3 pp in real world deployments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations