Challenge: Earnings calls are a key source of financial information about public companies. extracting information from earnings calls is difficult.
Approach: They propose to use LLMs to perform open-ended extraction from unstructured call transcripts to provide a baseline for this valuable domain through the consistent tracking of emergent KPIs.
Outcome: The proposed method provides a baseline for this valuable domain through the consistent tracking of emergent KPIs.

Similar Papers

ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call Transcripts (2022.emnlp-main)

Copied to clipboard

Challenge: ECTSum is a dataset for bullet-point summarization of earnings calls hosted by publicly traded companies.
Approach: They propose a dataset with transcripts of earnings calls and bullet point summaries derived from Reuters articles.
Outcome: The proposed dataset compares transcripts of earnings calls hosted by publicly traded companies with experts-written bullet point summaries derived from Reuters articles .
From Facts to Insights: A Study on the Generation and Evaluation of Analytical Reports for Deciphering Earnings Calls (2025.coling-main)

Copied to clipboard

Challenge: Existing studies have focused on the generation and evaluation of analytical reports derived from Earnings Calls (ECs).
Approach: They propose to use Large Language Models to generate and evaluate analytical reports derived from Earnings Calls (ECs) they propose to introduce specialized agents that introduce diverse viewpoints and desirable topics into the report generation process.
Outcome: The proposed model improves the quality of reports in different settings, while human-written reports remain preferred in the majority of cases.
Forecasting Earnings Surprises from Conference Call Transcripts (2023.findings-acl)

Copied to clipboard

Challenge: Earnings conference calls contain over 5,000 words of text and large amounts of industry jargon . this length and domain-specific language present problems for generic pretrained language models.
Approach: They propose a task of predicting earnings surprises from earnings call transcripts and propose linguistic models that use a long document dataset to test financial understanding.
Outcome: The proposed model can predict earnings surprises from earnings conference calls with reasonable accuracy and shows that it is possible to interpret the data with different interpretability methods.
ConEC: Earnings Call Dataset with Real-world Contexts for Benchmarking Contextual Speech Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on contextual speech recognition (ASR) systems focuses on recognizing words that are not frequently seen in training data, such as rare words, but word error rate on rare words remains over 20%.
Approach: They propose to use public-domain earnings calls and supplementary materials to evaluate contextual ASR approaches grounded on real-world applications.
Outcome: The proposed frameworks are noisier than artificially synthesized contexts that contain the ground truth, yet still make great room for future improvement of contextual ASR technology.
Financial Event Extraction Using Wikipedia-Based Weak Supervision (D19-51)

Copied to clipboard

Challenge: Existing methods for detecting financial and economic events from text have relied on a knowledge-base of financial events, or corresponding financial figures.
Approach: They propose to use Wikipedia sections to extract weak labels for sentences describing economic events from text.
Outcome: The proposed method can extract weak labels for sentences describing economic events from Wikipedia sentences.
SEC-FinTables: Evaluating Large Language Models for Detecting Logical Inconsistencies on Tabular Data (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are increasingly deployed in high-stakes domains where logical inconsistencies are unrecognized.
Approach: They propose a benchmarking system that decomposes inconsistency detection into granular subtasks and a protocol that decompiles it into subtask.
Outcome: The proposed model decomposes inconsistencies into subtasks and identifies them in 103,395 real-world and error-injected table instances.
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction (2026.acl-long)

Copied to clipboard

Challenge: Existing models for earnings surprise prediction rely on expensive, proprietary data.
Approach: They propose to use textual transcripts and audio recordings to build a dataset for earnings surprise prediction.
Outcome: The proposed dataset includes 2,688 unique conference calls from 2019 to 2021.
Decoding the Market’s Pulse: Context-Enriched Agentic Retrieval Augmented Generation for Predicting Post-Earnings Price Shocks (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for forecasting large stock price movements after corporate earnings calls are prone to **narrative bias** Existing approaches lack temporal-causal reasoning and are unable to predict large stock prices.
Approach: They propose a retrieval-augmented framework that deploys a team of cooperative LLM agents . they retrieve structured evidence from a Causal-Temporal Knowledge Graph built from financial statements and earnings calls .
Outcome: The proposed framework outperforms larger LLMs and fine-tuned models in macro-F1, MCC, and Sharpe for the same forecasting horizon.
Crossing Domains without Labels: Distant Supervision for Term Extraction (2025.emnlp-industry)

Copied to clipboard

Challenge: Current state-of-the-art methods require expensive human annotation and struggle with domain transfer, limiting their practical deployment.
Approach: They propose a benchmark spanning seven diverse domains to evaluate ATE performance . they propose psuedo-labels and post-hoc heuristics to ensure generalizability .
Outcome: The proposed model outperforms supervised cross-domain encoder models and few-shot learning baselines on the document- and corpus-levels and its GPT-4o teacher on the benchmark.
From Annotation to Adaptation: Metrics, Synthetic Data, and Aspect Extraction for Aspect-Based Sentiment Analysis with Large Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Using a synthetic sports feedback dataset, we evaluate open-weight LLMs’ ability to extract aspect-polarity pairs.
Approach: They propose a metric to facilitate the evaluation of aspect extraction with generative models.
Outcome: The proposed metric improves the performance of open-weight LLMs in the Aspect-Based Sentiment Analysis task.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations