Papers by Chung-Chi Chen
From Facts to Insights: A Study on the Generation and Evaluation of Analytical Reports for Deciphering Earnings Calls (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies have focused on the generation and evaluation of analytical reports derived from Earnings Calls (ECs). |
| Approach: | They propose to use Large Language Models to generate and evaluate analytical reports derived from Earnings Calls (ECs) they propose to introduce specialized agents that introduce diverse viewpoints and desirable topics into the report generation process. |
| Outcome: | The proposed model improves the quality of reports in different settings, while human-written reports remain preferred in the majority of cases. |
Financial Opinion Mining (2021.emnlp-tutorials)
Copied to clipboard
| Challenge: | This tutorial will provide an overview of financial opinion mining and provide research directions. |
| Approach: | This tutorial will introduce financial opinion mining and examine possible research directions. |
| Outcome: | This tutorial aims to provide an overview of financial opinion mining and figure out research directions. |
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized inference across diverse natural language tasks, with larger models performing better but at higher computational costs. |
| Approach: | They propose a confidence-driven strategy that dynamically selects the most suitable model based on confidence estimates. |
| Outcome: | The proposed approach reduces token usage by approximately 60% and improves cost efficiency on the Massive Multitask Language Understanding (MMLU) benchmark. |
Issues and Perspectives from 10,000 Annotated Financial Social Media Data (2020.lrec-1)
Copied to clipboard
| Challenge: | In the NLP community, many researchers have begun to use machine learning on financial and economic data. |
| Approach: | They present a dataset with 10,000 financial tweets annotated by experts from the front desk and the middle desk in a bank’s treasury. |
| Outcome: | The annotated financial tweets of a bank's front desk and middle desk are compared against a general sentiment dictionary and a domain-specific dictionary. |
Argument-Based Sentiment Analysis on Forward-Looking Statements (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing models for argument mining are limited in interpreting future-oriented arguments. |
| Approach: | They propose a categorization of argument units into claims, premises, and scenarios coupled with a unique sentiment analysis framework. |
| Outcome: | The proposed framework outperforms existing models in most tasks and is more efficient than existing methods. |
Semantics-Preserved Data Augmentation for Aspect-Based Sentiment Analysis (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for data augmentation address data deficiencies and semantic consistency, but they ignore the second issue. |
| Approach: | They propose a semantics-preserving data augmentation approach that preserves the semantics of a textual sequence. |
| Outcome: | The proposed method achieves better performance on publicly available datasets and stock price/risk movement prediction scenarios. |
Term-Driven Forward-Looking Claim Synthesis in Earnings Calls (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing arguments synthesis models excel in summarizing arguments, but lack accurate forward-looking perspectives. |
| Approach: | They propose a task called "forward-looking claim planning" that incorporates forward-looking perspectives. |
| Outcome: | The proposed method improves the existing models and improves performance. |
ML-Promise: A Multilingual Dataset for Corporate Promise Verification (2025.emnlp-main)
Copied to clipboard
| Challenge: | Promises shape perceptions and drive decisions, but verification of their fulfillment is difficult due to complexity and volume of commitments . authors propose a new approach to verifying promises in environmental, social, and governance reports . complexity of promises, complexity of evidence, difficulty in verifying their fulfillment a pressing need for new approaches . |
| Approach: | They propose a multilingual dataset that includes English, French, Chinese, Japanese, and Korean . they propose ML-Promise to facilitate in-depth verification of corporate promises . |
| Outcome: | The proposed approach includes promise identification, evidence assessment, and evaluation of timing for verification in multiple languages. |
The Impact of Language on Arithmetic Proficiency: A Multilingual Investigation with Cross-Agent Checking Computation (2024.naacl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) have garnered significant attention over the past year . previous studies have evaluated LLMs' performance in solving math word problems, but there is little discussion on whether they comprehend the operations they generate. |
| Approach: | They challenge the notion that arithmetic is language-independent and compare models with cross-agent collaborations to find significant limitations in their performance. |
| Outcome: | The proposed model outperforms collaborative approaches in basic arithmetic tasks. |
GADFA: Generator-Assisted Decision-Focused Approach for Opinion Expressing Timing Identification (2025.coling-main)
Copied to clipboard
| Challenge: | Existing models generate text on demand, but in real-life situations, individuals do not continuously generate text or voice opinions. |
| Approach: | They propose a novel task to identify news-triggered opinion expressing timing by using a dataset generated by professional stock analysts. |
| Outcome: | The proposed model can generate opinion on stock analysts' actions and improves performance in various opinion understanding tasks. |
Enhancing Society-Undermining Disinformation Detection through Fine-Grained Sentiment Analysis Pre-Finetuning (2024.findings-eacl)
Copied to clipboard
| Challenge: | a new method for disinformation detection is needed to address the issue of disinformation, authors argue . a series of rigorous experiments establishes a notable connection between disinformation and fine-grained sentiment labels . |
| Approach: | They propose a method leveraging pre-finetuning concept for efficient detection and removal of disinformation that may undermine society. |
| Outcome: | The proposed method improves performance across languages and languages, showing promising results. |
NumHG: A Dataset for Number-Focused Headline Generation (2024.lrec-main)
Copied to clipboard
| Challenge: | a lack of fine-grained annotations for accurate numeral generation in headlines is a major roadblock . a new dataset, the NumHG, provides over 27,000 annotated numeral-rich news articles for detailed investigation . |
| Approach: | They propose a dataset that provides annotated numerals for headline generation . they evaluate five well-performing headline-generation models using human evaluation . |
| Outcome: | The proposed dataset provides annotated numeral-rich news articles for detailed investigation. |
Dynamic Graph Transformer for Implicit Tag Recognition (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing studies focus on using explicit information in articles and do not consider the implicit information. |
| Approach: | They propose a dynamic graph transformer that distills the textual information and the entity relations on the fly. |
| Outcome: | The proposed model can extract the textual information and the entity relations on the fly. |
Fidelity-Enriched Contrastive Search: Reconciling the Faithfulness-Diversity Trade-Off in Text Generation (2023.emnlp-main)
Copied to clipboard
| Challenge: | Language models often generate fluent and convincing content but can lack consistency with the provided source, resulting in potential inaccuracies. |
| Approach: | They propose a new decoding method that augments the contrastive search framework with context-aware regularization terms to promote tokens that are semantically similar to the provided source while penalizing repetitiveness in the generated text. |
| Outcome: | The proposed method improves faithfulness across various language models while maintaining output diversity comparable to well-performing decoding algorithms. |
Can GPT-4 Sway Experts’ Investment Decisions? (2025.findings-naacl)
Copied to clipboard
| Challenge: | In the post-Turing era, evaluating large language models involves assessing generated text based on readers’ decisions rather than merely its indistinguishability from human-produced content. |
| Approach: | They propose to use GPT-4 to evaluate generated text from the aspects of grammar, convincingness, logical coherence, and usefulness to determine its validity. |
| Outcome: | The proposed model can generate persuasive analyses affecting the decisions of amateurs and experts. |
Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement (2026.acl-industry)
Copied to clipboard
Wataru Hirota, Tomoki Taniguchi, Tomoko Ohkuma, Kosuke Takahashi, Takahiro Omi, Kosuke Arima, Takuto Asakura, Chung-Chi Chen, Tatsuya Ishigaki
| Challenge: | Large language models (LLMs) make it easy to generate large numbers of product ideas. |
| Approach: | They propose to use a dataset of 3,000 individual scores across 300 patent-grounded product ideas to assess whether an automatic judge approximates an aggregate consensus. |
| Outcome: | The proposed model evaluators disagree on fine-grained ordinal scores, suggesting structured heterogeneity rather than random noise. |
Learning Strategies for Robust Argument Mining: An Analysis of Variations in Language and Domain (2024.lrec-main)
Copied to clipboard
| Challenge: | Argument mining is a complex process that requires a large amount of resources and time. |
| Approach: | They propose to analyze arguments in three different languages and domains to understand their robustness to natural language variations. |
| Outcome: | The proposed systems are more robust to natural language variations than existing arguments mining systems. |
Improving Numeracy by Input Reframing and Quantitative Pre-Finetuning Task (2023.findings-eacl)
Copied to clipboard
| Challenge: | Innumeracy is a problem in pretrained language models, but it is not discussed in this paper . Numerals are an indispensable part of narratives and provide much fine-grained information. |
| Approach: | They propose a method to solve innumeracy in pretrained language models by exploring the notation of numbers. |
| Outcome: | The proposed method improves performance in three benchmark datasets containing quantitative-related tasks. |
Numeracy-600K: Learning Numeracy for Detecting Exaggerated Information in Market Comments (P19-1)
Copied to clipboard
| Challenge: | Numeracy is the ability to predict the magnitude of a numeral at some specific position in a text description. |
| Approach: | They propose to use a dataset to test whether neural network models can learn numeracy . numerability is the ability to predict the magnitude of a numeral at some specific position in a text description. |
| Outcome: | The proposed task can predict the magnitude of a numeral at a specific position in a text description. |
DBQR-QA: A Question Answering Dataset on a Hybrid of Database Querying and Reasoning (2024.findings-acl)
Copied to clipboard
| Challenge: | Question answering (QA) is a fundamental task in the field of Natural Language Processing (NLP). |
| Approach: | They propose a database querying and reasoning dataset for question answering that is designed to accommodate sequential questions and multi-hop queries. |
| Outcome: | The proposed dataset better mirrors the dynamics of real-world information retrieval and analysis with a particular focus on the financial reports of US companies. |