Papers by Sameena Shah
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Recent large language models such as ChatGPT and GPT-4 have shown exceptional capabilities of generalist models . however, their applicability and effectiveness in specific domains like finance needs a better understanding . |
| Approach: | They conduct empirical studies to compare the performance of ChatGPT and GPT-4 on financial text analytical problems using eight benchmark datasets from five categories of tasks. |
| Outcome: | The proposed models outperform the state-of-the-art models on a wide range of financial text analytical tasks. |
ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering (2022.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large pre-trained language models have brought the NLP field into a new era. |
| Approach: | They propose a large-scale dataset to study the chain of numerical reasoning in conversational question answering. |
| Outcome: | The proposed dataset should push forward the exploration of real-world, complex reasoning tasks as the next research focus. |
AliGATr: Graph-based layout generation for form understanding (2024.findings-emnlp)
Copied to clipboard
| Challenge: | State of the art forms understanding models often rely on poorly calibrated output probabilities and low performance on relation extraction tasks. |
| Approach: | They propose a graph-based model that uses a generative objective to represent complex grid-like layouts that are often found in forms. |
| Outcome: | The proposed model performs better on the KIE and RE tasks and is more accurate than existing models. |
Where is this coming from? Making groundedness count in the evaluation of Document VQA models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Document Visual Question Answering (VQA) models have come close to or matching human performance on some benchmarks. |
| Approach: | They propose a method that accounts for the semantic and multimodal groundedness of a model’s outputs and can be parameterized so that users can configure the score according to their preferences. |
| Outcome: | The proposed method produces scores that are a better indicator of a model’s robustness and tends to give higher rewards to better-calibrated answers. |
FinQA: A Dataset of Numerical Reasoning over Financial Data (2021.emnlp-main)
Copied to clipboard
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, William Yang Wang
| Challenge: | Popular, large, pre-trained models fall far short of expert humans in acquiring finance knowledge and in complex multi-step numerical reasoning on that knowledge. |
| Approach: | They propose a large-scale dataset with Question-Answering pairs over financial reports written by financial experts to facilitate analytical progress. |
| Outcome: | The proposed dataset is the first of its kind and is available on github. |
Improving compositional generalization for multi-step quantitative reasoning in question answering (2022.emnlp-main)
Copied to clipboard
| Challenge: | Quantitative reasoning is an important aspect of question answering when numeric and verbal cues interact to indicate sophisticated, multi-step programs. |
| Approach: | They propose a method that encourages QA models to adjust attention patterns and capture input/output alignments that are meaningful to the reasoning task. |
| Outcome: | The proposed approach improves program accuracy and renders models more robust against overfitting as the number of reasoning steps grows. |
Using counterfactual contrast to improve compositional generalization for multi-step quantitative reasoning (2023.acl-long)
Copied to clipboard
| Challenge: | In quantitative question answering, compositional generalization is one of the main challenges of state of the art models. |
| Approach: | They propose a method that uses counterfactual scenarios to generate samples with compositional contrast. |
| Outcome: | The proposed method improves the performance of three state of the art models on four recently released datasets and also improves OOD performance on unseen domains and unsealed compositions. |
Towards a new research agenda for multimodal enterprise document understanding: What are we missing? (2024.findings-acl)
Copied to clipboard
| Challenge: | In this paper, we discuss the limitations of multimodal document understanding models in enterprise settings. |
| Approach: | They propose a research agenda that is aimed at driving the field towards higher impact in enterprise applications. |
| Outcome: | The proposed research agenda is aimed at driving the field towards higher impact in enterprise applications. |