Papers by Sudheer Chava

11 papers
KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation (2026.acl-long)

Copied to clipboard

Challenge: KG-MulQA extracts QA pairs at multiple complexity levels along three key dimensions: multi-hop retrieval, set operations, and answer plurality.
Approach: They propose a framework that extracts QA pairs at multiple complexity levels along three key dimensions: multi-hop retrieval, set operations, and answer plurality.
Outcome: The framework extracts QA pairs at multiple complexity levels along key dimensions . it enables fine-grained assessment of model performance across controlled difficulty levels.
ConfReady: A RAG based Assistant and Dataset for Conference Checklist Responses (2025.emnlp-demos)

Copied to clipboard

Challenge: ARR Responsible NLP Research checklist is designed to encourage best practices for responsible research . previous research has shown that self-reported checklist responses don't always accurately represent papers .
Approach: They propose a retrieval-augmented generation application that can be used to assist authors with conference checklists.
Outcome: The proposed application can be used to help authors with conference checklists and review their work.
CoCoHD: Congress Committee Hearing Dataset (2024.findings-emnlp)

Copied to clipboard

Challenge: Congressional hearings are crucial tools for both political parties to advance their agendas.
Approach: They propose a dataset covering congressional hearings from 1997 to 2024 across 86 committees, with 32,697 records.
Outcome: The proposed dataset covers hearings from 1997 to 2024 across 86 committees, with 32,697 records.
Saliency-Aware Interpolative Augmentation for Multimodal Financial Prediction (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in the Financial AI realm have expanded the scope of data and methods they use, such as textual and audio cues from financial earnings calls, but limitations exist.
Approach: They propose a Saliency-guided Hierarchical Mixup augmentation technique for multimodal financial prediction tasks.
Outcome: The proposed technique outperforms state-of-the-art methods by 3-7% on financial earnings and conference call datasets.
When FLUE Meets FLANG: Benchmarks and Large Pretrained Language Model for Financial Domain (2022.emnlp-main)

Copied to clipboard

Challenge: Pre-trained language models have shown impressive performance on a variety of tasks and domains.
Approach: They propose a domain specific financial LANGuage model which uses financial keywords and phrases for better masking.
Outcome: The proposed model outperforms existing models on a variety of tasks and domains.
Financial Language Model Evaluation (FLaME) (2025.findings-acl)

Copied to clipboard

Challenge: Language Models (LMs) have demonstrated impressive capabilities with core NLP tasks in finance, but their effectiveness is difficult to assess due to gaps in evaluation methodologies.
Approach: They propose to use a framework to evaluate language models against ‘reasoning-reinforced’ LMs to measure their performance on finance NLP tasks.
Outcome: The proposed frameworks are open-source and provide data and data for the study.
How Inclusively do LMs Perceive Social and Moral Norms? (2025.findings-naacl)

Copied to clipboard

Challenge: Language models (LMs) are used in decision-making systems and as interactive assistants.
Approach: They propose to prompt 11 LMs on rules-of-thumb and compare their outputs with 100 human annotators.
Outcome: The proposed model is compared with 100 human annotators to find out if they are inclusive of diverse human values.
Cryptocurrency Bubble Detection: A New Stock Market Dataset, Financial Task & Hyperbolic Models (2022.naacl-main)

Copied to clipboard

Challenge: speculative trading of highly volatile assets such as cryptocurrencies and meme stocks presents a new challenge in the financial realm.
Approach: They propose a multi-span bubble detection task based on social media hype and a set of sequence-to-sequence hyperbolic models . they use data from 9 exchanges over five years to test their models based upon the power-law dynamics of cryptocurrencies and user behavior on social networks.
Outcome: The proposed model is able to detect bubbles on a set of reddit and twitter posts spanning over two million tweets over five years .
Trillion Dollar Words: A New Financial Dataset, Task & Market Analysis (2023.acl-long)

Copied to clipboard

Challenge: a study of FOMC pronouncements shows how important the FOMC communications are . hawkish-dovish classification is difficult because of the negative connotations of words .
Approach: They propose to use a dataset to classify FOMC monetary policy stances . they construct a measure of monetary stance for the FOMC document release days .
Outcome: The proposed model is based on a best-performing model and is available on Huggingface and GitHub under CC BY-NC 4.0 license.
HYPHEN: Hyperbolic Hawkes Attention For Text Streams (2022.acl-short)

Copied to clipboard

Challenge: Existing methods for text stream modeling ignore fine-grained timing irregularities and time-varying scale-free properties of texts.
Approach: They propose a hyperbolic Hawkes Attention Network which learns a data-driven hyperbolical space and models irregular powerlaw excitations using a Hawke's process.
Outcome: The proposed model can model online text sequences in a geometry agnostic manner.
Tweet Based Reach Aware Temporal Attention Network for NFT Valuation (2022.findings-emnlp)

Copied to clipboard

Challenge: Non-Fungible Tokens (NFTs) are a relatively unexplored class of assets due to their extremely volatile nature.
Approach: They propose a reach-aware temporal learning approach to predict future NFT trends from a dataset consisting of over 1.3 million tweets and 180 thousand NFT transactions .
Outcome: The proposed model outperforms state-of-the-art models by an average of 36% on a dataset consisting of over 1.3 million tweets and 180 thousand NFT transactions spanning over 15 NFT collections.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations