Papers by Prannaya Gupta

2 papers
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk Appetite? (2025.emnlp-industry)

Copied to clipboard

Challenge: Our analysis was conducted on proprietary systems and open-weight models . FINRISKEVAL analyzed 1,720 profiles spanning a broad spectrum of possible risk categories .
Approach: They evaluated proprietary AI systems and open-weight models to assess investment risk appetite using carefully curated user profiles.
Outcome: The proposed models exhibit significant variance when user attributes that should not influence risk computation are changed.
WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models (2024.emnlp-demo)

Copied to clipboard

Challenge: Potential harms include training data leakage, biases in responses and decision-making, and unauthorized use for purposes such as terrorism and the generation of sexually explicit content.
Approach: WalledEval is a comprehensive AI safety testing toolkit designed to evaluate large language models.
Outcome: The framework supports both LLM and judge benchmarking and incorporates custom mutators to test safety against various text-style mutations such as future tense and paraphrasing.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations