Papers by Manish Nagireddy
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions (2024.findings-emnlp)
Copied to clipboard
Erik Miehling, Manish Nagireddy, Prasanna Sattigeri, Elizabeth Daly, David Piorkowski, John Richards
| Challenge: | Modern language models exhibit some inherent shortcomings, particularly in conversational settings. |
| Approach: | They propose a set of maxims for describing effective human-AI conversation that include quantity, quality, relevance, manner, benevolence, and transparency. |
| Outcome: | The proposed maxims are applied to human-AI interactions and are based on extensive research from the social science and AI communities. |
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models have been shown to have worse abstention abilities than reasoning models . a new class of abstraction methods is developed to improve absttention performance . |
| Approach: | They propose a class of abstention methods that generate reasoning trace and reconstruct most likely query from it. |
| Outcome: | The proposed method beats baselines in 33 out of 36 settings. |
DAMAGeR: Deploying Automatic and Manual Approaches to GenAI Red-teaming (2025.naacl-tutorial)
Copied to clipboard
| Challenge: | In this tutorial, we will review and apply current automatic and manual red-teaming techniques for GenAI models. |
| Approach: | This tutorial will review automatic and manual red-teaming techniques for GenAI models . |
| Outcome: | This tutorial will review and apply current automatic and manual red-teaming techniques for GenAI models. |
Value Alignment from Unstructured Text (2024.emnlp-industry)
Copied to clipboard
Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Manish Nagireddy, Pierre Dognin, Kush Varshney
| Challenge: | Currently, alignment of large language models to value systems relies on the availability of supervised and preference data. |
| Approach: | They propose a systematic approach for aligning large language models to values in unstructured text data using synthetic data generation techniques. |
| Outcome: | The proposed approach shows improved performance on the Mistral-7B-Instruct model compared to other approaches, as quantified through the use of automatic metrics and win rates. |
Granite Guardian: Comprehensive LLM Safeguarding (2025.naacl-industry)
Copied to clipboard
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling, Martín Santillán Cooper, Kieran Fraser, Giulio Zizzo, Muhammad Zaid Hameed, Mark Purcell, Michael Desmond, Qian Pan, Inge Vejsbjerg, Elizabeth M. Daly, Michael Hind, Werner Geyer, Ambrish Rawat, Kush R. Varshney, Prasanna Sattigeri
| Challenge: | a suite of advanced models is designed to detect and mitigate risks associated with prompts and responses. |
| Approach: | a team of researchers develop a model family to detect and mitigate risks associated with prompts and responses. the model family is based on the Granite 3.0 language models. |
| Outcome: | a new model family is designed to detect and mitigate risks associated with prompts and responses. |
Multi-Level Explanations for Generative Language Models (2025.acl-long)
Copied to clipboard
Lucas Monteiro Paes, Dennis Wei, Hyo Jin Do, Hendrik Strobelt, Ronny Luss, Amit Dhurandhar, Manish Nagireddy, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri, Werner Geyer, Soumya Ghosh
| Challenge: | Large language models (LLMs) are being used for context-grounded tasks like summarizing meetings and answering doctors' questions. |
| Approach: | They propose a technique to provide explanations for context-grounded text generation by assigning scores to parts of the context to quantify their influence on the model output. |
| Outcome: | The proposed framework can provide more faithful explanations of generated output than available alternatives, including LLM self-explanations. |