Papers by Manish Nagireddy

6 papers
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions (2024.findings-emnlp)

Copied to clipboard

Challenge: Modern language models exhibit some inherent shortcomings, particularly in conversational settings.
Approach: They propose a set of maxims for describing effective human-AI conversation that include quantity, quality, relevance, manner, benevolence, and transparency.
Outcome: The proposed maxims are applied to human-AI interactions and are based on extensive research from the social science and AI communities.
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models have been shown to have worse abstention abilities than reasoning models . a new class of abstraction methods is developed to improve absttention performance .
Approach: They propose a class of abstention methods that generate reasoning trace and reconstruct most likely query from it.
Outcome: The proposed method beats baselines in 33 out of 36 settings.
DAMAGeR: Deploying Automatic and Manual Approaches to GenAI Red-teaming (2025.naacl-tutorial)

Copied to clipboard

Challenge: In this tutorial, we will review and apply current automatic and manual red-teaming techniques for GenAI models.
Approach: This tutorial will review automatic and manual red-teaming techniques for GenAI models .
Outcome: This tutorial will review and apply current automatic and manual red-teaming techniques for GenAI models.
Value Alignment from Unstructured Text (2024.emnlp-industry)

Copied to clipboard

Challenge: Currently, alignment of large language models to value systems relies on the availability of supervised and preference data.
Approach: They propose a systematic approach for aligning large language models to values in unstructured text data using synthetic data generation techniques.
Outcome: The proposed approach shows improved performance on the Mistral-7B-Instruct model compared to other approaches, as quantified through the use of automatic metrics and win rates.
Granite Guardian: Comprehensive LLM Safeguarding (2025.naacl-industry)

Copied to clipboard

Challenge: a suite of advanced models is designed to detect and mitigate risks associated with prompts and responses.
Approach: a team of researchers develop a model family to detect and mitigate risks associated with prompts and responses. the model family is based on the Granite 3.0 language models.
Outcome: a new model family is designed to detect and mitigate risks associated with prompts and responses.
Multi-Level Explanations for Generative Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are being used for context-grounded tasks like summarizing meetings and answering doctors' questions.
Approach: They propose a technique to provide explanations for context-grounded text generation by assigning scores to parts of the context to quantify their influence on the model output.
Outcome: The proposed framework can provide more faithful explanations of generated output than available alternatives, including LLM self-explanations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations