Papers by Tanmay Rajpurohit

7 papers
QualEval: Qualitative Evaluation for Model Improvement (2024.naacl-long)

Copied to clipboard

Challenge: Quantitative evaluation metrics are inadequate for large language models due to complexity of tasks and cannot provide actionable diagnostics.
Approach: They propose a quantitative evaluation tool called QualEval that uses automated qualitative evaluation as a vehicle for model improvement.
Outcome: The proposed method improves the performance of the Llama 2 model by 15% compared to baselines.
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation (2025.acl-long)

Copied to clipboard

Challenge: blatant lying or unintentional hallucination are common in large language models.
Approach: They build a testbed mimicking a legislative environment where a corporate lobbyist module is proposing amendments to bills that benefit a specific company while evading identification by strong LLM detectors.
Outcome: The proposed model can be used to detect deception in legislative environments and to optimize its phrasing to avoid detection by strong detectors.
Toxicity in chatgpt: Analyzing persona-assigned language models (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have shown incredible capabilities and transcended the natural language processing community.
Approach: They evaluate toxicity in over half a million generations of ChatGPT by assigning it a persona . they find that outputs engage in incorrect stereotypes, harmful dialogue, hurtful opinions .
Outcome: a new study shows that assigning a persona to a chatbot can increase toxicity in half a million generations.
Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches for distilling large language models into smaller, more efficient student models are based on educational science principles such as knowledge tracing and personalized learning.
Approach: They propose a method for distilling large language models into smaller, more efficient student models that are aligned with educational science principles such as knowledge tracing and personalized learning.
Outcome: The proposed approach outperforms LLMs on three benchmarks while employing significantly fewer parameters.
LILA: A Unified Benchmark for Mathematical Reasoning (2022.emnlp-main)

Copied to clipboard

Challenge: Towards evaluating and improving AI systems in this domain, we propose a mathematical reasoning benchmark based on 23 diversetasks .
Approach: They propose a mathematical reasoning benchmark that includes 23 diverse tasks . they extend the benchmark by collecting task instructions and solutions in the form of Python programs .
Outcome: The proposed model improves on multi-tasking while the best performing model only achieves 60.40%.
PersonaGym: Evaluating Persona Agents and LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Persona agents are LLM agents conditioned to act according to an assigned persona . evaluating how faithfully these agents adhere to their personas remains a challenge .
Approach: a new study evaluates persona agents' ability to act according to an assigned persona . a persona agent's person score is a human-aligned automatic metric that can be used to evaluate a model .
Outcome: a new evaluation framework and a human-aligned automatic metric show that persona agents can perform better.
C-STS: Conditional Semantic Textual Similarity (2023.emnlp-main)

Copied to clipboard

Challenge: Semantic textual similarity (STS) is a cornerstone task in natural language processing, but it is inherently ambiguous.
Approach: They propose a task called conditional STS which measures similarity conditioned on an aspect elucidated in natural language.
Outcome: The proposed task reduces subjectivity and ambiguity and enables fine-grained similarity evaluation using diverse conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations