Papers by Santu Karmaker

6 papers
The Path Not Taken: Duality in Reasoning about Program Execution (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on predicting program properties tied to specific inputs, resulting in a narrow view of dynamic code reasoning and data contamination.
Approach: They instantiate dual-path reasoning in a benchmark and evaluate 13 LLMs.
Outcome: The proposed model provides a robust and discriminative proxy for dynamic code understanding.
Teaching an Old LLM Secure Coding: Localized Preference Optimization on Distilled Preferences (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to improve security of LLM generated code are ineffective and lack localized regions of code.
Approach: They propose a method for distilling a preference dataset of insecure and secure code pairs from frontier LLMs and a security reasoning that explains the issues and the fix.
Outcome: The proposed method reduces code insecurity while improving overall code quality.
Benchmarking LLMs on Semantic Overlap Summarization (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are the most capable text generation models in a variety of tasks and fields.
Approach: They benchmark Large Language Models (LLMs) on SOS and introduce PrivacyPolicyPairs (3P) a dataset of 135 high-quality privacy policy documents is used to evaluate the model.
Outcome: The proposed dataset complements existing resources and broadens domain coverage.
ALIGN-SIM: A Task-Free Test Bed for Evaluating and Interpreting Sentence Embeddings through Semantic Similarity Alignment (2024.findings-emnlp)

Copied to clipboard

Challenge: Sentence embeddings play a pivotal role in a wide range of NLP tasks . evaluating and interpreting these dense vectors remains an open challenge to date .
Approach: They propose a task-free test bed for evaluating and interpreting sentence embeddings . they examined five classical and eight LLM-induced sentence embedders based on semantic similarity alignment criteria .
Outcome: The proposed test bed consists of five semantic similarity alignment criteria . it shows that none of the embeddings aligned with the criteria compared to other benchmarks .
LLMs as Meta-Reviewers’ Assistants: A Case Study (2025.naacl-long)

Copied to clipboard

Challenge: Meta-reviews are a critical step in the overall scientific peer-reviewed process, which focuses on understanding the consensus of expert opinions on a scholarly work and making informed judgments on its scientific merit.
Approach: They propose to use large language models to generate a controlled multi-perspective-summary (MPS) of their opinions to help meta-reviewers better comprehend multiple experts' perspectives.
Outcome: The proposed model can help meta-reviewers better comprehend multiple experts’ perspectives by generating a controlled multi-perspective-summary (MPS) of their opinions.
Large Language Models for IT Automation Tasks: Are We There Yet? (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks rely on synthetic tasks that fail to capture the needs of practitioners who use IT automation tools.
Approach: They evaluate 14 open-source and 3 proprietary LLMs and find that GPT-4.1-Mini achieves the best pass@10 rate of 23.9%, while Claude-3.5-Sonnet achieves best pass @1 performance.
Outcome: The evaluated LLMs perform poorly in 126 tasks and show that they lack state reconciliation capabilities and lack module knowledge.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations