Papers by Shramay Palta

4 papers
It’s Not Easy Being Wrong: Large Language Models Struggle with Process of Elimination Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Recent research aims to unlock the reasoning capabilities of large language models (LLMs) chain-of-thought (COT) prompting can help LLMs reason toward correct answers, but its efficacy in reasoning toward incorrect answers is unexplored.
Approach: They propose a task where large language models reason toward incorrect answers using chain-of-thought prompting.
Outcome: The proposed task underperforms the strategy of choosing the correct answer on commonsense and scientific reasoning datasets.
Arguments that Alter Minds: LLM Rationales Sway Human (and LLM) Notions of Plausibility (2026.acl-long)

Copied to clipboard

Challenge: Experiments with LLMs reveal similar patterns of influence on human plausibility judgments of commonsense benchmark answers.
Approach: They find that human plausibility judgments of commonsense benchmark answers are affected by implausibility arguments for or against an answer.
Outcome: The results show that human judges find LLM rationales convincing and that human annotators agree on the most plausible answer when the plausibility gap is wide.
FORK: A Bite-Sized Test Set for Probing Culinary Cultural Biases in Commonsense Reasoning Models (2023.findings-acl)

Copied to clipboard

Challenge: a recent study shows that commonsense knowledge is universally shared by most people . early efforts to schematize commonsensical knowledge as scripts provide examples of unintended biases .
Approach: They propose a set of questions for probing cultural biases and assumptions in commonsense reasoning systems . they test commonsensibleQA-style questions on food-related customs in the u.s.
Outcome: The proposed questions show that they are better at detecting biases in commonsense reasoning systems than on non-US cultures.
Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Many commonsense reasoning questions require a hard selection of a single correct answer . ambiguity and semantic mismatches are common in many MCQs .
Approach: They collect plausibility judgments on 5 000 commonsense reasoning questions . they find that the answer rated most plausible does not match the benchmark gold answers .
Outcome: Experiments with LLMS reveal low accuracy and high variation in performance on the subset . high plausibility rating for the most plausible answer is highlighted in bold .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations