Papers by Arindam Mitra

6 papers
Combining Knowledge Hunting and Neural Language Models to Solve the Winograd Schema Challenge (P19-1)

Copied to clipboard

Challenge: Existing methods to solve Winograd Schema Challenge use only knowledge embedded in text . this limits the performance of such models on the WSC problems.
Approach: They propose to augment existing language models with a commonsense knowledge hunting module and an explicit reasoning module to extract the needed knowledge from text.
Outcome: The proposed system improves on the language model based methods by 5.53% and 7.7% on the dataset.
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents (2025.findings-acl)

Copied to clipboard

Challenge: Recent success in large multimodal models (LMMs) has sparked promising applications of agents capable of autonomously completing complex web tasks.
Approach: They propose a scalable recipe to synthesize the largest and most diverse trajectory-level dataset to date.
Outcome: The proposed model synthesizes the largest and most diverse trajectory-level dataset to date, with 94K successful multimodal web trajectories, 720K screenshots, and 33M web elements.
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing work investigating the logical reasoning ability of large language models has focused only on a couple of inference rules of propositional and first-order logics.
Approach: They propose to use a natural language question-answering dataset to evaluate the logical reasoning ability of large language models.
Outcome: The proposed model performs poorly on a range of natural language questions using chain-of-thought prompting.
Careful Selection of Knowledge to Solve Open Book Question Answering (P19-1)

Copied to clipboard

Challenge: Open book question answering requires deeper reasoning involving linguistic understanding and common knowledge.
Approach: They propose a dataset that mimics open book question answering to achieve 72.0% accuracy.
Outcome: The proposed dataset achieves 72.0% accuracy, an 11.6% improvement over the current state of the art.
NumGLUE: A Suite of Fundamental yet Challenging Mathematical Reasoning Tasks (2022.acl-long)

Copied to clipboard

Challenge: Existing AI systems fail to perform basic mathematical reasoning when presented in a slightly different manner.
Approach: They propose a multi-task benchmark that evaluates the performance of AI systems on eight different tasks that at their core require simple arithmetic understanding.
Outcome: The proposed benchmark compares the performance of AI systems on eight different tasks.
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter? (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies evaluate only the final predicted answer of a puzzle, without providing any finer metrics to evaluate them.
Approach: They propose to use a grid-based evaluation dataset to evaluate LLMs' reasoning abilities and a new error taxonomy to evaluate their reasoning chains.
Outcome: The proposed model outperforms existing prompting methods on a wide range of natural language understanding tasks previously thought to be exclusive to humans.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations