Papers by Amitava Das

14 papers
FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Disinformation can cause disruption in the share market, panic and anxiety in society, and even death during crises.
Approach: a new dataset is being developed to help combat disinformation . the dataset is a multimodal fake news dataset with 5W question-answering .
Outcome: FACTIFY 3M is the largest dataset and benchmark for multimodal fact verification.
Counter Turing Test (CT2): Investigating AI-Generated Text Detection for Hindi - Ranking LLMs based on Hindi AI Detectability Index (ADI_hi) (2024.findings-emnlp)

Copied to clipboard

Challenge: a growing number of large language models are being used to detect AI-generated text . a recent study has found that some techniques to bypass detection are fragile .
Approach: They propose to use 26 LLMs to evaluate their proficiency in generating Hindi text . they propose to introduce a Hindi AI Detectability Index to assess and rank LLM models based on their detectability levels.
Outcome: The proposed methods are effective in English, but struggle in Hindi . the proposed methods show that they are susceptible to fragility .
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting (2025.coling-main)

Copied to clipboard

Challenge: Proportional analogies are used to assess linguistic and cognitive abilities.
Approach: They propose a dataset for proportional analogy completion and evaluate its performance in large-scale learning environments.
Outcome: The proposed model achieves 55% accuracy in knowledge-enhanced prompts.
Minority Positive Sampling for Switching Points - an Anecdote for the Code-Mixing Language Modeling (2020.lrec-1)

Copied to clipboard

Challenge: Multilingual people code-mix using English phonetic typing and insertion of anglicisms in their native language.
Approach: They propose to use minority positive sampling to selectively induce more sample to achieve better performance.
Outcome: The proposed model performs better than other models, but switching points are the main challenge .
YinYang-Align: A new Benchmark for Competing Objectives and Introducing Multi-Objective Preference based Text-to-Image Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Recent controversies highlight the need for robust alignment mechanisms in text-to-image systems.
Approach: They propose a framework to evaluate T2I systems across six contradictory alignment objectives . objectives highlight key trade-offs such as artistic freedom and cultural sensitivity .
Outcome: The proposed framework achieves superior alignment across all objectives.
Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI) (2023.emnlp-main)

Copied to clipboard

Challenge: a number of issues have arisen regarding the risk and consequences of AI-generated text detection.
Approach: They propose a counter-turing test to evaluate the robustness of existing AGTD methods . they propose ADI, a quantifiable spectrum to assess detectability of LLMs .
Outcome: The proposed method evaluates the robustness of existing AGTD methods . it shows that larger LLMs tend to have lower ADI, indicating they are less detectable .
On the Relationship between Sentence Analogy Identification and Sentence Structure Encoding in Large Language Models (2024.findings-eacl)

Copied to clipboard

Challenge: Analogies facilitate the transfer of meaning and knowledge from one domain to another.
Approach: They propose to use large language models to encode syntactic and semantic structures of sentences to identify sentence analogies.
Outcome: The LLMs which capture syntactic structures better, also have higher abilities in identifying sentence analogies.
DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is a cornerstone for preference alignment but is constrained by fixed divergence measures and limited feature transformations.
Approach: They propose a new enhancement of Direct Preference Optimization that integrates kernel methods to overcome these challenges.
Outcome: The proposed model improves divergence measures and features by using kernels . the proposed model achieves state-of-the-art generalization in factuality, safety, reasoning, and instruction following .
Tutorial Proposal: Hallucination in Large Language Models (2024.lrec-tutorials)

Copied to clipboard

Challenge: Grasping the intricacies of hallucination in LLMs can be daunting, especially for those new to the field.
Approach: This tutorial aims to bridge the gap between the field and the field of hallucination . it will explore the key aspects of hallucinonation, including benchmarking, detection, and mitigation techniques .
Outcome: This tutorial will explore the key aspects of hallucination in LLMs . it will also explore the specific constraints and shortcomings of current approaches .
Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations (2025.emnlp-main)

Copied to clipboard

Challenge: a new metric measures the quality of large language models (LLMs) that detects hidden misalignments and jailbreak risks.
Approach: They propose a decoding-invariant metric that measures latent safety failures . they propose 'Alignment Quality Index' to measure latent activations in latent space .
Outcome: The proposed metric detects latent safety failures overlooked by behavioral benchmarks and jailbreaks.
The Troubling Emergence of Hallucination in Large Language Models - An Extensive Definition, Quantification, and Prescriptive Remediations (2023.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models have generated widespread acclaim, but hallucination has also emerged as a by-product.
Approach: They propose a fine-grained discourse on profiling hallucination based on its degree, orientation, and category . they categorize hallucines into six types: acronym ambiguity, generated golem, virtual voice, geographic erratum, time wrap .
Outcome: The proposed method categorizes hallucination into six types based on their degree, orientation, and category .
SEPSIS: I Can Catch Your Lies – A New Paradigm for Deception Detection (2025.acl-srw)

Copied to clipboard

Challenge: a new framework categorizes deception into three forms: lies of omission, lies of commission, and lies of influence . a novel framework for deception detection leveraging NLP techniques is proposed .
Approach: They propose a framework that categorizes deception into three forms: lies of omission, lies of commission, and lies of influence.
Outcome: The proposed framework achieves an impressive F1 score of 0.87 across all layers . it can be used to investigate lies of omission, lies of commission and lies of influence .
FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Contemporary fact-checking systems focus on estimating truthfulness using numerical scores which are not human-interpretable.
Approach: They propose a 5W framework for question-answer-based fact explainability that can assist human fact-checkers in asking relevant questions . they propose masked language model which generates QA pairs for claims and a baseline QA system that automatically locates those answers from evidence documents.
Outcome: The proposed framework can assist human fact-checkers in asking relevant questions related to a fact, which can then be validated separately to reach a final verdict.
ANALOGICAL - A Novel Benchmark for Long Text Analogy Evaluation in Large Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Modern large language models are evaluated on extrinsic measures based on benchmarks such as GLUE and SuperGLUE.
Approach: They propose a benchmark to intrinsically evaluate large language models across a taxonomy of analogies of long text with six levels of complexity.
Outcome: The proposed benchmark evaluates LLMs across a taxonomy of analogies of long text with six levels of complexity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations