Papers by Megha Chakraborty

8 papers
FACTIFY3M: A benchmark for multimodal fact verification with explainability through 5W Question-Answering (2023.emnlp-main)

Copied to clipboard

Challenge: Disinformation can cause disruption in the share market, panic and anxiety in society, and even death during crises.
Approach: a new dataset is being developed to help combat disinformation . the dataset is a multimodal fake news dataset with 5W question-answering .
Outcome: FACTIFY 3M is the largest dataset and benchmark for multimodal fact verification.
Parallel Communities Across the Surface Web and the Dark Web (2025.findings-emnlp)

Copied to clipboard

Challenge: Sense of Community is a social motivation that is reflected in the social behavior of humans.
Approach: They compile a large collection of parallel community datasets comprising over 7 million posts and comments from Reddit and 200,000 posts and comment from Dread, a dark web discussion forum, covering similar topics.
Outcome: The results show that users on Reddit exhibit a stronger sense of community membership despite the dark web’s restricted accessibility.
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking (2023.findings-emnlp)

Copied to clipboard

Challenge: Social media posts are noisy and pervasive, resulting in difficult to identify precise and prominent claims that require verification.
Approach: They propose a task called Claim Normalization that decomposes complex and noisy social media posts into more straightforward and understandable forms, termed normalized claims.
Outcome: The proposed model outperforms baselines across evaluation measures and errors.
Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI) (2023.emnlp-main)

Copied to clipboard

Challenge: a number of issues have arisen regarding the risk and consequences of AI-generated text detection.
Approach: They propose a counter-turing test to evaluate the robustness of existing AGTD methods . they propose ADI, a quantifiable spectrum to assess detectability of LLMs .
Outcome: The proposed method evaluates the robustness of existing AGTD methods . it shows that larger LLMs tend to have lower ADI, indicating they are less detectable .
LESA: Linguistic Encapsulation and Semantic Amalgamation Based Generalised Claim Detection from Online Content (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on claim detection is built on the basis of a 'segregation' of claims across different domains.
Approach: They propose a generalized generalized model that captures syntactic features through part-of-speech and dependency embeddings, as well as contextual features through a fine-tuned language model.
Outcome: The proposed model outperforms baselines on six claim datasets by an average of 3 claim-F1 points and 2 claim-f1 points on the general-domain experiments.
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection (2025.emnlp-main)

Copied to clipboard

Challenge: a survey examines the interplay between factual accuracy and cognitive biases . misinformation is more than just the existence of incorrect information, it also entails complex relationships between the information and the entities that consume it.
Approach: They examine the interplay between traditional fact-checking and psychological concepts such as cognitive biases, social dynamics, and emotional responses.
Outcome: The findings highlight limitations of current methods and identify opportunities for improvement . they also outline future research directions to create more robust frameworks .
Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter (2022.emnlp-main)

Copied to clipboard

Challenge: Current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of misinformation and fake news.
Approach: They propose a large-scale Twitter corpus with token-level claim spans on more than 7.5k tweets and a model that automatically detects and extracts the snippets of misinformation.
Outcome: The proposed model outperforms baseline systems on several evaluation metrics, improving by 1.5 points.
FACTIFY-5WQA: 5W Aspect-based Fact Verification through Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Contemporary fact-checking systems focus on estimating truthfulness using numerical scores which are not human-interpretable.
Approach: They propose a 5W framework for question-answer-based fact explainability that can assist human fact-checkers in asking relevant questions . they propose masked language model which generates QA pairs for claims and a baseline QA system that automatically locates those answers from evidence documents.
Outcome: The proposed framework can assist human fact-checkers in asking relevant questions related to a fact, which can then be validated separately to reach a final verdict.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations