Challenge: Existing automated fact-checking systems are predominantly developed for English . Existing systems focus on claim verification, but UrduFactQA targets factuality .
Approach: They propose two hand-annotated benchmarks to enable fact-checking and factual consistency evaluation in Urdu.
Outcome: The proposed benchmarks are the first of their kind for Urdu and are available online.

Similar Papers

OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) generate naturallysounding answers over a broad range of human inquiries, but they still produce content that deviates from real-world facts.
Approach: They propose a framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a document.
Outcome: The proposed framework assesses the factuality of free-form responses in open domains and evaluates factually of LLMs.
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs (2024.emnlp-demo)

Copied to clipboard

Challenge: Large language models (LLMs) often produce content that deviates from real-world facts.
Approach: They developed a unified framework to assess the factuality of large language models . open-sourced framework is publicly available as a Python library and web service .
Outcome: OpenFactCheck is open-sourced and publicly released as a Python library and web service.
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models generate naturally sounding answers over a broad range of human inquiries, but they often generate answers that contradict real-world facts.
Approach: They propose a framework for annotating and evaluating the factuality of large language models . they propose 'factcheck-bench' which provides a multi-stage annotation scheme .
Outcome: The proposed framework outperforms several popular LLM fact-checkers in claim, sentence, and document levels.
FactSearch: An Interactive Agentic Fact Search System for Verifying Large Language Model Outputs (2026.acl-demo)

Copied to clipboard

Challenge: Existing tool-augmented verification systems depend on opaque search APIs, introducing uncontrolled variability into factuality evaluation.
Approach: They propose a reproducibility-oriented agentic fact search system for claim-level factuality verification built on a locally aggregated open-source search infrastructure.
Outcome: The proposed system decomposes model outputs into atomic factual claims, generates targeted search queries, retrieves supporting evidence via a self-hosted meta-search engine, and performs modular verification within a fully configurable pipeline.
Large Language Models Require Curated Context for Reliable Political Fact-Checking—Even with Reasoning and Web Search (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results.
Approach: They evaluate 15 large language models on 6,000 claims fact-checked by PolitiFact . standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains .
Outcome: The models predict claim veracity and a curated RAG system improved macro F1 by 233% on average across model variants.
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences.
Approach: They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention.
Outcome: The proposed model improves accuracy and robustness even with a small model.
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs).
Approach: They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs.
Outcome: The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints.
FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior evaluation pipelines fail to evaluate factuality of long-form LLMs due to inefficiency and costly human assessment.
Approach: They propose a fast and strong evaluation pipeline that can evaluate factuality of long-form LLMs . they propose 'faStFact' to reduce cost of web searching and inference calling .
Outcome: The proposed evaluation pipeline achieves highest alignment with human evaluation and efficiency among existing baselines.
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing fact-checking evaluation methods rely on static datasets and classification metrics, which fail to evaluate justification production and uncover the nuanced limitations of LLMs.
Approach: They propose a framework that adaptively and dynamically assesses LLMs’ fact-checking capabilities by incorporating justification production alongside verdict prediction.
Outcome: Experiments show that the framework differentiates among state-of-the-art LLMs, providing valuable insights into model strengths and limitations in model-centric fact-checking analysis.
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for fact-checking text generated by large language models are expensive and time-consuming.
Approach: They propose a plug-and-play framework that harnesses large language models for efficient fact-checking in a few-shot manner.
Outcome: The proposed framework is compared with state-of-the-art models and shows that it can be used to speed up fact-checking in a few-shot manner.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations