UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking (2025.findings-emnlp)
Copied to clipboard
Sarfraz Ahmad, Hasan Iqbal, Momina Ahsan, Numaan Naeem, Muhammad Ahsan Riaz Khan, Arham Riaz, Muhammad Arslan Manzoor, Yuxia Wang, Preslav Nakov
| Challenge: | Existing automated fact-checking systems are predominantly developed for English . Existing systems focus on claim verification, but UrduFactQA targets factuality . |
| Approach: | They propose two hand-annotated benchmarks to enable fact-checking and factual consistency evaluation in Urdu. |
| Outcome: | The proposed benchmarks are the first of their kind for Urdu and are available online. |
Similar Papers
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs (2025.coling-main)
Copied to clipboard
Yuxia Wang, Minghan Wang, Hasan Iqbal, Georgi N. Georgiev, Jiahui Geng, Iryna Gurevych, Preslav Nakov
| Challenge: | Large language models (LLMs) generate naturallysounding answers over a broad range of human inquiries, but they still produce content that deviates from real-world facts. |
| Approach: | They propose a framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a document. |
| Outcome: | The proposed framework assesses the factuality of free-form responses in open domains and evaluates factually of LLMs. |
OpenFactCheck: A Unified Framework for Factuality Evaluation of LLMs (2024.emnlp-demo)
Copied to clipboard
| Challenge: | Large language models (LLMs) often produce content that deviates from real-world facts. |
| Approach: | They developed a unified framework to assess the factuality of large language models . open-sourced framework is publicly available as a Python library and web service . |
| Outcome: | OpenFactCheck is open-sourced and publicly released as a Python library and web service. |
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers (2024.findings-emnlp)
Copied to clipboard
Yuxia Wang, Revanth Gangi Reddy, Zain Mujahid, Arnav Arora, Aleksandr Rubashevskii, Jiahui Geng, Osama Mohammed Afzal, Liangming Pan, Nadav Borenstein, Aditya Pillai, Isabelle Augenstein, Iryna Gurevych, Preslav Nakov
| Challenge: | Large language models generate naturally sounding answers over a broad range of human inquiries, but they often generate answers that contradict real-world facts. |
| Approach: | They propose a framework for annotating and evaluating the factuality of large language models . they propose 'factcheck-bench' which provides a multi-stage annotation scheme . |
| Outcome: | The proposed framework outperforms several popular LLM fact-checkers in claim, sentence, and document levels. |
FactSearch: An Interactive Agentic Fact Search System for Verifying Large Language Model Outputs (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing tool-augmented verification systems depend on opaque search APIs, introducing uncontrolled variability into factuality evaluation. |
| Approach: | They propose a reproducibility-oriented agentic fact search system for claim-level factuality verification built on a locally aggregated open-source search infrastructure. |
| Outcome: | The proposed system decomposes model outputs into atomic factual claims, generates targeted search queries, retrieves supporting evidence via a self-hosted meta-search engine, and performs modular verification within a fully configurable pipeline. |
Large Language Models Require Curated Context for Reliable Political Fact-Checking—Even with Reasoning and Web Search (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models (LLMs) have raised hopes for automated end-to-end fact-checking, but prior studies report mixed results. |
| Approach: | They evaluate 15 large language models on 6,000 claims fact-checked by PolitiFact . standard models perform poorly, reasoning offers minimal benefits, and web search provides only moderate gains . |
| Outcome: | The models predict claim veracity and a curated RAG system improved macro F1 by 233% on average across model variants. |
From RAG to Agentic RAG for Faithful Islamic Question Answering (2026.findings-acl)
Copied to clipboard
Gagan Bhatia, Hamdy Mubarak, Mustafa Jarrar, George Mikros, Fadi Zaraket, Mahmoud Alhirthani, Mutaz al-Khatib, Logan Cochrane, Kareem Mohamed Darwish, Rashid Yahiaoui, Firoj Alam
| Challenge: | Large Language Models (LLMs) are increasingly used for Islamic question answering, where ungrounded responses may carry serious religious consequences. |
| Approach: | They propose a bilingual, bilingual, Arabic/English benchmark with atomic single-gold answers that measures hallucination and abstention. |
| Outcome: | The proposed model improves accuracy and robustness even with a small model. |
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
FaStFact: Faster, Stronger Long-Form Factuality Evaluations in LLMs (2025.findings-emnlp)
Copied to clipboard
Yingjia Wan, Haochen Tan, Xiao Zhu, Xinyu Zhou, Zhiwei Li, Qingsong Lv, Changxuan Sun, Jiaqi Zeng, Yi Xu, Jianqiao Lu, Yinhong Liu, Zhijiang Guo
| Challenge: | Prior evaluation pipelines fail to evaluate factuality of long-form LLMs due to inefficiency and costly human assessment. |
| Approach: | They propose a fast and strong evaluation pipeline that can evaluate factuality of long-form LLMs . they propose 'faStFact' to reduce cost of web searching and inference calling . |
| Outcome: | The proposed evaluation pipeline achieves highest alignment with human evaluation and efficiency among existing baselines. |
FACT-AUDIT: An Adaptive Multi-Agent Framework for Dynamic Fact-Checking Evaluation of Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Existing fact-checking evaluation methods rely on static datasets and classification metrics, which fail to evaluate justification production and uncover the nuanced limitations of LLMs. |
| Approach: | They propose a framework that adaptively and dynamically assesses LLMs’ fact-checking capabilities by incorporating justification production alongside verdict prediction. |
| Outcome: | Experiments show that the framework differentiates among state-of-the-art LLMs, providing valuable insights into model strengths and limitations in model-centric fact-checking analysis. |
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for fact-checking text generated by large language models are expensive and time-consuming. |
| Approach: | They propose a plug-and-play framework that harnesses large language models for efficient fact-checking in a few-shot manner. |
| Outcome: | The proposed framework is compared with state-of-the-art models and shows that it can be used to speed up fact-checking in a few-shot manner. |