Challenge: a new dataset evaluates the credibility of statements made by Israeli public figures and politicians . a dataset of 1021 statements is used to assess the credibility and accuracy of statements .
Approach: They propose a dataset to evaluate the credibility of statements by Israeli politicians . they use annotated statements manually annotating them for their credibility status .
Outcome: The proposed model outperforms models based on statement and context, and achieves a 48.3 F1 score.

Similar Papers

A Fine-grained Sentiment Dataset for Norwegian (2020.lrec-1)

Copied to clipboard

Challenge: Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines.
Approach: They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement .
Outcome: The proposed dataset is the first of its kind for Norwegian and is available online.
EU DisinfoTest: a Benchmark for Evaluating Language Models’ Ability to Detect Disinformation Narratives (2024.findings-emnlp)

Copied to clipboard

Challenge: Disinformation narratives can be deceptive and disinformative, designed to sow division, distrust, and fear.
Approach: They propose to evaluate the efficacy of Language Models in identifying disinformation narratives using a Human-in-the-Loop methodology.
Outcome: The EU DisinfoTest evaluates language models on their ability to perform zero-shot classification of disinformation narratives versus credible narratives.
X-Fact: A New Benchmark Dataset for Multilingual Fact Checking (2021.acl-short)

Copied to clipboard

Challenge: Several fact-checking initiatives, such as PolitiFact, expend manual labor to investigate and determine the truthfulness of viral statements.
Approach: They propose a multilingual dataset for factual verification of naturally existing claims . they use a benchmark to evaluate the multilingual models .
Outcome: The proposed model achieves an F-score of around 40%, suggesting it is a challenging benchmark for multilingual fact-checking models.
Verifying Annotation Agreement without Multiple Experts: A Case Study with Gujarati SNACS (2023.findings-acl)

Copied to clipboard

Challenge: a small fraction of the about 7,000 languages of the world have datasets or linguistic tools . linguistic datasets are a foundation of NLP research, but they are not always reliable . authors propose weak verifiers to help estimate dataset quality .
Approach: They propose four weak verifiers to help estimate dataset quality . they propose to use Gujarati as a low-resource language to test for dataset quality.
Outcome: The proposed methods concur with a double-annotation study in Gujarati.
Adversarial NLI: A New Benchmark for Natural Language Understanding (2020.acl-main)

Copied to clipboard

Challenge: a new large-scale NLI benchmark dataset is presented to test models on a variety of popular NLIs.
Approach: They propose a large-scale NLI benchmark dataset that is iteratively compared with a human-and-model-in-the-loop procedure.
Outcome: The proposed method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate.
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs).
Approach: They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs.
Outcome: The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints.
HebID: Detecting Social Identities in Hebrew-language Political Text (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing NLP datasets focus on coarse-grained identity categories . existing datasets are mostly English-centric and focus on fine-grain categories based on cultural contexts.
Approach: They introduce the first multilabel Hebrew corpus for social identity detection . they use Hebrew-tuned encoders alongside 2B-9B-parameter decoders .
Outcome: The proposed classifier is based on a national public survey and uses Hebrew-tuned encoders to analyze political discourse and political speeches.
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)

Copied to clipboard

Challenge: a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles .
Approach: They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection.
Outcome: The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection .
The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools (2022.lrec-1)

Copied to clipboard

Challenge: a growing number of studies address the spoken language understanding domain through a simple task like speech intent detection.
Approach: They focus on the french MEDIA SLU dataset, which is distributed since 2005 . they propose a recipe for its use, including data preparation, training and evaluation scripts .
Outcome: The MEDIA SLU dataset is used as a benchmark dataset for a large number of research projects.
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)

Copied to clipboard

Challenge: Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior .
Approach: They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context.
Outcome: The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations