The Truth, The Whole Truth, and Nothing but the Truth: A New Benchmark Dataset for Hebrew Text Credibility Assessment (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a new dataset evaluates the credibility of statements made by Israeli public figures and politicians . a dataset of 1021 statements is used to assess the credibility and accuracy of statements . |
| Approach: | They propose a dataset to evaluate the credibility of statements by Israeli politicians . they use annotated statements manually annotating them for their credibility status . |
| Outcome: | The proposed model outperforms models based on statement and context, and achieves a 48.3 F1 score. |
Similar Papers
A Fine-grained Sentiment Dataset for Norwegian (2020.lrec-1)
Copied to clipboard
| Challenge: | Using a dataset for fine-grained sentiment analysis in Norwegian, we examine the annotation effort and provide an overview of the developed annotation guidelines. |
| Approach: | They propose a dataset for fine-grained sentiment analysis in Norwegian . they provide an overview of the developed annotation guidelines and analyze inter-annotator agreement . |
| Outcome: | The proposed dataset is the first of its kind for Norwegian and is available online. |
EU DisinfoTest: a Benchmark for Evaluating Language Models’ Ability to Detect Disinformation Narratives (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Disinformation narratives can be deceptive and disinformative, designed to sow division, distrust, and fear. |
| Approach: | They propose to evaluate the efficacy of Language Models in identifying disinformation narratives using a Human-in-the-Loop methodology. |
| Outcome: | The EU DisinfoTest evaluates language models on their ability to perform zero-shot classification of disinformation narratives versus credible narratives. |
X-Fact: A New Benchmark Dataset for Multilingual Fact Checking (2021.acl-short)
Copied to clipboard
| Challenge: | Several fact-checking initiatives, such as PolitiFact, expend manual labor to investigate and determine the truthfulness of viral statements. |
| Approach: | They propose a multilingual dataset for factual verification of naturally existing claims . they use a benchmark to evaluate the multilingual models . |
| Outcome: | The proposed model achieves an F-score of around 40%, suggesting it is a challenging benchmark for multilingual fact-checking models. |
Verifying Annotation Agreement without Multiple Experts: A Case Study with Gujarati SNACS (2023.findings-acl)
Copied to clipboard
| Challenge: | a small fraction of the about 7,000 languages of the world have datasets or linguistic tools . linguistic datasets are a foundation of NLP research, but they are not always reliable . authors propose weak verifiers to help estimate dataset quality . |
| Approach: | They propose four weak verifiers to help estimate dataset quality . they propose to use Gujarati as a low-resource language to test for dataset quality. |
| Outcome: | The proposed methods concur with a double-annotation study in Gujarati. |
Adversarial NLI: A New Benchmark for Natural Language Understanding (2020.acl-main)
Copied to clipboard
| Challenge: | a new large-scale NLI benchmark dataset is presented to test models on a variety of popular NLIs. |
| Approach: | They propose a large-scale NLI benchmark dataset that is iteratively compared with a human-and-model-in-the-loop procedure. |
| Outcome: | The proposed method can be applied in a never-ending learning scenario, becoming a moving target for NLU, rather than a static benchmark that will quickly saturate. |
TruthTrap: A Bilingual Benchmark for Evaluating Factually Correct Yet Misleading Information in Question Answering (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to answer factual, information-seeking questions (ISQs). |
| Approach: | They propose to use a dataset to evaluate large language models to generate human-like text on ISQs in two languages, English and Farsi, and then use it to evaluate nine LLMs. |
| Outcome: | The proposed dataset shows that accuracy drops by 25% when models encounter misleading yet factual hints. |
HebID: Detecting Social Identities in Hebrew-language Political Text (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing NLP datasets focus on coarse-grained identity categories . existing datasets are mostly English-centric and focus on fine-grain categories based on cultural contexts. |
| Approach: | They introduce the first multilabel Hebrew corpus for social identity detection . they use Hebrew-tuned encoders alongside 2B-9B-parameter decoders . |
| Outcome: | The proposed classifier is based on a national public survey and uses Hebrew-tuned encoders to analyze political discourse and political speeches. |
BREAKING! Presenting Fake News Corpus for Automated Fact Checking (P19-2)
Copied to clipboard
| Challenge: | a new study shows that fake news spreads faster than mainstream articles on the same topic . however, there is no dataset containing compelling fake and questionable news articles . |
| Approach: | They introduce manually verified corpus of compelling fake and questionable news articles on the USA politics . they plan to extend the corpus in the future and use it for automated fake news detection. |
| Outcome: | The proposed model is based on linguistic features and will be extended in the future . it will be used to improve the existing model and improve the tools in the field of fake news detection . |
The Spoken Language Understanding MEDIA Benchmark Dataset in the Era of Deep Learning: data updates, training and evaluation tools (2022.lrec-1)
Copied to clipboard
Gaëlle Laperrière, Valentin Pelloin, Antoine Caubrière, Salima Mdhaffar, Nathalie Camelin, Sahar Ghannay, Bassam Jabaian, Yannick Estève
| Challenge: | a growing number of studies address the spoken language understanding domain through a simple task like speech intent detection. |
| Approach: | They focus on the french MEDIA SLU dataset, which is distributed since 2005 . they propose a recipe for its use, including data preparation, training and evaluation scripts . |
| Outcome: | The MEDIA SLU dataset is used as a benchmark dataset for a large number of research projects. |
Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency (2025.emnlp-main)
Copied to clipboard
Svetlana Maslenkova, Clement Christophe, Marco AF Pimentel, Tathagata Raha, Muhammad Umar Salman, Ahmed Al Mahrooqi, Avani Gupta, Shadab Khan, Ronnie Rajan, Praveenkumar Kanithi
| Challenge: | Current dataset curation and bias assessment practices lack transparency . current approaches lack a thorough understanding of how data characteristics influence model behavior . |
| Approach: | They propose a comprehensive bias evaluation framework that integrates general benchmarks with a healthcare-specific methodology to probe for biases in a sensitive healthcare context. |
| Outcome: | The proposed approach to bias evaluation leverages established benchmarks and a healthcare-specific methodology. |