Challenge: Existing fact-checking resources cover COVID-19 related information in news, but there is no dataset providing fact- checked COVId-19 related tweets with detailed annotations for biomedical entities, relations and relevant evidence.
Approach: They propose a fact-checked corpus of tweets with annotations for biomedical entities, relations and relevant evidence for COVID-19 related tweets.
Outcome: The proposed dataset provides fact-checked COVID-19 related tweets with detailed annotations for biomedical entities, relations and relevant evidence.

Similar Papers

Extracting a Knowledge Base of COVID-19 Events from Social Media (2022.coling-1)

Copied to clipboard

Challenge: a flood of COVID-19 related information has appeared on social media since December 2019 . this includes reports on public figures who have tested positive/negative for the virus .
Approach: They construct a corpus of 10,000 tweets with annotated public reports of five COVID-19 events, using slot-filling questions to fill in slots.
Outcome: The proposed method can be quickly applied to develop knowledge bases for new domains in response to emerging crises, including natural disasters or future disease outbreaks.
Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society (2021.findings-emnlp)

Copied to clipboard

Challenge: a dataset of 16K manually annotated tweets is used to analyze disinformation . the democratic nature of social media has raised questions about the quality and the factuality of the information that is shared on these platforms.
Approach: They use a dataset of manually annotated tweets to analyze COVID-19 disinformation . they show that tweets contain fake cures, rumors, conspiracy theories and xenophobia .
Outcome: The proposed dataset shows that it is useful in monolingual vs. multilingual settings.
COVID-19 and Misinformation: A Large-Scale Lexical Analysis on Twitter (2021.acl-srw)

Copied to clipboard

Challenge: Social media is used by individuals and organisations as a platform to spread misinformation.
Approach: They compile a large corpus of tweets related to coronavirus and perform an analysis to discover patterns with respect to vocabulary usage.
Outcome: The proposed model based on lexical features is effective in identifying misinformation-related tweets with accuracy over 80%.
Human-in-the-loop Evaluation for Early Misinformation Detection: A Case Study of COVID-19 Treatments (2023.acl-long)

Copied to clipboard

Challenge: Existing evaluations of human-in-the-loop systems to combat misinformation are often set up automatically using datasets that were retrospectively constructed.
Approach: They propose a human-in-the-loop evaluation framework for fact-checking novel misinformation claims and identifying social media messages that support them.
Outcome: The proposed framework is based on modern NLP methods for human-in-the-loop fact-checking in the domain of COVID-19 treatments.
COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 Pandemic (2021.acl-long)

Copied to clipboard

Challenge: a new method for fact-checking is needed to detect disinformation on the web . a dataset COVID-Fact contains 4,086 claims concerning the COVId-19 pandemic .
Approach: They propose a FEVER-like dataset COVID-Fact of 4,086 claims concerning the COVId-19 pandemic . they automatically detect true claims and their source articles and generate counter-claims using automatic methods .
Outcome: The proposed method reduces the cost of building domain-specific datasets for detecting misinformation . the proposed dataset contains 4,086 claims concerning the COVID-19 pandemic .
Recovering Patient Journeys: A Corpus of Biomedical Entities and Relations on Twitter (BEAR) (2022.lrec-1)

Copied to clipboard

Challenge: Existing medical social media corpora focus on a small set of entities and relations . existing text mining and information extraction methods focus on scientific text generated by researchers but their access to individual patient experiences or patient-doctor interactions is limited.
Approach: The dataset consists of 2,100 medical tweets with approx. 6,000 entities and 2,200 relations.
Outcome: The proposed dataset consists of 2,100 tweets with approx. 6,000 entities and 2,200 relations.
Detecting Contradictory COVID-19 Drug Efficacy Claims from Biomedical Literature (2023.acl-short)

Copied to clipboard

Challenge: During times of pandemic, treatment options are limited, and developing new drug treatments is infeasible in the short-term.
Approach: They propose to use a natural language inference problem to automatically identify contradictory claims about COVID-19 drug efficacy.
Outcome: The proposed models help domain experts distill and assess evidence concerning remdisivir and hydroxychloroquine.
Check-COVID: Fact-Checking COVID-19 News Claims with Scientific Evidence (2023.findings-acl)

Copied to clipboard

Challenge: Existing fact-checking benchmarks require systems to verify claims from everyday text against evidence from scientific journal articles.
Approach: They propose a benchmark system that checks claims from news against scientific journal articles and veracity labels.
Outcome: The new benchmark achieves F1 scores of 76.99 and 69.90 on both a fact-checking specific system and GPT-3.5, respectively.
CMTA: COVID-19 Misinformation Multilingual Analysis on Twitter (2021.acl-srw)

Copied to clipboard

Challenge: myths, sensationalism, rumours and misinformation, generated intentionally or unintentionally, spread rapidly through social networks during the COVID-19 pandemic . evaluation of tweets for recognizing misinformation can create beneficial understanding to review the top quality and also the readability of online information concerning the COV-19.
Approach: They propose a multilingual COVID-19 related tweet analysis method that uses a deep learning model for multilingual tweet misinformation detection and classification.
Outcome: The proposed method outperforms monolingual models in the misinformation detection task and shows that it can be used to improve the quality and readability of online information.
Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter (2022.emnlp-main)

Copied to clipboard

Challenge: Current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of misinformation and fake news.
Approach: They propose a large-scale Twitter corpus with token-level claim spans on more than 7.5k tweets and a model that automatically detects and extracts the snippets of misinformation.
Outcome: The proposed model outperforms baseline systems on several evaluation metrics, improving by 1.5 points.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations