Papers by Chengkai Li

14 papers
DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for question answering (QA) are lacking in a high-stakes environment.
Approach: They propose a rigorously verified benchmark of 3,000 expert-annotated questions . they propose 'keypoint-based evaluation protocol' emphasizing factual completeness over verbosity .
Outcome: Experiments with 20 models reveal substantial divergences from general-purpose models such as MMLU-Pro.
A Dashboard for Mitigating the COVID-19 Misinfodemic (2021.eacl-demos)

Copied to clipboard

Challenge: a new public dashboard aims to understand the impact of the COVID-19 misinfodemic on Twitter . the dashboard uses a curated catalog of COVId-19 related facts and debunks of misinformation .
Approach: They propose a public dashboard that matches tweets with COVID-19 misinformation . they also propose experiments to analyze the spread of misinformation on twitter .
Outcome: The proposed dashboard uses a curated catalog of COVID-19 related facts and debunks misinformation . it shows the most prevalent information from the catalog among Twitter users in user-selected geographic regions .
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media (2025.findings-acl)

Copied to clipboard

Challenge: Social media's global reach and ease of use have transformed how millions of users exchange opinions, news, and factual claims in real-time, making it fertile ground for misinformation.
Approach: They propose a framework that leverages large language models to construct taxonomies of factual claims from social media by generating topics at multiple levels of granularity.
Outcome: The proposed framework produces clear, coherent, and comprehensive taxonomies on three diverse datasets and outperforms other frameworks in most metrics.
MemWeaver: Weaving Hybrid Memories for Traceable Long-Horizon Agentic Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods rely on unstructured retrieval or coarse abstractions, which lead to temporal conflicts, brittle reasoning, and limited traceability.
Approach: They propose a unified memory framework that consolidates long-term agent experiences into three interconnected components that combine structured knowledge and evidence to construct compact yet information-dense contexts for reasoning.
Outcome: The proposed framework significantly improves multi-hop and temporal reasoning accuracy while reducing input context length by over 95% compared to long-context baselines.
CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Automated Fact-Checking has largely focused on verifying general knowledge against static corpora.
Approach: They propose a benchmark to verify colloquial legal claims against Supreme Court precedents . the benchmark leverages large language models to synthesize claims from expert case summaries . they say the benchmark is a step forward in the field of legal fact verification .
Outcome: The proposed benchmark bridges the gap between layperson assertions and technical jurisprudence while accounting for temporal validity.
RATSD: Retrieval Augmented Truthfulness Stance Detection from Social Media Posts Toward Factual Claims (2025.findings-naacl)

Copied to clipboard

Challenge: Social media provides a valuable lens for assessing public perceptions and opinions.
Approach: They propose a method that leverages large language models with retrieval-augmented generation to analyze tweets in relation to claims.
Outcome: The proposed method outperforms state-of-the-art methods on a new dataset . it shows that it outperformed existing methods and achieves a significant increase in Macro-F1 score on TSD-CT.
Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Graph-to-text models trained on small-scale datasets or datasets with limited variety of graph shapes are not adequate for more realistic large-scale, open-domain settings.
Approach: They propose a novel approach that, given a graph-sentence pair in GraphNarrative, trims the sentence to eliminate portions that are not present in the corresponding graph.
Outcome: The proposed model can be trained on existing datasets and is available on github.
Modeling Factual Claims with Semantic Frames (2020.lrec-1)

Copied to clipboard

Challenge: In recent years, the proliferation of misinformation has reached a staggering pace eroding people's confidence in politics and even affected democracies.
Approach: They propose an extension of the Berkeley FrameNet for the structured and semantic modeling of factual claims.
Outcome: The proposed extension provides 2,540 fully annotated sentences and can be used to understand how these frames are intended to work and to train machine learning models.
ClaimPortal: Integrated Monitoring, Searching, Checking, and Analytics of Factual Claims on Twitter (P19-3)

Copied to clipboard

Challenge: ClaimPortal is a web-based platform for monitoring, searching, checking and analyzing factual claims on Twitter from the American political domain.
Approach: They present a web-based platform for monitoring, searching, checking and analyzing English factual claims on Twitter from the American political domain.
Outcome: The proposed platform can monitor, search, check, and analyze English factual claims on Twitter from the political domain.
Robust Frame-Semantic Models with Lexical Unit Trees and Negative Samples (2024.acl-long)

Copied to clipboard

Challenge: Using a RoBERTa-based filter, we achieve an F1 score of 0.775, surpassing the previous state-of-the-art solution by +0.012.
Approach: They propose a new prefix tree modification to enable robust support for multi-word lexical units and a RoBERTa-based filter to achieve an F1 score of 0.775.
Outcome: The proposed model achieves an F1 score of 0.775, surpassing the state-of-the-art model by +0.012.
Task-Oriented Automatic Fact-Checking with Frame-Semantics (2025.findings-acl)

Copied to clipboard

Challenge: Existing work on automatic fact-checking relies on unstructured data and large language models to produce fact- check verdicts and explanations.
Approach: They propose a new paradigm for automatic fact-checking that leverages frame semantics to enhance the structured understanding of claims and guide the process of fact- checking them.
Outcome: The proposed paradigm improves evidence retrieval and explainability for fact-checking by leveraging frame semantics.
Can LLMs Extract Frame-Semantic Arguments? (2025.emnlp-main)

Copied to clipboard

Challenge: Frame-semantic parsing is a critical task in natural language understanding . however, the ability of large language models to extract frame-sensical arguments remains unexplored .
Approach: They propose a framework to extract frame-semantic arguments from large language models . they use JSON representations to enhance performance, but smaller models can achieve competitive results .
Outcome: The proposed model achieves state-of-the-art on ambiguous targets while limiting generalization to out-of domain data.
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes (2026.acl-long)

Copied to clipboard

Challenge: Existing preference-based approaches fail to address this challenge by exploiting language priors to bypass visual grounding.
Approach: They propose a framework that leverages scene graphs as structured visual information to perform controllable structural interventions.
Outcome: The proposed framework improves answer accuracy and reasoning faithfulness across seven visual reasoning benchmarks.
ClaimLens: Automated, Explainable Fact-Checking on Voting Claims Using Frame-Semantics (2024.emnlp-demo)

Copied to clipboard

Challenge: Existing fact-checking solutions lack transparency and explainability . a lack of transparency can make it difficult for users to trust and understand the reasoning behind the outcomes.
Approach: They propose an automated fact-checking system focused on voting-related factual claims that leverages frame-semantic parsing to provide structured and interpretable fact verification.
Outcome: The proposed system can extract relevant information from voting-related factual claims using public records and Vote semantic frame.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations