Papers by Zhengping Liu

5 papers
Core: Robust Factual Precision with Informative Sub-Claim Identification (2025.findings-acl)

Copied to clipboard

Challenge: Using the Decompose-Then-Verify framework, such as FActScore, can be manipulated by adding obvious or repetitive subclaims to artificially inflate scores.
Approach: They propose a decomposition-based tool called Core to filter subclaims based on their uniqueness and informativeness.
Outcome: The proposed evaluation framework supports easy and modular use of Core and various decomposition strategies.
Calibrating Zero-shot Cross-lingual (Un-)structured Predictions (2022.emnlp-main)

Copied to clipboard

Challenge: Existing need for model calibration when natural language models are deployed in critical tasks.
Approach: They compare model calibration methods in a context of zero-shot cross-lingual transfer with pre-trained language models.
Outcome: The proposed method fails to calibrate more complex confidence estimations in structured predictions compared to expressive alternatives like Gaussian Process Calibration.
RORA: Robust Free-Text Rationale Evaluation (2024.acl-long)

Copied to clipboard

Challenge: Existing metrics rely on degree to which rationale supports a label, but they fail to evaluate rationales that inadvertently leak the label.
Approach: They propose a RObust free-text RAtionale evaluation against label leakage that quantifies the new information supplied by a rationale to justify the label.
Outcome: The proposed evaluation outperforms existing methods in evaluating human-written, synthetic, or model-generated rationales, particularly demonstrating robustness against label leakage.
DHP Benchmark: Are LLMs Good NLG Evaluators? (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly serving as evaluators in Natural Language Generation (NLG) tasks.
Approach: They propose a framework that measures the discernment of Large Language Models (LLMs) across diverse NLG tasks.
Outcome: The proposed framework provides quantitative discernment scores for LLMs across four NLG tasks.
PubMedQA: A Dataset for Biomedical Research Question Answering (D19-1)

Copied to clipboard

Challenge: PubMedQA is a biomedical question answering dataset based on PubMed abstracts . 68.1% accuracy is achieved, compared to single human performance of 78.0% .
Approach: They propose a biomedical question answering dataset from PubMed abstracts . the dataset is annotated by experts and has 1k instances of QA .
Outcome: The proposed model achieves 68.1% accuracy compared to human performance of 78.0% and majority-baseline of 55.2%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations