Papers by Zaiqiao Meng

24 papers
Mixture-of-Partitions: Infusing Large Biomedical Knowledge Graphs into BERT (2021.emnlp-main)

Copied to clipboard

Challenge: Infusing factual knowledge into pre-trained models is fundamental for many knowledge-intensive tasks.
Approach: They propose an infusion approach that partitions a large knowledge graph into smaller sub-graphs and infuses their specific knowledge into various BERT models using lightweight adapters.
Outcome: The proposed approach improves the underlying BERTs and achieves new SOTA performance on six downstream tasks.
Time to Revisit Exact Match (2025.findings-emnlp)

Copied to clipboard

Challenge: Temporal question answering is an established method for evaluating temporal reasoning in large language models.
Approach: They propose a numerical estimation task where all questions require a numeric, temporal answer, allowing us to evaluate models beyond EM.
Outcome: The proposed model responses are based on a numerical estimation task and are distilled from Test of Time and TempTabQA.
EvoAgentX: An Automated Framework for Evolving Agentic Workflows (2025.emnlp-demos)

Copied to clipboard

Challenge: Existing MAS frameworks often require manual workflow configuration and lack native support for dynamic evolution and performance optimization.
Approach: They propose an open-source platform that automates generation, execution, and evolutionary optimization of multi-agent workflows.
Outcome: The proposed platform automates generation, execution, and evolutionary optimization of multi-agent workflows.
RadEval: A framework for radiology text evaluation (2025.emnlp-demos)

Copied to clipboard

Challenge: Evaluating automated radiology report generation systems remains a fundamental challenge in the development of safe, accurate, and clinically useful medical AI.
Approach: They propose a unified, open-source framework for evaluating radiology texts that consolidates a diverse range of metrics from classic ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (GREEN).
Outcome: The framework consolidates a diverse range of metrics from ngram overlap (BLEU) and contextual measures (BERTScore) to clinical concept-based scores (F1CheXbert, F1RadGraph, RaTEScore, SRR-BERT, TemporalEntityF1) and advanced LLMbased evaluators (GREEN).
CCD: Mitigating Hallucinations in Radiology MLLMs via Clinical Contrastive Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal large language models generate medical hallucinations due to over-sensitivity to clinical sections.
Approach: They propose a framework that integrates structured clinical signals from task-specific radiology expert models.
Outcome: The proposed framework improves overall performance on radiology report generation (RRG) on the MIMIC-CXR dataset, it yields up to 17% improvement in RadGraph-F1.
Biomedical Named Entity Recognition via Dictionary-based Synonym Generalization (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for biomedical named entity recognition require laborious human effort.
Approach: They propose a Synonym Generalization framework that recognizes biomedical concepts using span-based predictions.
Outcome: The proposed framework outperforms dictionary-based approaches on a wide range of benchmarks.
TRACE the Evidence: Constructing Knowledge-Grounded Reasoning Chains for Retrieval-Augmented Generation (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing retrievers are not perfect and often include irrelevant documents in the retrieved set.
Approach: They propose to construct knowledge-grounded reasoning chains from retrieved documents to integrate supporting evidence into RAG models.
Outcome: The proposed model achieves an average performance improvement of 14.03% on three multi-hop QA datasets.
Few-Shot Table-to-Text Generation with Prototype Memory (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural table-to-text generation models are data-hungry and require large amounts of training data to learn the mapping between tables and texts.
Approach: They propose a framework for table-to-text generation under the few-shot scenario that uses retrieved prototypes and a prototype selector to bridge the structural gap between tables and texts.
Outcome: The proposed framework significantly improves the model performance on three benchmark datasets with state-of-the-art models.
FusionDTI: Fine-grained Binding Discovery with Token-level Fusion for Drug-Target Interaction (2025.findings-emnlp)

Copied to clipboard

Challenge: despite advances in DTI models, models often struggle to capture fine-grained interactions between drugs and proteins.
Approach: They propose a novel drug-target interaction model that uses a token-level module to learn fine-grained information for drug-target interactions.
Outcome: The proposed model learns fine-grained information for drug-target interaction . it mitigates sequence fragment invalidation and incorporates the structure-aware vocabulary of target proteins .
Can We Edit LLMs for Long-Tail Biomedical Knowledge? (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing knowledge editing methods can enhance LLMs' performance on long-tail biomedical knowledge, but their performance on high-frequency popular knowledge remains inferior to that on high frequency popular knowledge.
Approach: They conduct the first comprehensive study to investigate the effectiveness of knowledge editing methods for editing long-tail biomedical knowledge.
Outcome: The proposed methods improve LLMs' performance on long-tail biomedical knowledge, but their performance on high-frequency popular knowledge remains inferior even after editing.
Rewire-then-Probe: A Contrastive Recipe for Probing Biomedical Knowledge of Pre-trained Language Models (2022.acl-long)

Copied to clipboard

Challenge: Despite the growing progress of probing knowledge for pre-trained language models, specialised areas such as the biomedical domain are vastly under-explored.
Approach: They propose a biomedical knowledge probing benchmark, MedLAMA, constructed based on the Unified Medical Language System (UMLS) Metathesaurus.
Outcome: The proposed approach pushes the acc@10 to 28%, but the performance gap remains notable.
Libra: Leveraging Temporal Images for Biomedical Radiology Analysis (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for radiology report generation rely on single-image analysis or rule-based heuristics to process multiple images.
Approach: They propose a temporal-aware MLLM tailored for chest X-ray report generation that combines a radiology-specific image encoder with a novel Temporal Alignment Connector.
Outcome: The proposed model sets new standards in clinical relevance and lexical accuracy on the MIMIC-CXR dataset.
REANO: Optimising Retrieval-Augmented Reader Models through Knowledge Graph Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing knowledge graphs suffer from incompleteness and lack information critical for answering given questions.
Approach: They propose to enhance the open domain question answering model with a knowledge graph generation module that generates KGs from the passages and an answer predictor.
Outcome: The proposed model improves the exact match score by 2.7% on the EntityQuestion dataset, with an average improvement of 1.8% across all the datasets.
To Generate or to Retrieve? On the Effectiveness of Artificial Contexts for Medical Open-Domain Question Answering (2024.acl-long)

Copied to clipboard

Challenge: Medical open-domain question answering requires substantial access to specialized knowledge.
Approach: They propose a framework that generates multiple-choice questions from a set of open-book parameters and a small-scale reader that can outcompete closed-book questions by 706x using fewer parameters.
Outcome: The proposed framework outperforms closed-book models on MedQA-USMLE, MedMCQA, and MMLU while using up to 706x fewer parameters.
SPARKLE: A Structured and Plug-and-play Agentic Retrieval Policy for Adaptive RAG Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for integrating external knowledge rely on frozen large language models without explicit supervision or require costly LLM finetuning.
Approach: They propose a structured and plug-and-play agentic retrieval policy with an additional proxy model to control the retrieval process.
Outcome: Experiments on three in-domain and four out-of-domain QA benchmarks show that SPARKLE outperforms state-of the-art adaptive RAG models, achieving average improvements of 9.17% and 2.85%, respectively.
KiRAG: Knowledge-Driven Iterative Retriever for Enhancing Retrieval-Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: Iterative retrieval-augmented generation models are difficult to use for multihop question answering (QA) . their retrieval processes can be disrupted by irrelevant documents or factually inaccurate chain-of-thoughts .
Approach: They propose a knowledge-driven iterative retriever model that decomposes documents into knowledge triples and performs iterativ retrieval with these triples to enable a factually reliable retrieval process.
Outcome: The proposed model outperforms existing iRAG models with an average improvement of 9.40% in R@3 and 5.14% in F1 on multi-hop QA datasets.
Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck (2026.acl-long)

Copied to clipboard

Challenge: Existing studies have identified a position bias in Large Language Models that causes them to overlook information at certain positions.
Approach: They propose a semantic probe to disentangle position bias in Large Language Models . they propose MFAI to steer attention towards selected positions .
Outcome: The proposed model can locate and integrate information at certain positions even in noisy, long-context settings.
TaCL: Improving BERT Pre-training with Token-aware Contrastive Learning (2022.findings-naacl)

Copied to clipboard

Challenge: Existing pre-trained MLMs produce an anisotropic distribution of token representations . this is not ideal for tasks that require discriminative semantic meanings of distinct tokens - a problem that exists in pre-training models .
Approach: They propose a continual pre-training approach that encourages BERT to learn an isotropic distribution of token representations.
Outcome: The proposed approach improves on a wide range of English and Chinese benchmarks.
Revisiting Parameter-Efficient Tuning: Are We Really There Yet? (2022.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models (PLMs) are used as backbones to be combined with additional parameters and finetuned on downstream tasks in an end-to-end manner.
Approach: They propose to use a fraction of parameters to tune pretrained language models (PLMs) this is the first comprehensive investigation into the training and evaluation of PETuning methods.
Outcome: The proposed methods have been validated and tested with a rigorous evaluation protocol and have shown that they are unstable and inconsistent.
GenKIE: Robust Generative Multimodal Document Key Information Extraction (2023.findings-emnlp)

Copied to clipboard

Challenge: Key information extraction (KIE) is a key application for information retrieval and text mining.
Approach: They propose a novel generative end-to-end model, named GenKIE, to address the KIE task.
Outcome: The proposed model generalizes over different types of documents and achieves state-of-the-art results.
Self-Alignment Pretraining for Biomedical Entity Representations (2021.naacl-main)

Copied to clipboard

Challenge: Existing approaches to self-supervised learning of biomedical entities are limited in the biomedic domain.
Approach: They propose a pretraining scheme that self-aligns the representation space of biomedical entities.
Outcome: The proposed framework achieves state-of-the-art on six MEL benchmarking datasets.
MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity Recognition (2023.acl-long)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a fundamental NLP task that aims at classifying mention spans into entity types.
Approach: They propose a variational memory-augmented few-shot named entity recognition model that uses a memory module to store information from source domain and retrieve relevant information from the memory to augment few-shot task in target domain.
Outcome: The proposed model can adapt the learned knowledge from source domain to target domain and achieve superior performance on English and Chinese cross domain few-shot NER datasets.
Can Pretrained Language Models (Yet) Reason Deductively? (2023.eacl-main)

Copied to clipboard

Challenge: Acquiring factual knowledge with Pretrained Language Models (PLMs) has attracted increasing attention, showing promising performance in many knowledge-intensive tasks.
Approach: They conduct a comprehensive evaluation of the learnable deductive reasoning capability of pretrained language models and compare their performance against simple adversarial surface form edits.
Outcome: The models are able to generalise learned logic rules and perform inconsistently against simple adversarial surface form edits, but catastrophically forget the previously learnt knowledge.
Can We Instruct LLMs to Compensate for Position Bias? (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies reveal that position bias in large language models (LLMs) leads to difficulty in accessing information retrieved from the retriever.
Approach: They propose to direct LLMs to allocate more attention towards a selected segment of the context through prompting.
Outcome: The proposed approach improves the performance of large language models by promoting instruction with an exact document index.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations