Papers by Ziwen Li

4 papers
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for mathematical reasoning are becoming less effective due to performance saturation.
Approach: They propose to use a mathematical reasoning benchmark with Olympiad difficulty to evaluate top-tier LLMs.
Outcome: The proposed benchmarks are cross-validated by experts to meet IMO difficulty standards and entirely original problems to prevent performance leakages from data memorization.
GraPPI: A Retrieve-Divide-Solve GraphRAG Framework for Large-scale Protein-protein Interaction Exploration (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models and Retrieval-Augmented Generation frameworks have accelerated drug discovery, but integrating models into workflows remains challenging.
Approach: They propose a large-scale knowledge graph-based retrieve-divide-solve agent pipeline RAG framework to support large-level PPI signaling pathway exploration.
Outcome: The proposed framework is based on large-scale knowledge graphs and can be used to analyze protein-protein interactions.
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models (2024.acl-demos)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved performance across tasks and domains . instruction tuning is a crucial technique to enhance the capabilities of LLMs - but there is no standard open-source instruction processing framework available for the community .
Approach: They propose an open-source instruction tuning framework for Large Language Models that modularizes instruction generation, selection, prompting and their combination and interaction.
Outcome: The proposed framework is open-source and available on Github.
RAGPPI: Retrieval-Augmented Generation Benchmark for Protein–Protein Interactions in Drug Discovery (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models and Retrieval-Augmented Generation (RAG) frameworks have supported Target ID, but no benchmark exists for identifying biological impacts of PPIs.
Approach: They propose to build a factual question-answer benchmark of 4,420 question-announced pairs that focus on the potential biological impacts of PPIs.
Outcome: The proposed benchmark is based on 4,420 question-answer pairs with expert-driven data annotation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations