Papers by Sen Wu

9 papers
PLAWBENCH: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning.
Approach: They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios.
Outcome: The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics.
Template-Based Named Entity Recognition Using BART (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for fewshot NER do not make full use of knowledge transfer in NER model parameters.
Approach: They propose a template-based method for NER that treats NER as a language model ranking problem in a sequence-to-sequence framework.
Outcome: The proposed method achieves 92.55% F1 score on the CoNLL03 task and significantly better than fine-tuning BERT 10.88%, 15.34%, and 11.73% F1 scores on the MIT Movie, the ATIS, and the MATLAB task.
StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation (2026.acl-long)

Copied to clipboard

Challenge: Domain-specific datasets of harmful prompts are scarce and often rely on manual construction. Existing efforts to improve domain knowledge and reduce harmful prompt generation are lacking.
Approach: They propose a framework that transforms domain knowledge into actionable constraints and increases the implicitness of generated harmful prompts.
Outcome: The proposed framework yields high-quality datasets combining strong domain relevance with implicitness, enabling more realistic red-teaming and advancing LLM safety research.
Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) often struggle with the issue of generating inaccurate or fabricated content even when they possess correct knowledge.
Approach: They propose a decoding method that mitigates hallucinations without extra training . they propose entropy eNhanced decoding that leverages inner probability changes .
Outcome: The proposed method improves the truthfulness and informativeness of generation while maintaining robust QA accuracy.
Challenges to Open-Domain Constituency Parsing (2022.findings-acl)

Copied to clipboard

Challenge: Existing findings on cross-domain constituency parsing are only made on a limited number of domains.
Approach: They manually annotate a high-quality constituency treebank containing five domains and analyze challenges to open-domain constituency parsing using a set of linguistic features.
Outcome: The proposed model significantly improves the performance of the proposed model on the domain-variant features.
Cross-Domain Data Integration for Named Entity Disambiguation in Biomedical Text (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for named entity disambiguation are limited by coarse-grained structural resources in biomedical knowledge bases and training datasets that provide low coverage over uncommon resources.
Approach: They propose a method that integrates structural knowledge from general text knowledge bases to the medical domain.
Outcome: The proposed method improves disambiguation accuracy on two benchmark medical NED datasets by up to 57 points.
Metadata Shaping: A Simple Approach for Knowledge-Enhanced Language Models (2022.findings-acl)

Copied to clipboard

Challenge: Existing methods to capture entity knowledge with factual knowledge are limited . despite its simplicity, metadata shaping is quite effective .
Approach: They propose a method which inserts substrings corresponding to readily available entity metadata into examples at train and inference time based on mutual information.
Outcome: The proposed method exceeds the baseline model by 4.3 F1 points and achieves state-of-the-art results.
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for tuning large language models from dense to MoE face significant data requirements and require large-scale post-training.
Approach: They propose an upcycling instruction tuning approach for tuning a dense pre-trained model into a MoE instruction model using genetic algorithm and parameter merging.
Outcome: The proposed approach improves the performance of large language models with a small amount of seed data and improves their scaling.
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions (2026.acl-long)

Copied to clipboard

Challenge: Existing long-horizon memory benchmarks use multi-turn dialogues or synthetic user histories . despite rapid progress on long-term memory evaluation, there are gaps in existing benchmarks .
Approach: They propose a long-form autobiographical narrative benchmark that reconstructs each narrative into a flashback-aware, time-anchored stream and evaluates models with evidence-linked questions.
Outcome: The proposed benchmarks build from long-form autobiographical narratives . they show that retrieval-augmented systems improve factual accuracy while errors persist on temporally grounded explanations and higher-level inferences.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations