Papers by Kevin Lu

7 papers
ELISA-EDL: A Cross-lingual Entity Extraction, Linking and Localization System (N18-5)

Copied to clipboard

Challenge: ELISA-EDL is a cross-lingual entity extraction, linking and localization system for Wikipedia languages.
Approach: They propose a cross-lingual entity extraction, linking and localization system for English speakers . it extracts entities from unstructured text in any of 282 Wikipedia languages and links them to English knowledge bases .
Outcome: The proposed system extracts entity mentions from Wikipedia and links them to English knowledge bases and visualizes locations related to disaster topics on a world heatmap.
OmniCode: A Benchmark for Evaluating Software Development Agents (2026.findings-acl)

Copied to clipboard

Challenge: popular coding benchmarks focus on narrowly scoped tasks such as competition programming and patch generation.
Approach: They propose a software engineering benchmark that aims to provide a broader set of tasks beyond code or patch generation.
Outcome: The proposed framework performs well on bug fixing for Python, test generation, code review fixing, and style fixing with popular agent frameworks such as SWE-Agent.
Character-Based Models for Adversarial Phone Extraction: Preventing Human Sex Trafficking (D19-55)

Copied to clipboard

Challenge: Illicit activity on the Web often obscures information between client and seller, such as the seller’s phone number.
Approach: They propose to use a dataset to model adversarial noise in a text extraction system and propose a visual character language model to interpret unseen unicode characters.
Outcome: The proposed model improves number recognition by 89% over a CRF with a CNN and shows that unicode characters can be translated to unicoding.
PINEAPPLE: Personifying INanimate Entities by Acquiring Parallel Personification Data for Learning Enhanced Generation (2022.coling-1)

Copied to clipboard

Challenge: Personifications are figures of speech that endow inanimate entities with properties and actions typically seen as requiring animacy.
Approach: They propose to use personification data to train a parallel corpus of personifications . they propose to combine personification-related literalizations with automatic ones .
Outcome: The proposed personification system can generate diverse and creative personifications . it can generate personification-related qualities such as interestingness and animacy .
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts (2025.emnlp-main)

Copied to clipboard

Challenge: Current large language models struggle to answer questions that span tens of thousands of tokens.
Approach: They evaluate 1–4 hop QA over 64k–128k-token excerpts from 83 novels . they find consistent accuracy drops with increased hops and context length .
Outcome: The novelhopqa benchmark evaluates 1–4 hop QA over 64k–128k-token excerpts from 83 public-domain novels.
Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies have evaluated and shown limitations in specific capabilities such as visual understanding, but a systematic evaluation of VLMs’ fundamental WM abilities remains absent.
Approach: They propose a framework that assesses perception and prediction to provide an atomic evaluation of VLMs as WMs.
Outcome: The proposed framework assesses perception and prediction abilities on 15 latest VLMs and compares them to human-level models.
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning (2025.naacl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) are primarily trained on high-resource natural languages, limiting their effectiveness in low-resourced settings and in tasks requiring deep logical reasoning.
Approach: They propose to use a dataset of logical propositions from Lean into a custom logical language to evaluate LLMs' logical reasoning and generalization capabilities in a controlled environment.
Outcome: The proposed model improves accuracy and accuracy beyond 20,000 training samples.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations