Papers by Huitong Pan

5 papers
SciER: An Entity and Relation Extraction Dataset for Datasets, Methods, and Tasks in Scientific Documents (2024.emnlp-main)

Copied to clipboard

Challenge: Scientific information extraction (SciIE) is critical for converting unstructured knowledge from scholarly articles into structured data.
Approach: They propose to use a scientific entity and relation extraction dataset to capture interactions between entities in full texts.
Outcome: The proposed dataset captures the intricate use and interactions among entities in full texts and provides an out-of-distribution test set to offer a more realistic evaluation.
DynClean: Training Dynamics-based Label Cleaning for Distantly-Supervised Named Entity Recognition (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to identify entities using distant annotations are expensive and time-consuming.
Approach: They propose a training dynamics-based label cleaning approach to characterize distant annotations and an automatic threshold estimation strategy to locate errors in distant labels.
Outcome: The proposed method outperforms several advanced DS-NER approaches across four datasets.
Taxonomy-Driven Knowledge Graph Construction for Domain-Specific Scientific Applications (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for constructing domain-specific knowledge graphs neglect curated taxonomies and LLMs fail to extract KGs in specialized domains.
Approach: They propose a taxonomy-driven framework for constructing domain-specific knowledge graphs . they use structured taxonomies, Large Language Models and Retrieval-Augmented Generation .
Outcome: The proposed framework can be adapted for other specialized domains.
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions (2024.lrec-main)

Copied to clipboard

Challenge: SciDMT is an enhanced and expanded corpus for scientific mention detection . existing corpora are limited by their small volume and entity linking capabilities .
Approach: They propose to enhance SciDMT, an annotated scientific corpus for scientific mention detection.
Outcome: The proposed corpus is the largest for scientific entity mention detection . it is based on deep learning architectures like SciBERT and GPT-3.5 .
DMDD: A Large-Scale Dataset for Dataset Mentions Detection (2023.tacl-1)

Copied to clipboard

Challenge: Existing corpora for dataset mention detection are limited in size and naming diversity.
Approach: They propose a dataset for dataset mention detection that is the largest publicly available corpus for this task.
Outcome: The proposed dataset is the largest publicly available corpus for dataset mention detection . it identifies open problems in dataset mention recognition and linking .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations