Papers by Yefan Tao

2 papers
Textual Dataset Distillation via Language Model Embedding (2024.findings-emnlp)

Copied to clipboard

Challenge: prevailing methods for dataset distillation generate distilled data as embedding vectors, which are not human-readable.
Approach: They propose a model-agnostic, data-efficient method that leverages Language Model embeddings . their method offers enhanced flexibility and improved transferability .
Outcome: The proposed method achieves comparable performance with faster processing times compared to other methods . it offers enhanced flexibility and improved transferability, expanding the range of potential applications .
BPID: A Benchmark for Personal Identity Deduplication (2024.emnlp-industry)

Copied to clipboard

Challenge: Data deduplication is a critical task in data management and mining, focused on consolidating duplicate records that refer to the same entity.
Approach: They propose to use a dataset with 1,000,000 unlabeled synthetic PII profiles and a subset of 10,000 pairs curated and labeled as matches or non-matches.
Outcome: The proposed datasets contain synthetic profiles built from publicly available sources that do not represent real individuals.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations