Papers by Adit Krishnan

4 papers
Learning from Natural Language Explanations for Generalizable Entity Matching (2024.emnlp-main)

Copied to clipboard

Challenge: Entity matching is the task of linking records from different sources that refer to the same real-world entity.
Approach: They propose to "distill" LLM reasoning into smaller entity matching models via natural language explanations.
Outcome: The proposed model distillation approach achieves strong performance on out-of-domain generalization tests (10.85% F-1).
Audience-Centric Natural Language Generation via Style Infusion (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to text style transfer (TST) with large volumes of parallel or non-parallel data are limiting for two reasons: it is difficult to collect large volumes and some stylistic objectives are hard to define without audience feedback.
Approach: They propose a task of style infusion - infusing stylistic preferences of audiences into pretrained language generation models by leveraging pairwise human judgments to bootstrap a style analysis model and augment a seed set of judgments.
Outcome: The proposed method generates compelling stylized examples with generic text prompts while balancing fluency and style adoption.
CEV-LM: Controlled Edit Vector Language Model for Shaping Natural Language Generations (2024.eacl-long)

Copied to clipboard

Challenge: Existing control approaches primarily adjust the semantic (e.g., emotion, topics), structural (e-speech, parts-of-seech), and lexical (el-s-sp-s) properties of text, but are insufficient to accomplish complex objectives such as pacing which control the complexity and readability of the text.
Approach: They propose a lightweight semi-autoregressive language model that uses edit vectors to control three complementary metrics that quantify the shape of text.
Outcome: The proposed model provides significantly more targeted and precise control of speed, volume, and circuitousness while using less training data, and containing fewer parameters.
BPID: A Benchmark for Personal Identity Deduplication (2024.emnlp-industry)

Copied to clipboard

Challenge: Data deduplication is a critical task in data management and mining, focused on consolidating duplicate records that refer to the same entity.
Approach: They propose to use a dataset with 1,000,000 unlabeled synthetic PII profiles and a subset of 10,000 pairs curated and labeled as matches or non-matches.
Outcome: The proposed datasets contain synthetic profiles built from publicly available sources that do not represent real individuals.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations