Papers by Scott Miller

4 papers
DEGREE: A Data-Efficient Generation-Based Event Extraction Model (2022.naacl-main)

Copied to clipboard

Challenge: Existing models for event extraction require expensive human annotations.
Approach: They propose a data-efficient event extraction model that formulates event extraction as a conditional generation problem.
Outcome: The proposed model can be trained with only a few labeled examples.
Cross-lingual Joint Entity and Word Embedding to Improve Entity Linking and Parallel Sentence Mining (D19-61)

Copied to clipboard

Challenge: Entities can be used as effective signals to generate less ambiguous semantic representations and align multiple languages.
Approach: They propose a method to generate cross-lingual data that is a mix of entities and contextual words based on Wikipedia.
Outcome: The proposed method can generate cross-lingual data that is a mix of entities and contextual words based on Wikipedia . it provides reliable alignment on word/entity level and sentence level, and thus can be used for unsupervised cross-linguistic entity linking.
The Challenges of Optimizing Machine Translation for Low Resource Cross-Language Information Retrieval (D19-1)

Copied to clipboard

Challenge: Existing studies do not investigate the effectiveness of MT metrics in predicting performance of downstream IR models.
Approach: They examine the relationship between MT performance and IR quality in a CLIR-based system . they find that the choice of IR collection can significantly affect MT tuning decisions .
Outcome: The proposed model can predict CLIR performance better from MT quality, the authors show . the proposed model is based on a BLEU-based model with a bag of words constraint .
SARAL: A Low-Resource Cross-Lingual Domain-Focused Information Retrieval System for Effective Rapid Document Triage (P19-3)

Copied to clipboard

Challenge: a new cross-lingual information retrieval system for low-resource languages is available in less-frequently-taught languages . a multilingual system can search for relevant information in a haystack of documents in swahili or Somali . human-driven approaches to this problem are complicated in 'low-resourced' languages aaron sagar: "the key role played by humans in triaging results is complicated"
Approach: They propose an end-to-end cross-lingual information retrieval system for low-resource languages . the system enables English speakers to search foreign language repositories using English queries . it summarizes the retrieved documents in English with respect to a particular information need .
Outcome: The proposed system achieves top performance in the most recent IARPA MATERIAL CLIR+summarization evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations