Papers by Adrian Cheung

2 papers
Inconsistencies in Crowdsourced Slot-Filling Annotations: A Typology and Identification Methods (2020.coling-main)

Copied to clipboard

Challenge: Standard slot-filling models train or finetune on large datasets of carefully-annotated data that is domain specific.
Approach: They propose automatic methods to identify inconsistencies in crowd-annotated data . a slot-filling model can extract the tokens "New York" as a TO LOCATION slot in a query .
Outcome: The proposed methods reveal inconsistencies in data, though there is scope for improvement.
Iterative Feature Mining for Constraint-Based Data Collection to Increase Data Diversity and Model Robustness (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work on dialog has found that crowdsourced data can have limited diversity as workers tend to write simple variations from prompts.
Approach: They propose a general approach for guiding workers to write more diverse text by iteratively constraining their writing.
Outcome: The proposed approach improves performance on dialog tasks and improves on existing datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations