Papers by Margaret Mitchell

4 papers
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)

Copied to clipboard

Challenge: Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets.
Approach: They propose to document a dataset created by applying filters to a single snapshot of Common Crawl.
Outcome: The proposed dataset shows that blocklist filtering removes text from minority individuals and patents.
SEAL: Interactive Tool for Systematic Error Analysis and Labeling (2022.emnlp-demos)

Copied to clipboard

Challenge: Existing models that fail on tail data or rare groups are difficult to identify due to lack of explicit labels.
Approach: They propose a systematic error analysis and labeling tool that uses a two-step approach to identify high-error slices of data and then give human-understandable semantics to those underperforming slices.
Outcome: The proposed tool identifies high-error slices of data and gives human-understandable semantics to those underperforming slices.
Perturbation Sensitivity Analysis to Detect Unintended Model Biases (D19-1)

Copied to clipboard

Challenge: Recent research shows that data-driven NLP models may inadvertently capture, reflect and sometimes amplify various social biases present in the language data they are trained on.
Approach: They propose a generic evaluation framework that detects unintended model biases related to named entities and requires no new annotations or corpora.
Outcome: The proposed framework detects unintended model biases related to named entities and requires no new annotations or corpora.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations