Papers by William Agnew

    1 papers
    Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)

    Copied to clipboard

    Challenge: Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets.
    Approach: They propose to document a dataset created by applying filters to a single snapshot of Common Crawl.
    Outcome: The proposed dataset shows that blocklist filtering removes text from minority individuals and patents.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations