Validating and Exploring Large Geographic Corpora (2024.lrec-main)

Copied to clipboard

Challenge: a paper examines the impact of corpus creation decisions on multi-lingual web corpora . the goal is to understand the impact on downstream corporata with a focus on under-represented languages and populations.
Approach: This paper evaluates the impact of corpus creation decisions on multi-lingual web corpora . three cleaning methods are used to improve the quality of sub-corpora in the common crawl . the goal is to understand the impact on downstream corporan with a focus on under-represented languages .
Outcome: The results show that the validity of sub-corpora is improved with each stage of cleaning but that this improvement is unevenly distributed across languages and populations.

Similar Papers

Does Corpus Quality Really Matter for Low-Resource Languages? (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on multilingual pre-training has relied on automatically filtered versions of CommonCrawl.
Approach: They propose to use tailored crawling to identify and scrape websites with high-quality content to improve representation learning in Basque.
Outcome: The proposed corpus, called EusCrawl, has a much higher quality according to native annotators than the Basque portion of popular multilingual corpora like CC100 and mC4.
Quality Beyond A Glance: Revealing Large Quality Differences Between Web-Crawled Parallel Corpora (2025.coling-main)

Copied to clipboard

Challenge: Parallel corpora play a vital role in advanced multilingual natural language processing tasks, notably in machine translation (MT).
Approach: They manually and automatically evaluated four well-known publicly available parallel corpora across eleven language pairs.
Outcome: The results show that the four well-known parallel corpora have a substantial amount of noisy sentence pairs, while CCMatrix and CCAligned have low quality sentences.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)

Copied to clipboard

Challenge: Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets.
Approach: They propose to document a dataset created by applying filters to a single snapshot of Common Crawl.
Outcome: The proposed dataset shows that blocklist filtering removes text from minority individuals and patents.
Geographically-Balanced Gigaword Corpora for 50 Language Varieties (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora of text corpors over-represent inner-circle varieties from the US and UK . this paper uses country-level population demographics to correct implicit geographic and demographic biases .
Approach: They propose to use country-level population demographics to correct geographic biases . they use a population-based sampling technique to remove geographic bias from gigaword corpora .
Outcome: The proposed corpus family removes geographic biases by comparing the population-based sampling with the baseline corpus.
Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages (2024.lrec-main)

Copied to clipboard

Challenge: Large, curated, web-crawled corpora play a vital role in training language models . however, relatively little attention has been given to the quality of these corporata .
Approach: They compare four of the currently most relevant large, web-crawled corpora across eleven lower-resourced European languages to evaluate their quality.
Outcome: The CC100 corpus achieves the highest scores on the tests in 11 lower-resourced European languages.
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)

Copied to clipboard

Challenge: Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs.
Approach: They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus.
Outcome: The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats.
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction.
Approach: They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it.
Outcome: The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities.
What’s in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus (2021.acl-short)

Copied to clipboard

Challenge: Recent advances in NLP have been driven by Transformer-based language models.
Approach: They analyze the Common Crawl, a web corpus extensively used for training language models.
Outcome: The Common Crawl contains hate speech and sexually explicit content even after filtering procedures.
Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German (2020.lrec-1)

Copied to clipboard

Challenge: Despite the small pool of speakers, there are still few natural language processing corpora, studies or tools for Swiss German.
Approach: They propose to use a web scraper to generate the largest Swiss German text corpus . they show that the tool can be applied to other low-resource languages as well .
Outcome: The proposed tool significantly improves language modeling in Swiss German, the authors show .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations