| Challenge: | a paper examines the impact of corpus creation decisions on multi-lingual web corpora . the goal is to understand the impact on downstream corporata with a focus on under-represented languages and populations. |
| Approach: | This paper evaluates the impact of corpus creation decisions on multi-lingual web corpora . three cleaning methods are used to improve the quality of sub-corpora in the common crawl . the goal is to understand the impact on downstream corporan with a focus on under-represented languages . |
| Outcome: | The results show that the validity of sub-corpora is improved with each stage of cleaning but that this improvement is unevenly distributed across languages and populations. |
Similar Papers
Does Corpus Quality Really Matter for Low-Resource Languages? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on multilingual pre-training has relied on automatically filtered versions of CommonCrawl. |
| Approach: | They propose to use tailored crawling to identify and scrape websites with high-quality content to improve representation learning in Basque. |
| Outcome: | The proposed corpus, called EusCrawl, has a much higher quality according to native annotators than the Basque portion of popular multilingual corpora like CC100 and mC4. |
Quality Beyond A Glance: Revealing Large Quality Differences Between Web-Crawled Parallel Corpora (2025.coling-main)
Copied to clipboard
| Challenge: | Parallel corpora play a vital role in advanced multilingual natural language processing tasks, notably in machine translation (MT). |
| Approach: | They manually and automatically evaluated four well-known publicly available parallel corpora across eleven language pairs. |
| Outcome: | The results show that the four well-known parallel corpora have a substantial amount of noisy sentence pairs, while CCMatrix and CCAligned have low quality sentences. |
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)
Copied to clipboard
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner
| Challenge: | Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets. |
| Approach: | They propose to document a dataset created by applying filters to a single snapshot of Common Crawl. |
| Outcome: | The proposed dataset shows that blocklist filtering removes text from minority individuals and patents. |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022.tacl-1)
Copied to clipboard
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofetoluwa Adeyemi
| Challenge: | Lower-resource corpora have systematic issues, including mislabeled or nonstandard/ambiguous language codes. |
| Approach: | They manually audit the quality of 205 language-specific corpora released with five major public datasets. |
| Outcome: | The results show that lower-resource corpora have systematic issues even for non-proficient speakers. |
Geographically-Balanced Gigaword Corpora for 50 Language Varieties (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora of text corpors over-represent inner-circle varieties from the US and UK . this paper uses country-level population demographics to correct implicit geographic and demographic biases . |
| Approach: | They propose to use country-level population demographics to correct geographic biases . they use a population-based sampling technique to remove geographic bias from gigaword corpora . |
| Outcome: | The proposed corpus family removes geographic biases by comparing the population-based sampling with the baseline corpus. |
Do Language Models Care about Text Quality? Evaluating Web-Crawled Corpora across 11 Languages (2024.lrec-main)
Copied to clipboard
Rik van Noord, Taja Kuzman, Peter Rupnik, Nikola Ljubešić, Miquel Esplà-Gomis, Gema Ramírez-Sánchez, Antonio Toral
| Challenge: | Large, curated, web-crawled corpora play a vital role in training language models . however, relatively little attention has been given to the quality of these corporata . |
| Approach: | They compare four of the currently most relevant large, web-crawled corpora across eleven lower-resourced European languages to evaluate their quality. |
| Outcome: | The CC100 corpus achieves the highest scores on the tests in 11 lower-resourced European languages. |
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs. |
| Approach: | They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus. |
| Outcome: | The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats. |
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction. |
| Approach: | They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it. |
| Outcome: | The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities. |
What’s in the Box? An Analysis of Undesirable Content in the Common Crawl Corpus (2021.acl-short)
Copied to clipboard
| Challenge: | Recent advances in NLP have been driven by Transformer-based language models. |
| Approach: | They analyze the Common Crawl, a web corpus extensively used for training language models. |
| Outcome: | The Common Crawl contains hate speech and sexually explicit content even after filtering procedures. |
Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German (2020.lrec-1)
Copied to clipboard
| Challenge: | Despite the small pool of speakers, there are still few natural language processing corpora, studies or tools for Swiss German. |
| Approach: | They propose to use a web scraper to generate the largest Swiss German text corpus . they show that the tool can be applied to other low-resource languages as well . |
| Outcome: | The proposed tool significantly improves language modeling in Swiss German, the authors show . |