Automatic Creation of Text Corpora for Low-Resource Languages from the Internet: The Case of Swiss German (2020.lrec-1)
Copied to clipboard
| Challenge: | Despite the small pool of speakers, there are still few natural language processing corpora, studies or tools for Swiss German. |
| Approach: | They propose to use a web scraper to generate the largest Swiss German text corpus . they show that the tool can be applied to other low-resource languages as well . |
| Outcome: | The proposed tool significantly improves language modeling in Swiss German, the authors show . |
Similar Papers
ParaCrawl: Web-Scale Acquisition of Parallel Corpora (2020.acl-main)
Copied to clipboard
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza
| Challenge: | We describe methods to create the largest publicly available parallel corpora by crawling the web . parallel corpus is essential for building highquality machine translation systems . |
| Approach: | They describe methods to create largest publicly available parallel corpora by crawling web sites . they empirically compare alternative methods and publish benchmark data sets . |
| Outcome: | The proposed methods improve state-of-the-art results on common benchmarks, the authors show . the pipeline has been tested on Russian, Sinhala, Nepali, Tagalog, Swahili, and Somali . |
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)
Copied to clipboard
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner
| Challenge: | Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets. |
| Approach: | They propose to document a dataset created by applying filters to a single snapshot of Common Crawl. |
| Outcome: | The proposed dataset shows that blocklist filtering removes text from minority individuals and patents. |
Swiss-AL: A Multilingual Swiss Web Corpus for Applied Linguistics (2020.lrec-1)
Copied to clipboard
| Challenge: | Swiss-AL is a multilingual web corpus for Applied Linguistics that supports data-based and data-driven research on societal and political discourses in Switzerland. |
| Approach: | They propose a multilingual Swiss web corpus for Applied Linguistics that supports data-based research on societal and political discourses in Switzerland. |
| Outcome: | The Swiss Web Corpus for Applied Linguistics (SWS) is a multilingual collection of texts from selected web sources. |
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them. |
| Approach: | They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus . |
| Outcome: | The proposed corpus includes a broader range of domains and can be trained with a pre-trained model. |
Towards a Cleaner Document-Oriented Multilingual Crawled Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing web crawling pipelines are used to collect large corpora raw data, but the main way to collect such data is through manual data extraction. |
| Approach: | They propose to use a web crawler to extract and classify data from a multilingual web corpus and an automated annotation pipeline to improve it. |
| Outcome: | The proposed version of OSCAR could be used to pre-train large generative language models and other applications in Natural Language Processing and Digital Humanities. |
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction (2021.acl-demo)
Copied to clipboard
| Challenge: | Existing tools for text extraction and web corpus construction are not enough to extract and pre-process web data to meet scientific expectations with respect to text quality. |
| Approach: | They propose a text discovery and extraction tool published under open-source license that allows for main text, comments and metadata extraction while also providing building blocks for web crawling tasks. |
| Outcome: | The proposed tool performs significantly better than other open-source solutions on real-world data and in external benchmarks. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Labadain Crawler is a data collection pipeline designed to automate and optimize the process of constructing textual corpora from the web, with a specific target to low-resource languages. |
| Approach: | They propose a data collection pipeline built on top of Nutch, an open-source web crawler and data extraction framework, and a tokenizer and identifier for Tetun. |
| Outcome: | The proposed pipeline is based on Nutch, an open-source web crawler and data extraction framework, and is tested with Tetun, one of Timor-Leste’s official languages. |
Does Corpus Quality Really Matter for Low-Resource Languages? (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on multilingual pre-training has relied on automatically filtered versions of CommonCrawl. |
| Approach: | They propose to use tailored crawling to identify and scrape websites with high-quality content to improve representation learning in Basque. |
| Outcome: | The proposed corpus, called EusCrawl, has a much higher quality according to native annotators than the Basque portion of popular multilingual corpora like CC100 and mC4. |
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models. |
| Approach: | They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 . |
| Outcome: | The proposed corpus boosts the accuracy of machine translation models on various domains. |