| Challenge: | a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized . |
| Approach: | They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content . |
| Outcome: | The proposed corpus is based on a pipeline methodology and is available for querying and downloading. |
Similar Papers
BlogSet-BR: A Brazilian Portuguese Blog Corpus (L18-1)
Copied to clipboard
| Challenge: | Several efforts have been made to build a corpus based on user-generated content . however, there is still a lack of a large semi-structured corpus that also contains author profiles in Brazilian Portuguese. |
| Approach: | They propose to build a Brazilian Portuguese corpus with 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs. |
| Outcome: | The proposed corpus contains 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs. |
Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing (2020.coling-main)
Copied to clipboard
| Challenge: | Existing corpus for automatic post-editing of English and Brazilian Portuguese is limited. |
| Approach: | They introduce a corpus for Automatic Post-Editing of English and Brazilian Portuguese. |
| Outcome: | The proposed corpus improves on the English and Brazilian Portuguese languages. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area. |
| Approach: | They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach. |
| Outcome: | The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues. |
A New Annotated Portuguese/Spanish Corpus for the Multi-Sentence Compression Task (L18-1)
Copied to clipboard
| Challenge: | Existing corpus for Multi-sentence Compression (MSC) tasks is limited to English . a dataset is available for MSC tasks in the French language . |
| Approach: | They propose a new corpus for Multi-Sentence Compression task in Portuguese and Spanish. |
| Outcome: | The proposed corpus is compared with two state-of-the-art systems in Portuguese and Spanish. |
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)
Copied to clipboard
| Challenge: | DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing . |
| Approach: | They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus . |
| Outcome: | The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset. |
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities . |
| Approach: | They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese. |
| Outcome: | The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities . |
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)
Copied to clipboard
| Challenge: | A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language. |
| Approach: | They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far. |
| Outcome: | The proposed model has the largest vocabulary and the best evaluation scores published so far. |
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)
Copied to clipboard
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, Matt Gardner
| Challenge: | Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets. |
| Approach: | They propose to document a dataset created by applying filters to a single snapshot of Common Crawl. |
| Outcome: | The proposed dataset shows that blocklist filtering removes text from minority individuals and patents. |
Exploring a Choctaw Language Corpus with Word Vectors and Minimum Distance Length (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools to explore low resource languages that require no expert knowledge or substantial labor are limited. |
| Approach: | They introduce additions to the Choctaw corpus by using off-the-shelf tools word2vec and Linguistica to create new computational resources for the American indigenous language. |
| Outcome: | The proposed tools can be implemented with minimal labor in the American indigenous language Choctaw. |