The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized .
Approach: They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content .
Outcome: The proposed corpus is based on a pipeline methodology and is available for querying and downloading.

Similar Papers

BlogSet-BR: A Brazilian Portuguese Blog Corpus (L18-1)

Copied to clipboard

Challenge: Several efforts have been made to build a corpus based on user-generated content . however, there is still a lack of a large semi-structured corpus that also contains author profiles in Brazilian Portuguese.
Approach: They propose to build a Brazilian Portuguese corpus with 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
Outcome: The proposed corpus contains 2.1 billion words extracted from 7.4 million posts over 808 thousand different Brazilian blogs.
Building The First English-Brazilian Portuguese Corpus for Automatic Post-Editing (2020.coling-main)

Copied to clipboard

Challenge: Existing corpus for automatic post-editing of English and Brazilian Portuguese is limited.
Approach: They introduce a corpus for Automatic Post-Editing of English and Brazilian Portuguese.
Outcome: The proposed corpus improves on the English and Brazilian Portuguese languages.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
Towards AMR-BR: A SemBank for Brazilian Portuguese Language (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a recent and prominent meaning representation with good acceptance and several applications in the Natural Language Processing area.
Approach: They propose to build an AMR annotated corpus for Brazilian Portuguese using an alignment-based approach.
Outcome: The proposed corpus is based on the Little Prince book, which went into the public domain and explored some language-specific annotation issues.
A New Annotated Portuguese/Spanish Corpus for the Multi-Sentence Compression Task (L18-1)

Copied to clipboard

Challenge: Existing corpus for Multi-sentence Compression (MSC) tasks is limited to English . a dataset is available for MSC tasks in the French language .
Approach: They propose a new corpus for Multi-Sentence Compression task in Portuguese and Spanish.
Outcome: The proposed corpus is compared with two state-of-the-art systems in Portuguese and Spanish.
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)

Copied to clipboard

Challenge: DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing .
Approach: They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus .
Outcome: The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset.
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities .
Approach: They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese.
Outcome: The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities .
Finely Tuned, 2 Billion Token Based Word Embeddings for Portuguese (L18-1)

Copied to clipboard

Challenge: A distributional semantics model is instrumental to improve the performance of many applications and processing tasks for any language.
Approach: They propose to develop an advanced distributional model for Portuguese with the largest vocabulary and best evaluation scores published so far.
Outcome: The proposed model has the largest vocabulary and the best evaluation scores published so far.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus (2021.emnlp-main)

Copied to clipboard

Challenge: Large text corpora are often introduced with minimal documentation . documenting collection process, composition, intended uses, and other are key for structured, task-specific datasets.
Approach: They propose to document a dataset created by applying filters to a single snapshot of Common Crawl.
Outcome: The proposed dataset shows that blocklist filtering removes text from minority individuals and patents.
Exploring a Choctaw Language Corpus with Word Vectors and Minimum Distance Length (2020.lrec-1)

Copied to clipboard

Challenge: Existing tools to explore low resource languages that require no expert knowledge or substantial labor are limited.
Approach: They introduce additions to the Choctaw corpus by using off-the-shelf tools word2vec and Linguistica to create new computational resources for the American indigenous language.
Outcome: The proposed tools can be implemented with minimal labor in the American indigenous language Choctaw.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations