The PALMA Corpora of African Varieties of Portuguese (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of urban varieties of Portuguese is being studied in Angola, Mozambique and So Tomé and Prncipe . the corpora are transcribed spoken data, complemented by metadata describing the setting of the audio recordings and sociolinguistic information about the speakers.
Approach: They present three new corpora of urban varieties of Portuguese spoken in Angola, Mozambique and So Tomé and Prncipe . they provide new, contemporary data for the study of each variety and for comparative research on African, Brazilian and European varieties .
Outcome: The corpora are transcribed spoken data and annotated with POS and lemma information . they are already being used for comparative research on possession and location .

Similar Papers

Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)

Copied to clipboard

Challenge: Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power .
Approach: They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus.
Outcome: The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences.
Corpora and Baselines for Humour Recognition in Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on the recognition of verbal humour in Portuguese has not been done . humor recognition is a sign of fluency in a language, and is not yet widely used in other languages.
Approach: They propose to create three corpora covering two styles of humour and four sources of non-humorous text that are used for testing computational models.
Outcome: The proposed models can be used to train and test models in Portuguese, and may be used as baselines for future projects.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)

Copied to clipboard

Challenge: a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized .
Approach: They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content .
Outcome: The proposed corpus is based on a pipeline methodology and is available for querying and downloading.
Natural Language Generation: Recently Learned Lessons, Directions for Semantic Representation-based Approaches, and the Case of Brazilian Portuguese Language (P19-2)

Copied to clipboard

Challenge: Natural Language Generation (NLG) is a promising area in Natural Language Processing (NLP) .
Approach: They present a review of the literature on Natural Language Generation in Brazilian Portuguese.
Outcome: The proposed approaches are based on the Abstract Meaning Representation formalism and have potential future directions.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)

Copied to clipboard

Challenge: Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications.
Approach: They analyze available dialogue corpora and propose possible approaches to create new ones.
Outcome: The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed .
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)

Copied to clipboard

Challenge: a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities .
Approach: They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese.
Outcome: The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities .
WaCadie: Towards an Acadian French Corpus (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora do not exist for many languages and language varieties, such as Acadian French.
Approach: They propose to build a corpus of Acadian French using web-as-corpus methodologies . they use domain crawling, social media scraping, and search engines to create corpus .
Outcome: The proposed corpus includes some traces of Acadian French, but it is not available for many languages and language varieties, such as Acadinian French.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations