Tjerk Hagemeijer, Amália Mendes, Rita Gonçalves, Catarina Cornejo, Raquel Madureira, Michel Généreux
| Challenge: | a corpus of urban varieties of Portuguese is being studied in Angola, Mozambique and So Tomé and Prncipe . the corpora are transcribed spoken data, complemented by metadata describing the setting of the audio recordings and sociolinguistic information about the speakers. |
| Approach: | They present three new corpora of urban varieties of Portuguese spoken in Angola, Mozambique and So Tomé and Prncipe . they provide new, contemporary data for the study of each variety and for comparative research on African, Brazilian and European varieties . |
| Outcome: | The corpora are transcribed spoken data and annotated with POS and lemma information . they are already being used for comparative research on possession and location . |
Similar Papers
Elote, Choclo and Mazorca: on the Varieties of Spanish (2024.naacl-long)
Copied to clipboard
| Challenge: | Spanish is the official language in 20 countries and the second most-spoken native language . available corpora treat it as one monolithic language, damping prediction power . |
| Approach: | They compile and curate datasets in different varieties of Spanish around the world at an unprecedented scale and create the CEREAL corpus. |
| Outcome: | The results show that Spanish is a multilingual language with a wide range of cultural and cultural influences. |
Corpora and Baselines for Humour Recognition in Portuguese (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on the recognition of verbal humour in Portuguese has not been done . humor recognition is a sign of fluency in a language, and is not yet widely used in other languages. |
| Approach: | They propose to create three corpora covering two styles of humour and four sources of non-humorous text that are used for testing computational models. |
| Outcome: | The proposed models can be used to train and test models in Portuguese, and may be used as baselines for future projects. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The brWaC Corpus: A New Open Resource for Brazilian Portuguese (L18-1)
Copied to clipboard
| Challenge: | a large corpus for Brazilian Portuguese is needed for NLP applications . the corpus is 2.7 billion tokens, and domain diversity is maximized . |
| Approach: | They propose to build a large Web corpus for Brazilian Portuguese with 2.7 billion tokens . they also propose an updated sentence-level approach for the strict removal of duplicated content . |
| Outcome: | The proposed corpus is based on a pipeline methodology and is available for querying and downloading. |
Natural Language Generation: Recently Learned Lessons, Directions for Semantic Representation-based Approaches, and the Case of Brazilian Portuguese Language (P19-2)
Copied to clipboard
| Challenge: | Natural Language Generation (NLG) is a promising area in Natural Language Processing (NLP) . |
| Approach: | They present a review of the literature on Natural Language Generation in Brazilian Portuguese. |
| Outcome: | The proposed approaches are based on the Abstract Meaning Representation formalism and have potential future directions. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |
A Brief Survey of Textual Dialogue Corpora (2022.lrec-1)
Copied to clipboard
| Challenge: | Several dialogue corpora are available for research purposes, but they do not cover all the necessities of real-world applications. |
| Approach: | They analyze available dialogue corpora and propose possible approaches to create new ones. |
| Outcome: | The proposed corpus of human-human dialogues is based on a list of available dialogue corpora . it covers speakers, size, languages, collection, annotations, and domains . some trends are identified and possible approaches are also discussed . |
The BDCamões Collection of Portuguese Literary Documents: a Research Resource for Digital Humanities and Language Technology (2020.lrec-1)
Copied to clipboard
| Challenge: | a new corpus of literary documents in Portuguese is presented . it includes close to 4 million words from over 200 complete documents . the corpus is suitable for research in language technology and digital humanities . |
| Approach: | They present the BDCames Collection of Portuguese Literary Documents, a new corpus of literary texts written in Portuguese. |
| Outcome: | The BDCames Collection of Portuguese Literary Documents is a new corpus of literary documents written in Portuguese . it includes close to 4 million words from over 200 complete documents from 83 authors in 14 genres . the corpus is suitable for research in language technology and language science and digital humanities . |
WaCadie: Towards an Acadian French Corpus (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora do not exist for many languages and language varieties, such as Acadian French. |
| Approach: | They propose to build a corpus of Acadian French using web-as-corpus methodologies . they use domain crawling, social media scraping, and search engines to create corpus . |
| Outcome: | The proposed corpus includes some traces of Acadian French, but it is not available for many languages and language varieties, such as Acadinian French. |