Papers by Amália Mendes

5 papers
The PALMA Corpora of African Varieties of Portuguese (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of urban varieties of Portuguese is being studied in Angola, Mozambique and So Tomé and Prncipe . the corpora are transcribed spoken data, complemented by metadata describing the setting of the audio recordings and sociolinguistic information about the speakers.
Approach: They present three new corpora of urban varieties of Portuguese spoken in Angola, Mozambique and So Tomé and Prncipe . they provide new, contemporary data for the study of each variety and for comparative research on African, Brazilian and European varieties .
Outcome: The corpora are transcribed spoken data and annotated with POS and lemma information . they are already being used for comparative research on possession and location .
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)

Copied to clipboard

Challenge: Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications.
Approach: They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages.
Outcome: The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages.
A Multi- versus a Single-classifier Approach for the Identification of Modality in the Portuguese Language (L18-1)

Copied to clipboard

Challenge: Comparative study of two different approaches to build an automatic classification system for Modality values in the Portuguese language.
Approach: They propose to use a single multi-class classifier with the full Portuguese language dataset that includes eleven modal verbs and a weighted average approach to build different classifiers for each verb.
Outcome: The proposed system is based on a Portuguese language dataset with 11 modal verbs and two different classifiers, one for each verb.
Error annotation in a Learner Corpus of Portuguese (L18-1)

Copied to clipboard

Challenge: Using the corpus architecture and the TEITOK platform, error tagging is a time-consuming task that has to be performed manually.
Approach: They propose a system that produces a final standoff, multilevel annotation with position-based tags that account for the main error types observed in the corpus.
Outcome: The proposed system annotates 47% of the corpus using the COPLE2 architecture and the TEITOK platform.
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)

Copied to clipboard

Challenge: lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels.
Approach: They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense.
Outcome: The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations