Papers by Amália Mendes
The PALMA Corpora of African Varieties of Portuguese (2022.lrec-1)
Copied to clipboard
Tjerk Hagemeijer, Amália Mendes, Rita Gonçalves, Catarina Cornejo, Raquel Madureira, Michel Généreux
| Challenge: | a corpus of urban varieties of Portuguese is being studied in Angola, Mozambique and So Tomé and Prncipe . the corpora are transcribed spoken data, complemented by metadata describing the setting of the audio recordings and sociolinguistic information about the speakers. |
| Approach: | They present three new corpora of urban varieties of Portuguese spoken in Angola, Mozambique and So Tomé and Prncipe . they provide new, contemporary data for the study of each variety and for comparative research on African, Brazilian and European varieties . |
| Outcome: | The corpora are transcribed spoken data and annotated with POS and lemma information . they are already being used for comparative research on possession and location . |
Multilingual Extension of PDTB-Style Annotation: The Case of TED Multilingual Discourse Bank (L18-1)
Copied to clipboard
| Challenge: | Existing corpora enriched with discourse annotations are scarce but exist . TED-MDB is hoped to be a source of parallel data for contrastive linguistic analysis and language technology applications. |
| Approach: | They propose a multilingual discourse treebank to provide a clear description of discourse structure and semantics in multiple languages. |
| Outcome: | The proposed corpus provides a clearly described level of discourse structure and semantics in multiple languages. |
A Multi- versus a Single-classifier Approach for the Identification of Modality in the Portuguese Language (L18-1)
Copied to clipboard
| Challenge: | Comparative study of two different approaches to build an automatic classification system for Modality values in the Portuguese language. |
| Approach: | They propose to use a single multi-class classifier with the full Portuguese language dataset that includes eleven modal verbs and a weighted average approach to build different classifiers for each verb. |
| Outcome: | The proposed system is based on a Portuguese language dataset with 11 modal verbs and two different classifiers, one for each verb. |
Error annotation in a Learner Corpus of Portuguese (L18-1)
Copied to clipboard
| Challenge: | Using the corpus architecture and the TEITOK platform, error tagging is a time-consuming task that has to be performed manually. |
| Approach: | They propose a system that produces a final standoff, multilevel annotation with position-based tags that account for the main error types observed in the corpus. |
| Outcome: | The proposed system annotates 47% of the corpus using the COPLE2 architecture and the TEITOK platform. |
A Lexicon of Discourse Markers for Portuguese – LDM-PT (L18-1)
Copied to clipboard
| Challenge: | lexicon of discourse markers for European Portuguese is composed of 252 pairs of discourse marker/rhetorical sense . lexical items have the function of structuring discourse and ensuring textual cohesion and coherence at intra-sentential and inter-sententential levels. |
| Approach: | They propose to create a lexicon of Portuguese discourse markers that contains 252 pairs of discourse markers/rhetorical sense. |
| Outcome: | The lexicon is compiled in an excel spread sheet and converted to an XML scheme compatible with the DiMLex format. |