Papers by Laura García-Sardiña
BaSCo: An Annotated Basque-Spanish Code-Switching Corpus for Natural Language Understanding (2022.lrec-1)
Copied to clipboard
| Challenge: | Basque-Spanish code-switching is a widespread phenomenon among bilingual speakers in the Basque Country. |
| Approach: | They propose to use annotated utterances to train bilingual chatbots in Basque and Spanish to cover the phenomenon of code-switching. |
| Outcome: | The proposed corpus is the first with annotated linguistic resources encompassing Basque-Spanish code-switching. |
HitzalMed: Anonymisation of Clinical Text in Spanish (2020.lrec-1)
Copied to clipboard
| Challenge: | HITZALMED is a web-framed tool that performs automatic detection of sensitive information in clinical texts using machine learning algorithms reported to be competitive for the task. |
| Approach: | This paper presents a web-framed tool that performs automatic detection of sensitive information in clinical texts using machine learning algorithms reported to be competitive for the task. |
| Outcome: | The proposed tool is available online and can be configured by the user. |
ES-Port: a Spontaneous Spoken Human-Human Technical Support Corpus for Dialogue Research in Spanish (L18-1)
Copied to clipboard
| Challenge: | ES-Port is a spontaneous spoken human-human dialogue corpus in Spanish that consists of 1170 dialogues from calls to the technical support department of a telecommunications provider. |
| Approach: | They describe the compilation process from transcription to anonymisation of sensitive data contained in the transcriptions. |
| Outcome: | The ES-Port corpus is a human-human dialogue corpus that consists of 1170 calls to the technical support department of a telecommunications provider. |