Papers by Sandra Aluísio

3 papers
MuPe Life Stories Dataset: Spontaneous Speech in Brazilian Portuguese with a Case Study Evaluation on ASR Bias against Speakers Groups and Topic Modeling (2025.coling-main)

Copied to clipboard

Challenge: Recent datasets for automatic speech recognition in Brazilian Portuguese lack diversity in terms of age groups, regional accents, and education levels.
Approach: They propose to use a dataset to analyze the impact of ASR in Brazilian Portuguese (BP) they demonstrate that current models are biased regarding age, education, and regional accents.
Outcome: The proposed dataset helps mitigate biases in current ASR models regarding education levels and age groups.
Evaluating Sentence Segmentation in Different Datasets of Neuropsychological Language Tests in Brazilian Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Using automated analysis of connected speech is a promising direction for diagnosing cognitive impairments.
Approach: They propose to use a novel model to segment impaired speech transcriptions . they propose to include a Linear Chain CRF and a self-attention mechanism .
Outcome: The proposed system performs better than the existing model with three new datasets used to diagnose cognitive impairments.
Using Eye-tracking Data to Predict the Readability of Brazilian Portuguese Sentences in Single-task, Multi-task and Sequential Transfer Learning Approaches (2020.coling-main)

Copied to clipboard

Challenge: Sentence complexity assessment is a relatively new task in Natural Language Processing.
Approach: They propose to use Brazilian Portuguese to evaluate sentences with linguistic features to improve readability.
Outcome: The proposed model reaches the state-of-the-art for Brazilian Portuguese with 97.8% accuracy with linguistic features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations