Papers by Lauren Cassidy
gaBERT — an Irish Language Model (2022.lrec-1)
Copied to clipboard
James Barry, Joachim Wagner, Lauren Cassidy, Alan Cowap, Teresa Lynn, Abigail Walsh, Mícheál J. Ó Meachair, Jennifer Foster
| Challenge: | We compare gaBERT to multilingual BERT and the monolingual Irish WikiBERT and show that gaBERt provides better representations for downstream parsing tasks. |
| Approach: | They propose a monolingual BERT model for the Irish language that provides better representations for a downstream parsing task. |
| Outcome: | The proposed model performs better than the multilingual BERT and the monolingual Irish WikiBERT on a parsing task. |
Treebanking User-Generated Content: A Proposal for a Unified Representation in Universal Dependencies (2020.lrec-1)
Copied to clipboard
Manuela Sanguinetti, Cristina Bosco, Lauren Cassidy, Özlem Çetinoğlu, Alessandra Teresa Cignarella, Teresa Lynn, Ines Rehbein, Josef Ruppenhofer, Djamé Seddah, Amir Zeldes
| Challenge: | Despite the increasing number of contributions on Part-of-Speech tagging and parsing, automatic processing of user-generated content (UGC) still represents a challenging task. |
| Approach: | They propose a set of guidelines for the annotation of user-generated texts within the Universal Dependencies framework. |
| Outcome: | The proposed annotation guidelines promote cross-linguistic consistency, which has always been in the spirit of UD. |
TwittIrish: A Universal Dependencies Treebank of Tweets in Modern Irish (2022.acl-long)
Copied to clipboard
| Challenge: | Modern Irish is a minority language lacking computational resources for accurate automatic syntactic parsing of user-generated content. |
| Approach: | They propose to use a treebank to facilitate natural language parsing of user-generated content in Irish. |
| Outcome: | The proposed treebank enables natural language processing of user-generated content in Irish. |