Papers by Robert Gaizauskas
SNuC: The Sheffield Numbers Spoken Language Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | SNuC is the first published corpus of spoken alphanumeric identifiers . it contains recordings and transcriptions of over 50 native British English speakers . |
| Approach: | They present a corpus of spoken alphanumeric identifiers of the sort typically used as serial and part numbers in the manufacturing sector. |
| Outcome: | The proposed corpus can be used to improve spoken alphanumeric identifier recognition. |
BLN600: A Parallel Corpus of Machine/Human Transcribed Nineteenth Century Newspaper Texts (2024.lrec-main)
Copied to clipboard
| Challenge: | Historical documents present unique challenges to automated digital transcription technologies, such as optical character recognition (OCR). |
| Approach: | They propose to use a publicly available nineteenth-century newspaper corpus to train and develop OCR and post-OCR correction methodologies for historical newspaper machine transcription. |
| Outcome: | The proposed corpus will be useful for training and development of OCR and post-OCR correction methodologies for historical newspaper machine transcription. |
A Language Modelling Approach to Quality Assessment of OCR’ed Historical Text (2022.lrec-1)
Copied to clipboard
| Challenge: | a language model-based approach is used to score the quality of OCR transcriptions in the British Library Newspapers corpus . a corpus of genre-adjacent texts captures the common and legal parlance of nineteenth-century London . |
| Approach: | They propose a language model-based approach to score the quality of OCR transcriptions in the British Library Newspapers corpus parts 1 and 2 . they aim to link newspapers of crime in nineteenth-century London to the Digital Panopticon . |
| Outcome: | The proposed approach is based on the Proceedings of the Old Bailey Online corpus, which captures the common and legal parlance of nineteenth-century London. |