| Challenge: | Currently, there is little to no data available to build natural language processing models for endangered languages. |
| Approach: | They propose a benchmark dataset of transcriptions for scanned books in three critically endangered languages and a method to improve OCR in these data-scarce settings. |
| Outcome: | The proposed method reduces the recognition error rate by 34% across the three endangered languages. |
Similar Papers
Phoneme transcription of endangered languages: an evaluation of recent ASR architectures in the single speaker scenario (2022.findings-acl)
Copied to clipboard
| Challenge: | Recent work on phonetic transcription is reported to be the bottleneck in endangered languages . however, when a single speaker is involved, small amounts of training are needed . |
| Approach: | They compare automatic speech recognition (ASR) approaches to speaker-dependent phonetic transcription using a common dataset of 11 languages. |
| Outcome: | The proposed system handles morphologically complex languages and writing systems for which no pronunciation dictionary exists. |
Endangered Languages meet Modern NLP (2020.coling-tutorials)
Copied to clipboard
| Challenge: | This tutorial will focus on NLP for endangered languages documentation and revitalization. |
| Approach: | This tutorial will focus on NLP for endangered languages documentation and revitalization . the goal is to motivate more NLP practitioners to work towards this important direction . |
| Outcome: | This tutorial will acquaint attendees with the process and the challenges of language documentation and revitalization. |
Low-resource Post Processing of Noisy OCR Output for Historical Corpus Digitisation (L18-1)
Copied to clipboard
| Challenge: | 7.6% of the words in the original OCR text contain an error; fully manual correction would take thousands of hours due to the size of the corpus. |
| Approach: | They propose a post-processing system to efficiently correct OCR errors in a 2.7 million word Faroese corpus. |
| Outcome: | The proposed method reduces the word error rate to 1.3% with around 65 hours of human annotator work. |
OCR Improves Machine Translation for Low-Resource Languages (2022.findings-acl)
Copied to clipboard
| Challenge: | Despite many recent successes, Machine Translation still lacks support or fails to achieve good performance for most low-resource languages. |
| Approach: | They propose a benchmark to evaluate OCR systems on low-resource languages and low- resource scripts. |
| Outcome: | The proposed benchmark evaluates state-of-the-art OCR systems on low-resource languages and low-rural scripts. |
Unsupervised Multi-View Post-OCR Error Correction With Language Models (2021.emnlp-main)
Copied to clipboard
| Challenge: | Prior work used text generation techniques or redundancy in similar passages for OCR error correction, which is not appropriate in cases of low corpus redundancies or weak document contextual information. |
| Approach: | They propose to use a pretrained language model to reconcile different OCR views in unsupervised way so that their combination contains fewer errors than each individual view. |
| Outcome: | The proposed model can reconcile multiple OCR views so that their combined version contains fewer errors than the best OCR view. |
A Benchmark and Dataset for Post-OCR text correction in Sanskrit (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Sanskrit is a classical language with 30 million manuscripts available for digitisation . however, it is considered to be low-resource when it comes to available digital resources. |
| Approach: | They propose to use a post-OCR text correction dataset to correct errors from OCR predictions from 30 different books in the Indian subcontinent. |
| Outcome: | The proposed model outperforms OCR models on graphemic and lexical levels and shows that it is more accurate than previous models. |
Lexically Aware Semi-Supervised Learning for OCR Post-Correction (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing methods for digitizing text in endangered languages rely on manual data curated by the user. |
| Approach: | They propose a semi-supervised learning method that utilizes raw images to improve performance. |
| Outcome: | The proposed method reduces errors by 15%–29% on four endangered languages. |
Cleaning Dirty Books: Post-OCR Processing for Previously Scanned Texts (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a large amount of work is required to clean digitized books for NLP analysis because of errors in the scanned text and duplicate volumes in the corpora. |
| Approach: | They propose methods to handle optical character recognition errors in scanned texts . they identify the canonical version for each of 17,136 repeatedly-scanned books . |
| Outcome: | The proposed method corrects over six times as many errors as it introduces, the authors show . the authors evaluate a collection of 19,347 texts from the Gutenberg dataset and 96,635 from the HathiTrust Library . |
Hire a Linguist!: Learning Endangered Languages in LLMs with In-Context Linguistic Descriptions (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing LLMs rarely perform well in unseen, endangered languages . Existing models such as Llama and GPT-4 lack a rich corpus of training data . |
| Approach: | They propose a training-free approach to enable an LLM to process unseen languages that hardly occur in its pre-training. |
| Outcome: | The proposed approach elevates translation capability from GPT-4’s 0 to 10.5 BLEU for 10 language directions. |
Enabling Interactive Transcription in an Indigenous Community (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for manual transcription are often in isolation from the speech community, and so we miss out on the opportunity to take advantage of the interests and skills of local people. |
| Approach: | They propose a transcription workflow which combines spoken term detection and human-in-the-loop to support speech transcription in almost-zero resource settings. |
| Outcome: | The proposed workflow is based on two endangered languages with zero-resource datasets. |