Learning From Failure: Data Capture in an Australian Aboriginal Community (2022.acl-long)
Copied to clipboard
| Challenge: | a prototype of a language data capture app for speakers was tested in an Aboriginal community . elicitation of word lists, phrases, etc. has been used for decades to collect data for Indigenous languages . many software tools are developed to support linguists' work . |
| Approach: | They propose to deploy an app for speakers to confirm system guesses in an approach to transcription based on word spotting. |
| Outcome: | The proposed app was tested in an Aboriginal community in australia . it was able to confirm system guesses without a transcription bottleneck . the results were compared with other apps in the community . |
Similar Papers
Fashioning Local Designs from Generic Speech Technologies in an Australian Aboriginal Community (2022.coling-1)
Copied to clipboard
| Challenge: | Recent research has focused on low-resource languages and the transcription bottleneck paradigm. |
| Approach: | They propose to use a spoken term detection system to train a speech recognition system in an Aboriginal community to reach better comprehension and engagement from Aboriginal participants. |
| Outcome: | The proposed system can be implemented in an Aboriginal community and reach better comprehension and engagement from Aboriginal participants. |
Enabling Interactive Transcription in an Indigenous Community (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods for manual transcription are often in isolation from the speech community, and so we miss out on the opportunity to take advantage of the interests and skills of local people. |
| Approach: | They propose a transcription workflow which combines spoken term detection and human-in-the-loop to support speech transcription in almost-zero resource settings. |
| Outcome: | The proposed workflow is based on two endangered languages with zero-resource datasets. |
Learnings from Technological Interventions in a Low Resource Language: A Case-Study on Gondi (2020.lrec-1)
Copied to clipboard
Devansh Mehta, Sebastin Santy, Ramaravind Kommiya Mothilal, Brij Mohan Lal Srivastava, Alok Sharma, Anurag Shukla, Vishnu Prasad, Venkanna U, Amit Sharma, Kalika Bali
| Challenge: | 40% of all the languages in the world face the danger of extinction in the near future . when a language dies out, future generations lose a vital part of the culture that is necessary to completely understand it. |
| Approach: | They propose to use 4 technology-driven methods of data collection to collect data on Gondi, a low-resource vulnerable language spoken by 2.3 million tribal people in south and central India. |
| Outcome: | The proposed methods collected 12,000 translated words and/or sentences and identified more than 650 community members whose help can be solicited for future translation efforts. |
Not always about you: Prioritizing community needs when developing endangered language technology (2022.acl-long)
Copied to clipboard
| Challenge: | low-resource languages lack the quantity of data needed to train statistical and machine learning tools and models. |
| Approach: | They propose to use language technology to support endangered languages' revitalization . they propose to work with indigenous speakers to develop technology for such training . |
| Outcome: | The authors discuss the challenges that researchers and indigenous speech community members face when working together to develop language technology to support endangered languages. |
”It’s how you do things that matters”: Attending to Process to Better Serve Indigenous Communities with Language Technologies (2024.eacl-short)
Copied to clipboard
| Challenge: | Indigenous languages are historically under-served by natural language processing (NLP) but this is changing with the recent scaling of large multilingual models and an increased focus by the NLP community on endangered languages. |
| Approach: | They propose to build NLP technologies for Indigenous languages that should primarily serve Indigenous communities. |
| Outcome: | The proposed approach is based on interviews with 17 researchers working in or with Aboriginal and/or Torres Strait Islander communities on language technology projects in Australia. |
Building Representative Corpora from Illiterate Communities: A Reviewof Challenges and Mitigation Strategies for Developing Countries (2021.eacl-main)
Copied to clipboard
| Challenge: | Existing methods for collecting data from high-income countries (HICs) make implicit assumptions about literacy and internet access, but in low-income and sub-Saharan Africa (SSA) such assumptions may not hold for LICs where the bulk of the population lives. |
| Approach: | They propose a set of practical mitigation strategies to address the under-representation of illiterate communities in NLP corpora. |
| Outcome: | The proposed methods address the under-representation of illiterate communities in NLP corpora and propose mitigation strategies to help future work. |
Centering the Speech Community (2024.eacl-long)
Copied to clipboard
| Challenge: | In remote speech communities, people interact with the outside world using a variety of an institutional language. |
| Approach: | They propose to use local languages to support their collaboration in a remote community in the far north of australia to explore the functional differences between oral and institutional languages. |
| Outcome: | The proposed language technologies are better aligned with local interests and aspirations than the first author's western framing of language as data for exploitation by machines. |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Local Word Discovery for Interactive Transcription (2021.emnlp-main)
Copied to clipboard
| Challenge: | a new computational task supports the construction of high quality texts and lexicons for low resource languages. |
| Approach: | They propose a computational task which is tuned to the available knowledge and interests in an Indigenous community. |
| Outcome: | The proposed method achieves a transcription density gain of 17% in a morphologically complex language . the proposed grammar includes a description of the phonology and morphosyntax . |
Decolonising Speech and Language Technology (2020.coling-main)
Copied to clipboard
| Challenge: | Indigenous peoples are increasingly unable to go on without speech and language technologies, says a researcher . a postcolonial approach to computational methods for supporting language vitality is needed, says the researcher - lil'watul Lorna Williams . |
| Approach: | They propose to examine colonising discourses in speech and language technology and propose a postcolonial approach to computational methods for supporting language vitality. |
| Outcome: | The paper reviews colonising discourses in speech and language technology and suggests new ways of working with Indigenous communities. |