Papers by Steven Moran
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)
Copied to clipboard
| Challenge: | BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Approach: | They present a database of phonological inventory data from 137 ancient and reconstructed languages. |
| Outcome: | The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages . |
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)
Copied to clipboard
| Challenge: | ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language. |
| Approach: | They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language . |
| Outcome: | The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages . |
Towards faithfully visualizing global linguistic diversity (L18-1)
Copied to clipboard
| Challenge: | Currently, visualizations of worldwide linguistic diversity are limited by point symbology . |
| Approach: | They propose a method to visualize linguistic diversity using point symbology instead of Mercator . instead of languages-as-points, they use Voronoi/Thiessen tessellations to model linguistic areas . |
| Outcome: | The proposed method is based on an Eckert IV projection instead of Mercator . instead of languages-as-points, it uses Voronoi/Thiessen tessellations to model linguistic areas . |
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)
Copied to clipboard
| Challenge: | a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research . |
| Approach: | a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary . |
| Outcome: | a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web . |
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)
Copied to clipboard
| Challenge: | a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages . |
| Approach: | They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity. |
| Outcome: | The proposed measure can be used to identify the types of languages that are not represented in a data set. |
Cross-linguistically Small World Networks are Ubiquitous in Child-directed Speech (L18-1)
Copied to clipboard
| Challenge: | In this paper we use network theory to model graphs of child-directed speech from caregivers of children from nine typologically and morphologically diverse languages. |
| Approach: | They use network theory to model child-directed speech from caregivers of children from nine typologically and morphologically diverse languages. |
| Outcome: | The proposed model adds to the repertoire of universal distributional patterns found in the input to children cross-linguistically. |
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)
Copied to clipboard
| Challenge: | ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free . |
| Approach: | They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics . |
| Outcome: | The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements. |
Is Word Segmentation Child’s Play in All Languages? (P19-1)
Copied to clipboard
| Challenge: | Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development . |
| Approach: | They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants. |
| Outcome: | The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid . |
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)
Copied to clipboard
| Challenge: | a deep understanding of language is being achieved through increased access to data from minority and low-resource languages. |
| Approach: | They present a diversity sample of text data for language comparison and multilingual natural language processing. |
| Outcome: | The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures . |
SegBo: A Database of Borrowed Sounds in the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing words from other languages. |
| Approach: | They propose to use a database to aggregate borrowed phonological segments from languages to create a new contrastive sound. |
| Outcome: | The proposed database is based on a cross-linguistic database of borrowed phonological segments. |