Papers by Steven Moran

10 papers
BDPROTO: A Database of Phonological Inventories from Ancient and Reconstructed Languages (L18-1)

Copied to clipboard

Challenge: BDPROTO is a database of phonological inventory data from 137 ancient and reconstructed languages.
Approach: They present a database of phonological inventory data from 137 ancient and reconstructed languages.
Outcome: The BDPROTO database is a publicly available, unicode-compliant resource . it contains phonological inventory data from 137 ancient and reconstructed languages .
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)

Copied to clipboard

Challenge: ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language.
Approach: They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language .
Outcome: The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages .
Towards faithfully visualizing global linguistic diversity (L18-1)

Copied to clipboard

Challenge: Currently, visualizations of worldwide linguistic diversity are limited by point symbology .
Approach: They propose a method to visualize linguistic diversity using point symbology instead of Mercator . instead of languages-as-points, they use Voronoi/Thiessen tessellations to model linguistic areas .
Outcome: The proposed method is based on an Eckert IV projection instead of Mercator . instead of languages-as-points, it uses Voronoi/Thiessen tessellations to model linguistic areas .
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)

Copied to clipboard

Challenge: a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research .
Approach: a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary .
Outcome: a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web .
A Measure for Transparent Comparison of Linguistic Diversity in Multilingual NLP Data Sets (2024.findings-naacl)

Copied to clipboard

Challenge: a new study aims to assess linguistic diversity of multilingual data sets against a reference language sample . linguistic diversity is typically measured as the number of languages included in the data set . but such measures do not consider structural properties of the included languages .
Approach: They propose to measure linguistic diversity against a reference language sample to maximise linguistic diversity.
Outcome: The proposed measure can be used to identify the types of languages that are not represented in a data set.
Cross-linguistically Small World Networks are Ubiquitous in Child-directed Speech (L18-1)

Copied to clipboard

Challenge: In this paper we use network theory to model graphs of child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Approach: They use network theory to model child-directed speech from caregivers of children from nine typologically and morphologically diverse languages.
Outcome: The proposed model adds to the repertoire of universal distributional patterns found in the input to children cross-linguistically.
Phonetic Segmentation of the UCLA Phonetics Lab Archive (2024.lrec-main)

Copied to clipboard

Challenge: ''big data'' does not exist for the majority of the world's languages . a corpus of audited phonetic transcriptions and phone-level alignments is available for free .
Approach: They present a corpus of audited phonetic transcriptions and phone-level alignments from the UCLA Phonetics Lab Archive . they discuss the utility of the corpus for general research and pedagogy in crosslinguistic phonetics .
Outcome: The VoxAngeles corpus improves the original corpus for phonetic typology and word- and phone duration measurements.
Is Word Segmentation Child’s Play in All Languages? (P19-1)

Copied to clipboard

Challenge: Existing word learning strategies for infants are cross-linguistically robust . infants do not know which language(s) will be found in their environment at the beginning of development .
Approach: They propose to use 11 conceptually diverse algorithms to learn word-like units in infants . they propose to employ cross-linguistically robust algorithms that can be used by all infants.
Outcome: The proposed algorithms perform above chance on 8 different languages . the results show that some of the algorithms are cross-linguistically valid .
TeDDi Sample: Text Data Diversity Sample for Language Comparison and Multilingual NLP (2022.lrec-1)

Copied to clipboard

Challenge: a deep understanding of language is being achieved through increased access to data from minority and low-resource languages.
Approach: They present a diversity sample of text data for language comparison and multilingual natural language processing.
Outcome: The TeDDi sample features 89 languages based on the typological diversity sample in the World Atlas of Language Structures .
SegBo: A Database of Borrowed Sounds in the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: phonological segment borrowing is a process through which languages acquire new contrastive speech sounds as the result of borrowing words from other languages.
Approach: They propose to use a database to aggregate borrowed phonological segments from languages to create a new contrastive sound.
Outcome: The proposed database is based on a cross-linguistic database of borrowed phonological segments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations