Papers by Robert Forkel

3 papers
Converting Legacy Data to CLDF: A FAIR Exit Strategy for Linguistic Web Apps (2024.lrec-main)

Copied to clipboard

Challenge: a number of web applications that enabled comparative linguistics research became obsolete . cross-linguistic data formats (CLDF) are available for use in linguistic research .
Approach: a new standard allows researchers to convert legacy linguistic web apps into FAIR data . the standard uses W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary .
Outcome: a new standard can be used to convert legacy linguistic web apps into FAIR datasets . the standard is built on the W3C recommendations Model for Tabular Data and Metadata on the Web and MetaData Vocabulary for Tabulary on the web .
CLDFBench: Give Your Cross-Linguistic Data a Lift (2020.lrec-1)

Copied to clipboard

Challenge: despite the increasing amount of cross-linguistic data, most datasets are not FAIR (findable, accessible, interoperable, and reproducible) . with the Cross-Linguistic Data Formats initiative, first standards for cross-language data have been presented and successfully tested.
Approach: They propose a framework for the retro-standardization of legacy data and the curation of new datasets that drastically simplifies the creation of CLDFs.
Outcome: The proposed framework simplifies the creation of CLDFs by providing a consistent, reproducible workflow that supports version control and long term archiving of research data and code.
Linguistic Survey of India and Polyglotta Africana: Two Retrostandardized Digital Editions of Large Historical Collections of Multilingual Wordlists (2024.lrec-main)

Copied to clipboard

Challenge: Linguistic Survey of India and Polyglotta Africana are two of the largest historical collections of multilingual wordlists.
Approach: They present a retro-standardized edition of the Linguistic Survey of India and the Polyglotta Africana, which are two of the largest historical collections of multilingual wordlists.
Outcome: The LSI and PA are the largest historical collections of multilingual wordlists . but no editions in which the original data is presented in standardized form have been produced so far .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations