Papers by Mark Liberman
Reflections on 30 Years of Language Resource Development and Sharing (2022.lrec-1)
Copied to clipboard
| Challenge: | Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development. |
| Approach: | They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs. |
| Outcome: | The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans. |
From ‘Solved Problems’ to New Challenges: A Report on LDC Activities (L18-1)
Copied to clipboard
Christopher Cieri, Mark Liberman, Stephanie Strassel, Denise DiPersio, Jonathan Wright, Andrea Mazzucchi
| Challenge: | This paper reports on the activities of the Linguistic Data Consortium . |
| Approach: | This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report . |
| Outcome: | The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world . |
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)
Copied to clipboard
Christopher Cieri, James Fiumara, Stephanie Strassel, Jonathan Wright, Denise DiPersio, Mark Liberman
| Challenge: | Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources. |
| Approach: | a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation. |
| Outcome: | 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation. |
Introducing NIEUW: Novel Incentives and Workflows for Eliciting Linguistic Data (L18-1)
Copied to clipboard
| Challenge: | a 2010 survey found that the language of the European Union, not even English, was not fully supplied . the absence of Language Resources stifles teaching and technology building, authors say . |
| Approach: | They propose to harness the power of alternative incentives to elicit linguistic data and annotation . they also describe changes to the workflows necessary to collect data from workforces attracted by incentives . |
| Outcome: | a new initiative to harness incentives to elicit linguistic data and annotation improves language resources . the NIEUW project is funded by the u.s. national science foundation . |