Christopher Cieri, Mark Liberman, Stephanie Strassel, Denise DiPersio, Jonathan Wright, Andrea Mazzucchi
| Challenge: | This paper reports on the activities of the Linguistic Data Consortium . |
| Approach: | This paper reports on the activities of the Linguistic Data Consortium . it summarizes the over 100 Language Resources released since the last report . |
| Outcome: | The report summarizes the over 100 Language Resources released since the last report . many of the LRs have been contributed by research groups around the world . |
Similar Papers
A Progress Report on Activities at the Linguistic Data Consortium Benefitting the LREC Community (2020.lrec-1)
Copied to clipboard
Christopher Cieri, James Fiumara, Stephanie Strassel, Jonathan Wright, Denise DiPersio, Mark Liberman
| Challenge: | Linguistic Data Consortium (LDC) activities include the collection, annotation, processing, distribution, archiving and curation of language resources. |
| Approach: | a new report sketches the activities of a data center devoted to supporting the work of LREC attendees . 96 new corpora released in 2018-2020 to date, a technology evaluation campaign and innovations to advance methodology for language data collection and annotation. |
| Outcome: | 96 new corpora released in 2018-2020 to date, new technology evaluation campaign and innovations to advance methodology of language data collection and annotation. |
Reflections on 30 Years of Language Resource Development and Sharing (2022.lrec-1)
Copied to clipboard
| Challenge: | Linguistic Data Consortium was founded in 1992 to solve the problem that limitations in access to shareable data was impeding progress in Human Language Technology research and development. |
| Approach: | They review the roles of the Linguistic Data Consortium over the past 30 years after describing the conditions that lead to an HLT winter followed by a reawakening and an insatiable hunger for LRs. |
| Outcome: | The authors review the roles of the Linguistic Data Consortium over the past 30 years and provide a preview into future plans. |
The LREC Workshops Map (L18-1)
Copied to clipboard
| Challenge: | a corpus of workshops titles and related presentations has been retrieved from the conference's website . data is used to analyze the research presented at the conferences over the years 1998-2016 . |
| Approach: | a paper aims to present an overview of the research presented at the LREC workshops over the years 1998-2016. |
| Outcome: | The aim of the present study is to shed light on the community represented by workshop participants over the years 1998-2016. |
How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) are impressive in solving tasks, but they can quickly be outdated after deployment. |
| Approach: | They provide a review of recent advances in aligning deployed large language models with the ever-changing world knowledge. |
| Outcome: | The proposed models can be used to perform various tasks directly through in-context learning or for further fine-tuning for domain-specific uses. |
Related Works in the Linguistic Data Consortium Catalog (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing metadata standards for Related Works are used to define relations between language resources. |
| Approach: | They describe the development and implementation of a Related Works schema and the steps to implementation. |
| Outcome: | The proposed schema has been implemented in the Linguistic Data Consortium's (LDC) Catalog. |
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (L18-1)
Copied to clipboard
| Challenge: | null |
| Approach: | null |
| Outcome: | null |
Language Resources to Support Language Diversity – the ELRA Achievements (2022.lrec-1)
Copied to clipboard
| Challenge: | ELRA and its operational agency ELDA have continued to increase their catalogue of Language Resources (LRs) over the past few years, ELLA and ELTA have contributed to improve the access to multilingual information in the context of the pandemic . |
| Approach: | ELRA and its operational agency ELDA have increased their catalogue of Language Resources (LRs) over the past few years, they have established partnerships for the production of various types of LRs. |
| Outcome: | ELRA and its operational agency ELDA have contributed to improve the access to multilingual information in the context of the pandemic, develop tools for the de-identification of texts in the legal and medical domains, and support the EU eTranslation Machine Translation system. |
Understanding the Gap: an Analysis of Research Collaborations in NLP and Language Documentation (2025.findings-acl)
Copied to clipboard
| Challenge: | despite 20 years of NLP work, practical use of this work remains vanishingly scarce. |
| Approach: | They propose to use interviews and surveys to examine the lack of NLP adoption in LD . they find that linguists and language communities have little or no use of Nlp in their work . |
| Outcome: | a new study shows that linguists and language researchers are not using NLP in LD . the findings highlight the importance of misaligned professional incentives and LD software . |
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce (2025.acl-long)
Copied to clipboard
| Challenge: | Language is a powerful means of communication and should be regarded as more than just a collection of tokens. |
| Approach: | They collect feedback from individuals directly involved in and impacted by NLP artefacts for medium- and low-resource languages and highlight key issues related to data quality, cultural appropriateness and ethics of common annotation practices. |
| Outcome: | The findings highlight key issues related to data quality, cultural appropriateness, and ethics of common annotation practices. |
Some Issues with Building a Multilingual Wordnet (2020.lrec-1)
Copied to clipboard
| Challenge: | Notable extensions include: confidence, corpus frequency, orthographic variants, lexicalized and non-lexicalised synsets and lemmas, new parts of speech, and more. |
| Approach: | They propose to integrate a new open multilingual wordnet format that tests the extensions introduced by the new format and integrates a set of tools to ensure the integrity of the Collaborative Interlingual Index. |
| Outcome: | The proposed format integrates a set of tools that test the extensions while ensuring the integrity of the Collaborative Interlingual Index (CILI). |