| Challenge: | Existing corpora on Samoyedic ( Uralic) languages include the INEL Kamas Corpus and the INL Selkup Corpus . |
| Approach: | They describe the Corpus Services framework, a collection of Java validation tools for language corpora compiled in XML-based data formats. |
| Outcome: | The proposed framework is integrated into the curation and publication workflows for EXMARaLDA-driven corpora of Northern Eurasian languages, as developed by the long-term project INEL . |
Similar Papers
Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models. |
| Approach: | They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters. |
| Outcome: | The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories . |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)
Copied to clipboard
| Challenge: | ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language. |
| Approach: | They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language . |
| Outcome: | The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages . |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
A Lightweight Modeling Middleware for Corpus Processing (L18-1)
Copied to clipboard
| Challenge: | Present-day empirical research in computational or theoretical linguistics has richly annotated and diverse corpus resources. |
| Approach: | They propose a framework for modeling arbitrary multi-modal corpus resources in a unified form for processing tools. |
| Outcome: | The proposed framework allows researchers to explore and query more diverse corpus resources and artifacts through a single interactive interface. |
An Empirical Evaluation of Annotation Practices in Corpora from Language Documentation (2020.lrec-1)
Copied to clipboard
| Challenge: | Language documentation projects have produced substantial amounts of primary data from a wide variety of endangered languages. |
| Approach: | They propose to use common annotation conventions in existing corpora to facilitate their future processing. |
| Outcome: | The proposed formats are based on the common ELAN and Toolbox formats and are used to facilitate their future processing. |
Introducing the CURLICAT Corpora: Seven-language Domain Specific Annotated Corpora from Curated Sources (2022.lrec-1)
Copied to clipboard
Tamás Váradi, Bence Nyéki, Svetla Koeva, Marko Tadić, Vanja Štefanec, Maciej Ogrodniczuk, Bartłomiej Nitoń, Piotr Pęzik, Verginica Barbu Mititelu, Elena Irimia, Maria Mitrofan, Dan Tufiș, Radovan Garabík, Simon Krek, Andraž Repar
| Challenge: | The CURLICAT CEF Telecom project aims to collect and deeply annotate a set of large corpora from selected domains. |
| Approach: | They present the results of the CURLICAT CEF Telecom project . they propose to collect and deeply annotate a set of large corpora from selected domains . |
| Outcome: | The CURLICAT CEF Telecom project provides a set of large corpora from selected domains . the corporatized corporates are tokenized, lemmatized and morphologically analysed . |
LexiDB: Patterns & Methods for Corpus Linguistic Database Management (2020.lrec-1)
Copied to clipboard
| Challenge: | LexiDB is a tool for storing, managing and querying corpus data. |
| Approach: | They propose to use LexiDB for storing, managing and querying corpus data. |
| Outcome: | The proposed methods outperform existing tools for corpus queries and storage. |
DECAF: A Dynamically Extensible Corpus Analysis Framework (2025.acl-demo)
Copied to clipboard
| Challenge: | DeCAF is an open-source Python library that enables the analysis and filtering of linguistically-annotated datasets down to the character level. |
| Approach: | They propose a framework that enables the analysis and filtering of linguistically-annotated datasets down to the character level. |
| Outcome: | The proposed framework analyzes a parsed version of the 115M-word BabyLM corpus and generates highly controlled and reproducible experimental settings targeting specific research questions. |
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)
Copied to clipboard
| Challenge: | a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks . |
| Approach: | They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs. |
| Outcome: | The proposed method is the most accurate and leads to lesser performance in downstream tasks. |