Challenge: NINJAL has developed several types of corpora for linguistic research . for each corpus NINJAL provided an online search environment, ‘Chunagon’ .
Approach: NINJAL has developed several types of corpora for linguistic research . for each corpus NINJAL provided an online search environment, ‘Chunagon’, which is a morphological-information-annotation-based concordance system made publicly available in 2011 . NINjal has now provided a system ‘Kotonoha’ based on the ‘Chunegon’ systems .
Outcome: NINJAL has provided a skewer-search system ‘Kotonoha’ based on ‘Chunagon’ systems.

Similar Papers

JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them.
Approach: They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus .
Outcome: The proposed corpus includes a broader range of domains and can be trained with a pre-trained model.
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)

Copied to clipboard

Challenge: Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models.
Approach: They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 .
Outcome: The proposed corpus boosts the accuracy of machine translation models on various domains.
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)

Copied to clipboard

Challenge: a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics .
Approach: They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development .
Outcome: This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development .
Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method (L18-1)

Copied to clipboard

Challenge: a mixed corpus composed of different dialects is sufficiently resourced to cluster them into dialects.
Approach: They propose a pipeline to derive clusters of dialects from a mixed corpus when their standard counterpart is sufficiently resourced.
Outcome: The proposed pipeline can identify dialectal content when its standard counterpart is sufficiently resourced and can then cluster it into four dialects.
ParaCrawl: Web-Scale Acquisition of Parallel Corpora (2020.acl-main)

Copied to clipboard

Challenge: We describe methods to create the largest publicly available parallel corpora by crawling the web . parallel corpus is essential for building highquality machine translation systems .
Approach: They describe methods to create largest publicly available parallel corpora by crawling web sites . they empirically compare alternative methods and publish benchmark data sets .
Outcome: The proposed methods improve state-of-the-art results on common benchmarks, the authors show . the pipeline has been tested on Russian, Sinhala, Nepali, Tagalog, Swahili, and Somali .
Facilitating Corpus Usage: Making Icelandic Corpora More Accessible for Researchers and Language Users (2020.lrec-1)

Copied to clipboard

Challenge: Gigaword corpus is a large text corpus used in natural language processing . large corpora are needed to achieve better performance in the field of NLP .
Approach: They propose a set of tools to facilitate the use of the Icelandic Gigaword Corpus . they provide n-grams based on the corpus, and a variety of pre-trained word embeddings models .
Outcome: The proposed tools facilitate the use of the Icelandic Gigaword corpus in the field of Natural Language Processing and other fields.
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's .
Approach: They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus.
Outcome: The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora.
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)

Copied to clipboard

Challenge: a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks .
Approach: They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs.
Outcome: The proposed method is the most accurate and leads to lesser performance in downstream tasks.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)

Copied to clipboard

Challenge: DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text.
Approach: They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports.
Outcome: The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations