KOTONOHA: A Corpus Concordance System for Skewer-Searching NINJAL Corpora (2020.lrec-1)
Copied to clipboard
Teruaki Oka, Yuichi Ishimoto, Yutaka Yagi, Takenori Nakamura, Masayuki Asahara, Kikuo Maekawa, Toshinobu Ogiso, Hanae Koiso, Kumiko Sakoda, Nobuko Kibe
| Challenge: | NINJAL has developed several types of corpora for linguistic research . for each corpus NINJAL provided an online search environment, ‘Chunagon’ . |
| Approach: | NINJAL has developed several types of corpora for linguistic research . for each corpus NINJAL provided an online search environment, ‘Chunagon’, which is a morphological-information-annotation-based concordance system made publicly available in 2011 . NINjal has now provided a system ‘Kotonoha’ based on the ‘Chunegon’ systems . |
| Outcome: | NINJAL has provided a skewer-search system ‘Kotonoha’ based on ‘Chunagon’ systems. |
Similar Papers
JParaCrawl: A Large Scale Web-Based English-Japanese Parallel Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Recent machine translation algorithms rely on parallel corpora, but only some resource-rich language pairs can benefit from them. |
| Approach: | They construct a parallel corpus for English-Japanese, which has 8.7 million sentence pairs . they use a web crawler to automatically align parallel sentences in the corpus . |
| Outcome: | The proposed corpus includes a broader range of domains and can be trained with a pre-trained model. |
JParaCrawl v3.0: A Large-scale English-Japanese Parallel Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing parallel corpora for English-Japanese are limited, limiting the accuracy of machine translation models. |
| Approach: | They propose a web-based English-Japanese parallel corpus with 21 million unique sentence pairs . this is more than twice as many as the previous corpus JParaCrawl v2.0 . |
| Outcome: | The proposed corpus boosts the accuracy of machine translation models on various domains. |
To Boldly Query What No One Has Annotated Before? The Frontiers of Corpus Querying (2020.acl-main)
Copied to clipboard
| Challenge: | a systematic review of corpora and query tools focuses on the query side . annotated corporata are the backbone of many fields in linguistics . |
| Approach: | They propose a chronology of the major interplay between corpus progression and query tool evolution . they focus on the query side and hints at exciting directions for future development . |
| Outcome: | This paper provides a broad overview of the history of corpora and query tools . it focuses on the query side and hints at exciting directions for future development . |
Creating dialect sub-corpora by clustering: a case in Japanese for an adaptive method (L18-1)
Copied to clipboard
| Challenge: | a mixed corpus composed of different dialects is sufficiently resourced to cluster them into dialects. |
| Approach: | They propose a pipeline to derive clusters of dialects from a mixed corpus when their standard counterpart is sufficiently resourced. |
| Outcome: | The proposed pipeline can identify dialectal content when its standard counterpart is sufficiently resourced and can then cluster it into four dialects. |
ParaCrawl: Web-Scale Acquisition of Parallel Corpora (2020.acl-main)
Copied to clipboard
Marta Bañón, Pinzhen Chen, Barry Haddow, Kenneth Heafield, Hieu Hoang, Miquel Esplà-Gomis, Mikel L. Forcada, Amir Kamran, Faheem Kirefu, Philipp Koehn, Sergio Ortiz Rojas, Leopoldo Pla Sempere, Gema Ramírez-Sánchez, Elsa Sarrías, Marek Strelec, Brian Thompson, William Waites, Dion Wiggins, Jaume Zaragoza
| Challenge: | We describe methods to create the largest publicly available parallel corpora by crawling the web . parallel corpus is essential for building highquality machine translation systems . |
| Approach: | They describe methods to create largest publicly available parallel corpora by crawling web sites . they empirically compare alternative methods and publish benchmark data sets . |
| Outcome: | The proposed methods improve state-of-the-art results on common benchmarks, the authors show . the pipeline has been tested on Russian, Sinhala, Nepali, Tagalog, Swahili, and Somali . |
Facilitating Corpus Usage: Making Icelandic Corpora More Accessible for Researchers and Language Users (2020.lrec-1)
Copied to clipboard
| Challenge: | Gigaword corpus is a large text corpus used in natural language processing . large corpora are needed to achieve better performance in the field of NLP . |
| Approach: | They propose a set of tools to facilitate the use of the Icelandic Gigaword Corpus . they provide n-grams based on the corpus, and a variety of pre-trained word embeddings models . |
| Outcome: | The proposed tools facilitate the use of the Icelandic Gigaword corpus in the field of Natural Language Processing and other fields. |
Wikinflection Corpus: A (Better) Multilingual, Morpheme-Annotated Inflectional Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Inflectional corpora with annotated morpheme boundaries are scarce in the NLP community . a generated, multilingual inflectional lexicon with morphological features is not as good as UniMorph's . |
| Approach: | They evaluate a multilingual inflectional corpus with morpheme boundaries from the English Wiktionary and the UniMorph project's inflection corpus. |
| Outcome: | The generated Wikinflection corpus is not as good as UniMorph's, but extracts significant amount of words from the intersection of the two corpora. |
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)
Copied to clipboard
| Challenge: | a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks . |
| Approach: | They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs. |
| Outcome: | The proposed method is the most accurate and leads to lesser performance in downstream tasks. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
The German Reference Corpus DeReKo: New Developments – New Opportunities (L18-1)
Copied to clipboard
| Challenge: | DeReKo contains 42 billion tokens, comprising a multitude of genres such as newspaper text, fiction, or specialised text. |
| Approach: | They discuss legal issues around the recent German copyright reform and recent corpus extensions in popular magazines, journals, historical texts, and web-based football reports. |
| Outcome: | The German Reference Corpus DeReKo contains more than 42 billion tokens and is growing at 3.1 billion word per year. |