Papers with NCC
The Norwegian Colossal Corpus: A Text Corpus for Training Large Norwegian Language Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Norwegian is one of many languages lacking sufficient textual data to train quality language models. |
| Approach: | They propose to release 49GB of clean Norwegian textual data containing over 7B words . they hope to foster the creation of better Norwegian language models and multilingual language models . |
| Outcome: | The Norwegian Colossal Corpus (NCC) contains 49GB of clean Norwegian textual data containing over 7B words. |