The GINCO Training Dataset for Web Genre Identification of Documents Out in the Wild (2022.lrec-1)
Copied to clipboard
| Challenge: | GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words. |
| Approach: | They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset . |
| Outcome: | The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task. |
Similar Papers
Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism Websites (2024.lrec-main)
Copied to clipboard
| Challenge: | a concept of move structure has been overlooked in genre analysis, but it is not widely used in natural language processing. |
| Approach: | They propose to incorporate move structure into a neural architecture for automatic genre identification. |
| Outcome: | The proposed approach can increase performance and reduce computational power. |
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)
Copied to clipboard
| Challenge: | Large text corpora are increasingly important for a wide variety of NLP tasks. |
| Approach: | They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages. |
| Outcome: | The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation. |
Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation (L18-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation is a crucial task in Natural Language Processing . supervised systems need to be trained on word-by-word basis, a problem that is beyond reach for resource-rich languages like English. |
| Approach: | They release six large-scale sense-annotated datasets in multiple languages to pave the way for supervised multilingual Word Sense Disambiguation. |
| Outcome: | The results show that large-scale sense annotations can be used as training sets for supervised systems. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation) (2022.findings-naacl)
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a process of identifying named entities in unstructured texts and classifying them through specific semantic categories. |
| Approach: | They propose a method for automatically producing NER annotations and introduce a manually-annotated test set. |
| Outcome: | The proposed method covers 10 languages, 15 NER categories and 2 textual genres and a manually-annotated test set. |
Estimating Confidence of Predictions of Individual Classifiers and TheirEnsembles for the Genre Classification Task (2022.lrec-1)
Copied to clipboard
| Challenge: | Genre identification is a kind of non-topic text classification. genre is defined as a functional space. |
| Approach: | They propose to use SOTA to identify genres in non-topic texts . genres are functional and cannot be expressed just by some keywords . |
| Outcome: | The proposed models show that they perform better than their individual models in large datasets. |
Data, Data Everywhere: A Guide for Pretraining Dataset Construction (2024.emnlp-main)
Copied to clipboard
Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings, Bo Liu, Aastha Jhunjhunwala, Zhilin Wang, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro
| Challenge: | Recent language models have impressive capabilities on a number of evaluation areas. |
| Approach: | They conduct systematic analysis of pretraining set construction to identify which methods yield the greatest gains in model accuracy. |
| Outcome: | The proposed method can be used to refine and improve a pretraining set. |
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data (2020.lrec-1)
Copied to clipboard
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, Edouard Grave
| Challenge: | Pre-training text representations have led to significant improvements in many areas of natural language processing. |
| Approach: | They propose a pipeline to extract monolingual datasets from Common Crawl . pipeline follows data processing introduced in fastText that deduplicates documents . |
| Outcome: | The proposed pipeline performs standard document deduplication and language identification similar to the pipeline introduced in fastText and a filtering step to select documents close to high quality corpora like Wikipedia. |
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets. |
| Approach: | They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level. |
| Outcome: | The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges. |
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)
Copied to clipboard
| Challenge: | Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties. |
| Approach: | They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each. |
| Outcome: | The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety. |