Challenge: GINCO is a new training dataset for automatic genre identification based on 1,125 crawled Slovenian web documents that consist of 650,000 words.
Approach: They propose to use 1,125 crawled Slovenian web documents to train a new genre classification system based on a GINCO training dataset .
Outcome: The proposed classifiers perform better on the 1,125 crawled Slovenian web documents than the existing models and achieve higher scores on the task.

Similar Papers

Contribution of Move Structure to Automatic Genre Identification: An Annotated Corpus of French Tourism Websites (2024.lrec-main)

Copied to clipboard

Challenge: a concept of move structure has been overlooked in genre analysis, but it is not widely used in natural language processing.
Approach: They propose to incorporate move structure into a neural architecture for automatic genre identification.
Outcome: The proposed approach can increase performance and reduce computational power.
Language ID in the Wild: Unexpected Challenges on the Path to a Thousand-Language Web Text Corpus (2020.coling-main)

Copied to clipboard

Challenge: Large text corpora are increasingly important for a wide variety of NLP tasks.
Approach: They propose to train automatic language identification models on up to 1,629 languages . they find that human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages.
Outcome: The proposed models achieve over 90% average F1 on 1,629 languages . human-judged accuracy for web-crawl text corpora is only around 5% for many lower-resource languages - suggesting a need for more robust evaluation.
Huge Automatically Extracted Training-Sets for Multilingual Word SenseDisambiguation (L18-1)

Copied to clipboard

Challenge: Word Sense Disambiguation is a crucial task in Natural Language Processing . supervised systems need to be trained on word-by-word basis, a problem that is beyond reach for resource-rich languages like English.
Approach: They release six large-scale sense-annotated datasets in multiple languages to pave the way for supervised multilingual Word Sense Disambiguation.
Outcome: The results show that large-scale sense annotations can be used as training sets for supervised systems.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
MultiNERD: A Multilingual, Multi-Genre and Fine-Grained Dataset for Named Entity Recognition (and Disambiguation) (2022.findings-naacl)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a process of identifying named entities in unstructured texts and classifying them through specific semantic categories.
Approach: They propose a method for automatically producing NER annotations and introduce a manually-annotated test set.
Outcome: The proposed method covers 10 languages, 15 NER categories and 2 textual genres and a manually-annotated test set.
Estimating Confidence of Predictions of Individual Classifiers and TheirEnsembles for the Genre Classification Task (2022.lrec-1)

Copied to clipboard

Challenge: Genre identification is a kind of non-topic text classification. genre is defined as a functional space.
Approach: They propose to use SOTA to identify genres in non-topic texts . genres are functional and cannot be expressed just by some keywords .
Outcome: The proposed models show that they perform better than their individual models in large datasets.
Data, Data Everywhere: A Guide for Pretraining Dataset Construction (2024.emnlp-main)

Copied to clipboard

Challenge: Recent language models have impressive capabilities on a number of evaluation areas.
Approach: They conduct systematic analysis of pretraining set construction to identify which methods yield the greatest gains in model accuracy.
Outcome: The proposed method can be used to refine and improve a pretraining set.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data (2020.lrec-1)

Copied to clipboard

Challenge: Pre-training text representations have led to significant improvements in many areas of natural language processing.
Approach: They propose a pipeline to extract monolingual datasets from Common Crawl . pipeline follows data processing introduced in fastText that deduplicates documents .
Outcome: The proposed pipeline performs standard document deduplication and language identification similar to the pipeline introduced in fastText and a filtering step to select documents close to high quality corpora like Wikipedia.
DocHieNet: A Large and Diverse Dataset for Document Hierarchy Parsing (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for document hierarchy parsing are limited due to the small scale and inconsistency of datasets.
Approach: They propose a document hierarchy parsing dataset to compensate for the data scarcity problem and propose 'dHP' framework to grasp fine-grained text content and coarse-grounded pattern at layout element level.
Outcome: The proposed framework grasps both fine-grained text content and coarse-grounded pattern at layout element level, enhancing the capacity of pre-trained text-layout models in handling multi-page and multi-level challenges.
Writing System and Speaker Metadata for 2,800+ Language Varieties (2022.lrec-1)

Copied to clipboard

Challenge: Currently, language technologies are easily available in only a small minority of the world's 7,000+ language varieties.
Approach: They propose to use an open-source dataset to provide the writing system(s) for each of the 2,800+ languages used in the world today and an estimated speaker count for each.
Outcome: The dataset provides the attested writing system(s) for each of these 2,800+ varieties, as well as an estimated speaker count for each variety.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations