Challenge: Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models.
Approach: They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters.
Outcome: The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories .

Similar Papers

Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)

Copied to clipboard

Challenge: a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks .
Approach: They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs.
Outcome: The proposed method is the most accurate and leads to lesser performance in downstream tasks.
A Repository of Corpora for Summarization (L18-1)

Copied to clipboard

Challenge: Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task.
Approach: They propose a repository containing corpora available to train and evaluate automatic summarization systems.
Outcome: The proposed system is based on a repository of corpora available for summarization tasks.
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)

Copied to clipboard

Challenge: Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs.
Approach: They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus.
Outcome: The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
A database of German definitory contexts from selected web sources (L18-1)

Copied to clipboard

Challenge: a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts.
Approach: They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction .
Outcome: The proposed method is based on a web corpus and a robust pattern-based extraction method.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit (2024.lrec-main)

Copied to clipboard

Challenge: Existing corpora are limited in scope and can be used to collect data on author attributes.
Approach: They propose to use subreddits, flairs, and self-reports as distant labels for author attributes (age, gender, nationality, personality, and political leaning) .
Outcome: The proposed method could be used to infer author attributes from public posts despite their discreetness and anonymity .
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)

Copied to clipboard

Challenge: Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday.
Approach: They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages.
Outcome: The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations