Know thy Corpus! Robust Methods for Digital Curation of Web corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for estimating the lexicon of Web corpora have not been used to train pre-trained models. |
| Approach: | They propose a framework for digital curation of Web corpora to provide robust estimation of their parameters. |
| Outcome: | The proposed framework provides robust estimation of Web corpora's composition and lexicon . the proposed framework is similar to the BNC and ELMO models, but lacks curated categories . |
Similar Papers
Investigating Web Corpus Filtering Methods for Language Model Development in Japanese (2024.naacl-srw)
Copied to clipboard
| Challenge: | a high quality web corpus is essential for large language models to be developed . strong filtering methods can lead to lesser performance in downstream tasks . |
| Approach: | They build classifiers and language models that can process large amounts of corpora quickly enough for pretraining LLMs. |
| Outcome: | The proposed method is the most accurate and leads to lesser performance in downstream tasks. |
A Repository of Corpora for Summarization (L18-1)
Copied to clipboard
| Challenge: | Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task. |
| Approach: | They propose a repository containing corpora available to train and evaluate automatic summarization systems. |
| Outcome: | The proposed system is based on a repository of corpora available for summarization tasks. |
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs. |
| Approach: | They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus. |
| Outcome: | The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)
Copied to clipboard
| Challenge: | Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding. |
| Approach: | They propose to use sense-annotated corpora for supervised Word Sense Disambiguation. |
| Outcome: | The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available. |
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets (2022.tacl-1)
Copied to clipboard
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Ortiz Suarez, Iroro Orife, Kelechi Ogueji, Andre Niyongabo Rubungo, Toan Q. Nguyen, Mathias Müller, André Müller, Shamsuddeen Hassan Muhammad, Nanda Muhammad, Ayanda Mnyakeni, Jamshidbek Mirzakhalov, Tapiwanashe Matangira, Colin Leong, Nze Lawson, Sneha Kudugunta, Yacine Jernite, Mathias Jenny, Orhan Firat, Bonaventure F. P. Dossou, Sakhile Dlamini, Nisansa de Silva, Sakine Çabuk Ballı, Stella Biderman, Alessia Battisti, Ahmed Baruwa, Ankur Bapna, Pallavi Baljekar, Israel Abebe Azime, Ayodele Awokoya, Duygu Ataman, Orevaoghene Ahia, Oghenefego Ahia, Sweta Agrawal, Mofetoluwa Adeyemi
| Challenge: | Lower-resource corpora have systematic issues, including mislabeled or nonstandard/ambiguous language codes. |
| Approach: | They manually audit the quality of 205 language-specific corpora released with five major public datasets. |
| Outcome: | The results show that lower-resource corpora have systematic issues even for non-proficient speakers. |
A database of German definitory contexts from selected web sources (L18-1)
Copied to clipboard
| Challenge: | a specialized web corpus and robust pattern-based extraction methods are used to detect definitory contexts. |
| Approach: | They propose to use a web corpus and a database to detect definitory contexts . they describe an experimental setting and front-end for pattern-based definition extraction . |
| Outcome: | The proposed method is based on a web corpus and a robust pattern-based extraction method. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
SOBR: A Corpus for Stylometry, Obfuscation, and Bias on Reddit (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora are limited in scope and can be used to collect data on author attributes. |
| Approach: | They propose to use subreddits, flairs, and self-reports as distant labels for author attributes (age, gender, nationality, personality, and political leaning) . |
| Outcome: | The proposed method could be used to infer author attributes from public posts despite their discreetness and anonymity . |
Building and curating conversational corpora for diversity-aware language science and technology (2022.lrec-1)
Copied to clipboard
| Challenge: | Language resources that capture language use in its natural habitat of social interaction are rare despite the obvious merits of studying the very environment where we all learn and use it everyday. |
| Approach: | They propose to build an analysis pipeline and best practice guidelines for building and curating corpora of everyday conversation in diverse languages. |
| Outcome: | The proposed pipeline can be used to collect and curate conversational corpora in 67 languages and varieties from 28 phyla. |