Challenge: a new corpus for detecting and linking survey variables is being developed . the corpus is multilingual and includes manually curated word and phrase alignments .
Approach: They propose to create a corpus for the evaluation of detecting and linking survey variables in social science publications.
Outcome: The proposed corpus is the first gold standard for the variable detection and linking task.

Similar Papers

A Multi-level Annotated Corpus of Scientific Papers for Scientific Document Summarization and Cross-document Relation Discovery (2020.lrec-1)

Copied to clipboard

Challenge: Recent studies have proposed to take advantage of the scientific paper's citation network to approach literature summarization.
Approach: They propose to annotate related work sections, cite papers and sentences using machine readable data and an additional layer of papers citing the references.
Outcome: The proposed corpus expands the existing data-set of related work sections and cites the papers cited in the related work section.
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions (2024.lrec-main)

Copied to clipboard

Challenge: SciDMT is an enhanced and expanded corpus for scientific mention detection . existing corpora are limited by their small volume and entity linking capabilities .
Approach: They propose to enhance SciDMT, an annotated scientific corpus for scientific mention detection.
Outcome: The proposed corpus is the largest for scientific entity mention detection . it is based on deep learning architectures like SciBERT and GPT-3.5 .
Albanian Part-of-Speech Tagging: Gold Standard and Evaluation (L18-1)

Copied to clipboard

Challenge: a corpus of more than 31,000 tokens is used for part-of-speech tagging in Albanian . a large number of multi-word units are difficult to tally, especially when they have articles or particles as their first part.
Approach: They propose a gold standard corpus for Albanian part-of-speech tagging and perform evaluation experiments with different statistical taggers.
Outcome: The proposed corpus can accurately represent the syntagmatic aspects of Albanian . the results show that the standard is accurate on both the full and coarse tagsets .
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Enriching a Lexicon of Discourse Connectives with Corpus-based Data (L18-1)

Copied to clipboard

Challenge: Existing annotation efforts for multiple languages have focused on discourse connectives, but we have limited it to the class of connectives marking contrast and the additional relations such connectives might convey.
Approach: They enrich a lexicon of italian COnnectives with real corpus data for connectives marking contrast relations in text.
Outcome: The proposed resource is a valuable tool for linguistic analyses of discourse relations and the training of a classifier for NLP applications.
Boosting Entity Linking Performance by Leveraging Unlabeled Documents (P19-1)

Copied to clipboard

Challenge: a new approach to entity linking relies on unlabeled documents and Wikipedia . a supervised approach uses only natural information, such as unlabed documents .
Approach: They propose a method which exploits only naturally occurring information . they construct a high recall list of candidate entities for each mention in an unlabeled document .
Outcome: The proposed model outperforms fully-supervised state-of-the-art systems on standard test sets.
A Short Survey on Sense-Annotated Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Word Sense Disambiguation (WSD) is a key task in Natural Language Understanding.
Approach: They propose to use sense-annotated corpora for supervised Word Sense Disambiguation.
Outcome: The proposed methods have been compared with knowledge-based approaches and have shown to be more efficient when they are available.
An Emotional Mess! Deciding on a Framework for Building a Dutch Emotion-Annotated Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing frameworks for emotion recognition are limited and do not allow for categorical versus dimensional oppositions.
Approach: They propose to use the emotions joy, love, anger, sadness and fear as well as dimensional models to annotate texts from different domains and topics.
Outcome: The proposed frameworks are well-suited to annotate texts from different domains and topics, but the connotation of the labels strongly depends on the origin of the texts.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations