Challenge: Quotation extraction is a useful task, but it is not widely studied in other languages.
Approach: They propose to annotate a manually annotated corpus of 1,676 newswire texts in French for quotation extraction and source attribution.
Outcome: The proposed system is compared to the most recent system for quotation extraction in the French language.

Similar Papers

DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to extract and attribute quotations from news data are difficult and require a lot of effort.
Approach: They propose a corpus of 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media.
Outcome: The proposed corpus contains 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media.
A French Corpus and Annotation Schema for Named Entity Recognition and Relation Extraction of Financial News (2020.lrec-1)

Copied to clipboard

Challenge: Strict regulatory regimes mandate financial institutions to rigorously monitor their customers' financial activities.
Approach: They propose to use an ontology of compliance-related concepts and relationships along with a corpus annotated according to it to train and evaluate named entity recognition algorithms.
Outcome: The proposed ontology allows for training and evaluating domain-specific named entity recognition and relation extraction algorithms.
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)

Copied to clipboard

Challenge: In literature, spoken interactions between characters are of central importance to the narrative.
Approach: They propose to annotate quotations, including their interpersonal structure, for English literary text.
Outcome: The proposed dataset provides a rich view of dialogue structures not available from other available corpora.
An Attribution Relations Corpus for Political News (L18-1)

Copied to clipboard

Challenge: Existing resources for recognizing attributions in context are limited in size and completeness.
Approach: They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction .
Outcome: The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date.
Dataset of Quotation Attribution in German News Articles (2024.lrec-main)

Copied to clipboard

Challenge: Lack of annotated data for quotation attribution in news articles severely limits the quality and usability of possible systems.
Approach: They propose a dataset for quotation attribution in German news articles using WIKINEWS and manually annotated quotes from 1000 articles.
Outcome: The proposed dataset provides curated, high-quality annotations across 1000 documents (250,000 tokens) in a fine-grained annotation schema enabling various downstream uses for the dataset.
CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction (2022.coling-1)

Copied to clipboard

Challenge: Existing solutions for quotation extraction use rule-based approaches and sequence labeling models.
Approach: They propose a Context and Former-Label Enhanced Net for quotation extraction.
Outcome: The proposed method achieves state-of-the-art performance on complicated quotation extraction on two public datasets and one proprietary dataset.
An Environment for Relational Annotation of Political Debates (P19-3)

Copied to clipboard

Challenge: Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities.
Approach: They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors.
Outcome: The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors.
CrudeOilNews: An Annotated Crude Oil News Corpus for Event Extraction (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of English crude oil news for event extraction is presented . the corpus contains 425 news articles with approximately 11k events annotated .
Approach: They present a corpus of English Crude Oil news for event extraction . it is the first of its kind for Commodity News and contributes to text mining .
Outcome: The proposed corpus of English crude oil news is the first of its kind for Commodity News . the annotated news articles are compared with the standard news articles .
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names.
Approach: They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations .
Outcome: The French TreeBank is the main source of morphosyntactic and syntactical annotations for French.
TIMELINE: Exhaustive Annotation of Temporal Relations Supporting the Automatic Ordering of Events in News Articles (2023.emnlp-main)

Copied to clipboard

Challenge: Existing temporal relation extraction models have low inter-annotator agreement due to lack of specificity of annotation guidelines . authors propose a method for annotating all temporal relations, including long-distance ones, which automates the process .
Approach: They propose a new annotation scheme that defines criteria for temporal relations to be annotated . scheme includes events even if they are not expressed as verbs, they argue .
Outcome: The proposed method reduces time and manual effort on the part of annotators.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations