FRACAS: a FRench Annotated Corpus of Attribution relations in newS (2024.lrec-main)
Copied to clipboard
| Challenge: | Quotation extraction is a useful task, but it is not widely studied in other languages. |
| Approach: | They propose to annotate a manually annotated corpus of 1,676 newswire texts in French for quotation extraction and source attribution. |
| Outcome: | The proposed system is compared to the most recent system for quotation extraction in the French language. |
Similar Papers
DirectQuote: A Dataset for Direct Quotation Extraction and Attribution in News Articles (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing methods to extract and attribute quotations from news data are difficult and require a lot of effort. |
| Approach: | They propose a corpus of 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. |
| Outcome: | The proposed corpus contains 19,760 paragraphs and 10,279 direct quotations manually annotated from online news media. |
A French Corpus and Annotation Schema for Named Entity Recognition and Relation Extraction of Financial News (2020.lrec-1)
Copied to clipboard
| Challenge: | Strict regulatory regimes mandate financial institutions to rigorously monitor their customers' financial activities. |
| Approach: | They propose to use an ontology of compliance-related concepts and relationships along with a corpus annotated according to it to train and evaluate named entity recognition algorithms. |
| Outcome: | The proposed ontology allows for training and evaluating domain-specific named entity recognition and relation extraction algorithms. |
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)
Copied to clipboard
| Challenge: | In literature, spoken interactions between characters are of central importance to the narrative. |
| Approach: | They propose to annotate quotations, including their interpersonal structure, for English literary text. |
| Outcome: | The proposed dataset provides a rich view of dialogue structures not available from other available corpora. |
An Attribution Relations Corpus for Political News (L18-1)
Copied to clipboard
| Challenge: | Existing resources for recognizing attributions in context are limited in size and completeness. |
| Approach: | They propose to use the largest and most complete attribution relations corpus to date . they propose to create sophisticated end-to-end solutions for attribution extraction . |
| Outcome: | The political news attribution relations corpus 2016 is the largest and most complete attribution relations corpuse to date. |
Dataset of Quotation Attribution in German News Articles (2024.lrec-main)
Copied to clipboard
| Challenge: | Lack of annotated data for quotation attribution in news articles severely limits the quality and usability of possible systems. |
| Approach: | They propose a dataset for quotation attribution in German news articles using WIKINEWS and manually annotated quotes from 1000 articles. |
| Outcome: | The proposed dataset provides curated, high-quality annotations across 1000 documents (250,000 tokens) in a fine-grained annotation schema enabling various downstream uses for the dataset. |
CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation Extraction (2022.coling-1)
Copied to clipboard
| Challenge: | Existing solutions for quotation extraction use rule-based approaches and sequence labeling models. |
| Approach: | They propose a Context and Former-Label Enhanced Net for quotation extraction. |
| Outcome: | The proposed method achieves state-of-the-art performance on complicated quotation extraction on two public datasets and one proprietary dataset. |
An Environment for Relational Annotation of Political Debates (P19-3)
Copied to clipboard
| Challenge: | Scalable text analysis techniques can open corpora to new questions in computational social sciences and digital humanities. |
| Approach: | They describe a tool that allows annotating newspaper text with rich information about claims (demands) raised by politicians and other actors. |
| Outcome: | The MARDY tool realizes the complete workflow necessary for annotating a large newspaper text collection with rich information about claims (demands) raised by politicians and other actors. |
CrudeOilNews: An Annotated Crude Oil News Corpus for Event Extraction (2022.lrec-1)
Copied to clipboard
| Challenge: | a corpus of English crude oil news for event extraction is presented . the corpus contains 425 news articles with approximately 11k events annotated . |
| Approach: | They present a corpus of English Crude Oil news for event extraction . it is the first of its kind for Commodity News and contributes to text mining . |
| Outcome: | The proposed corpus of English crude oil news is the first of its kind for Commodity News . the annotated news articles are compared with the standard news articles . |
Establishing a New State-of-the-Art for French Named Entity Recognition (2020.lrec-1)
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a task consisting in identifying text spans that denote named entities such as person, location and organization names. |
| Approach: | They manually annotated the French TreeBank with information related to named entities . they sketch the underlying annotation guidelines and provide a few figures about the annotations . |
| Outcome: | The French TreeBank is the main source of morphosyntactic and syntactical annotations for French. |
TIMELINE: Exhaustive Annotation of Temporal Relations Supporting the Automatic Ordering of Events in News Articles (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing temporal relation extraction models have low inter-annotator agreement due to lack of specificity of annotation guidelines . authors propose a method for annotating all temporal relations, including long-distance ones, which automates the process . |
| Approach: | They propose a new annotation scheme that defines criteria for temporal relations to be annotated . scheme includes events even if they are not expressed as verbs, they argue . |
| Outcome: | The proposed method reduces time and manual effort on the part of annotators. |