Challenge: Sensitive information detection is of great importance in a number of applications where unintended leaks of sensitive information may incur severe negative consequences.
Approach: They propose to use a corpus of sentences to evaluate sensitive information detection approaches . they employ human annotations and automatically infer labels from domain experts .
Outcome: The proposed models are based on a monsanto trial and are evaluated on sentence level.

Similar Papers

AMR Beyond the Sentence: the Multi-sentence AMR corpus (C18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is limited to capturing the semantics of individual sentences.
Approach: They propose a corpus that annotates coreference and similar phenomena on top of existing AMRs.
Outcome: The proposed corpus is compared with existing corpora on sentence-level semantics . it shows that it can be used for information extraction and question answering .
Beyond Dataset Creation: Critical View of Annotation Variation and Bias Probing of a Dataset for Online Radical Content Detection (2025.coling-main)

Copied to clipboard

Challenge: Existing datasets and models fail to address the complexities of multilingual data, authors say . detection of radical content on online platforms has become an increasingly pressing concern .
Approach: They propose a publicly available multilingual dataset annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic.
Outcome: The proposed dataset is annotated with radicalization levels, calls for action, and named entities in English, French, and Arabic.
A Legal Perspective on Training Models for Natural Language Processing (L18-1)

Copied to clipboard

Challenge: a significant concern in processing natural language data is the unclear legal status of the input and output data/resources.
Approach: They examine which legal rules apply at relevant steps and how they affect the legal status of the results.
Outcome: The proposed model training process is based on three scenarios . the analysis focuses on which legal rules apply and how they affect the legal status of the results .
ProGene - A Large-scale, High-Quality Protein-Gene Annotated Benchmark Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Genes and proteins are fundamental entities of molecular genetics and are important for precision medicine.
Approach: They propose to use a corpus of gene and protein names to cope with this class of named entities in a large-scale annotation campaign at the Jena University Language & Information Engineering lab.
Outcome: The proposed corpus is an overall subdomain-independent corpus . it consists of 3,308 MEDLINE abstracts with over 36k sentences and more than 960k tokens annotated with nearly 60k named entity mentions.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
NEWTS: A Corpus for News Topic-Focused Summarization (2022.findings-acl)

Copied to clipboard

Challenge: Existing benchmarking corpora provide concordant pairs of full and abridged versions of Web, news or professional content.
Approach: They propose a topical summarization corpus called NEWTS that is annotated via crowd-sourcing.
Outcome: The proposed model can condition summaries on a desired range of themes . the proposed model outperforms Lead-3 baselines on most benchmark datasets .
No offence, Bert - I insult only humans! Multilingual sentence-level attack on toxicity detection networks (2023.findings-emnlp)

Copied to clipboard

Challenge: a new sentence-level attack on toxic detection models is shown to work on seven languages . toxicity detection systems are used to silence the voices of criticism, causing echo chambers .
Approach: They propose a sentence-level attack that adds positive words to a hateful message . they show the attack works on seven languages from three different language families .
Outcome: The proposed attack is shown to work on seven languages from three different language families.
Towards an argumentative content search engine using weak supervision (C18-1)

Copied to clipboard

Challenge: Existing work focused on detecting claims within a small set of documents . however, pinpointing relevant claims within massive unstructured corpora, received little attention.
Approach: They propose to use a weak signal to develop a query for claim–sentence detection using a large text corpus.
Outcome: The proposed system outperforms previous results in terms of precision and coverage.
Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations (D19-3)

Copied to clipboard

Challenge: Proceedings of the system demonstrations session were presented at the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) EMNMP-IjCNLP 2019 has a Best Demo Award for the first time .
Approach: Proceedings of the system demonstrations session are available online . they were presented at the conference on empirical methods in natural language processing .
Outcome: The system demonstrations session received 110 submissions, 22 of which were either invalid or withdrawn by the authors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations