BeSt: The Belief and Sentiment Corpus (2022.lrec-1)

Copied to clipboard

Challenge: a corpus of propositional content is a set of cognitive attitudes of different agents towards a text . propositional attitudes are a cognitive attitude, including belief and sentiment, towards .
Approach: They propose a corpus which records cognitive state: who believes what, who has what sentiment . they use newswire and discussion forums in Chinese, English, and Spanish .
Outcome: The proposed corpus records who believes what (i.e., factuality) and who has what sentiment towards what.

Similar Papers

Opinion Mining Using Pre-Trained Large Language Models: Identifying the Type, Polarity, Intensity, Expression, and Source of Private States (2024.lrec-main)

Copied to clipboard

Challenge: Existing research on opinion mining has focused on a small subset of the MPQA 2.0 dataset . a recent study focused on the subjective expressions of people who express opinions, sentiments, and attitudes toward targets.
Approach: They propose to use MPQA 2.0 to analyze the entire dataset . they propose to provide a clean version of the MPQA Opinion Corpus in a more interpretable format .
Outcome: The proposed methods establish high baselines for future work.
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)

Copied to clipboard

Challenge: Until recently, language descriptions were available in paper form only, with indexes as the only search aid.
Approach: They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful.
Outcome: The proposed corpus is searchable through a couple of well-established corpus infrastructures.
A Large Scale Speech Sentiment Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpus for sentiment analysis uses text inputs, but voice inputs are becoming more important as smart assistants and mobile voice control become more prevalent.
Approach: They propose to extend the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment.
Outcome: The proposed corpus contains 49500 labeled speech segments covering 140 hours of audio.
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)

Copied to clipboard

Challenge: ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language.
Approach: They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language .
Outcome: The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages .
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)

Copied to clipboard

Challenge: ELQA corpus is metalinguistic—it consists of language about language.
Approach: They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity .
Outcome: The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models .
NoReC: The Norwegian Review Corpus (L18-1)

Copied to clipboard

Challenge: The Norwegian Review Corpus is a dataset of full-text reviews from major news sources.
Approach: This paper presents the Norwegian Review Corpus, created for document-level sentiment analysis.
Outcome: The corpus comprises more than 35,000 full-text reviews from a range of different domains.
Medical Entity Corpus with PICO elements and Sentiment Analysis (L18-1)

Copied to clipboard

Challenge: In this paper, we establish a PICO and a sentiment annotated corpus of clinical trial publications.
Approach: They propose to create a phrase-level PICO corpus and a sentence-level sentiment annotated corpus from clinical trial publications.
Outcome: The proposed corpus is annotated on a phrase-level and a sentiment annotation on the same corpus.
Sense and Sentiment (2022.lrec-1)

Copied to clipboard

Challenge: Existing sentiment lexicons and concept-based sentiment-tagged corpora are not accurate, and it is difficult to map sentiment scores accurately to different languages.
Approach: They examine existing sentiment lexicons and sense-based sentiment-tagged corpora to find out how sense and concept-based semantic relations effect sentiment scores.
Outcome: The proposed lexicon can be used to generate sentiment lexicos for English using the Open Multilingual Wordnet.
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)

Copied to clipboard

Challenge: Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers.
Approach: They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results .
Outcome: The proposed analysis is the first of its kind in the field of Natural Language Processing.
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)

Copied to clipboard

Challenge: Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs.
Approach: They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus.
Outcome: The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations