Jennifer Tracey, Owen Rambow, Claire Cardie, Adam Dalton, Hoa Trang Dang, Mona Diab, Bonnie Dorr, Louise Guthrie, Magdalena Markowska, Smaranda Muresan, Vinodkumar Prabhakaran, Samira Shaikh, Tomek Strzalkowski
| Challenge: | a corpus of propositional content is a set of cognitive attitudes of different agents towards a text . propositional attitudes are a cognitive attitude, including belief and sentiment, towards . |
| Approach: | They propose a corpus which records cognitive state: who believes what, who has what sentiment . they use newswire and discussion forums in Chinese, English, and Spanish . |
| Outcome: | The proposed corpus records who believes what (i.e., factuality) and who has what sentiment towards what. |
Similar Papers
Opinion Mining Using Pre-Trained Large Language Models: Identifying the Type, Polarity, Intensity, Expression, and Source of Private States (2024.lrec-main)
Copied to clipboard
Saeed Ahmadnia, Arash Yousefi Jordehi, Mahsa Hosseini Khasheh Heyran, SeyedAbolghasem Mirroshandel, Owen Rambow
| Challenge: | Existing research on opinion mining has focused on a small subset of the MPQA 2.0 dataset . a recent study focused on the subjective expressions of people who express opinions, sentiments, and attitudes toward targets. |
| Approach: | They propose to use MPQA 2.0 to analyze the entire dataset . they propose to provide a clean version of the MPQA Opinion Corpus in a more interpretable format . |
| Outcome: | The proposed methods establish high baselines for future work. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Large Scale Speech Sentiment Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpus for sentiment analysis uses text inputs, but voice inputs are becoming more important as smart assistants and mobile voice control become more prevalent. |
| Approach: | They propose to extend the Switchboard-1 Telephone Speech Corpus by adding sentiment labels from 3 different human annotators for every transcript segment. |
| Outcome: | The proposed corpus contains 49500 labeled speech segments covering 140 hours of audio. |
The ACQDIV Corpus Database and Aggregation Pipeline (2020.lrec-1)
Copied to clipboard
| Challenge: | ACQDIV corpus database and aggregation pipeline aims to identify universal cognitive processes that allow children to acquire any language. |
| Approach: | They present the ACQDIV corpus database and aggregation pipeline . the tool aims to identify universal cognitive processes that allow children to acquire any language . |
| Outcome: | The ACQDIV corpus database and aggregation pipeline is a tool developed by the European Research Council . the database represents 15 corpora from 14 typologically maximally diverse languages . |
ELQA: A Corpus of Metalinguistic Questions and Answers about English (2023.acl-long)
Copied to clipboard
| Challenge: | ELQA corpus is metalinguistic—it consists of language about language. |
| Approach: | They present a corpus of questions and answers in and about the English language . they use a free-form question answering task and multiple LLMs to analyze their capacity . |
| Outcome: | The ELQA corpus covers grammar, meaning, fluency, and etymology . the results can be used to investigate metalinguistic capabilities of NLU models . |
NoReC: The Norwegian Review Corpus (L18-1)
Copied to clipboard
Erik Velldal, Lilja Øvrelid, Eivind Alexander Bergem, Cathrine Stadsnes, Samia Touileb, Fredrik Jørgensen
| Challenge: | The Norwegian Review Corpus is a dataset of full-text reviews from major news sources. |
| Approach: | This paper presents the Norwegian Review Corpus, created for document-level sentiment analysis. |
| Outcome: | The corpus comprises more than 35,000 full-text reviews from a range of different domains. |
Medical Entity Corpus with PICO elements and Sentiment Analysis (L18-1)
Copied to clipboard
| Challenge: | In this paper, we establish a PICO and a sentiment annotated corpus of clinical trial publications. |
| Approach: | They propose to create a phrase-level PICO corpus and a sentence-level sentiment annotated corpus from clinical trial publications. |
| Outcome: | The proposed corpus is annotated on a phrase-level and a sentiment annotation on the same corpus. |
Sense and Sentiment (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing sentiment lexicons and concept-based sentiment-tagged corpora are not accurate, and it is difficult to map sentiment scores accurately to different languages. |
| Approach: | They examine existing sentiment lexicons and sense-based sentiment-tagged corpora to find out how sense and concept-based semantic relations effect sentiment scores. |
| Outcome: | The proposed lexicon can be used to generate sentiment lexicos for English using the Open Multilingual Wordnet. |
Beyond Metadata: What Paper Authors Say About Corpora They Use (2021.findings-acl)
Copied to clipboard
| Challenge: | Currently, dataset retrieval relies almost exclusively on metadata provided by the publishers. |
| Approach: | They propose to use metadata to extract review statements from scientific publications . they argue that a crucial piece of information is missing to inform the examination of search results . |
| Outcome: | The proposed analysis is the first of its kind in the field of Natural Language Processing. |
Quality Does Matter: A Detailed Look at the Quality and Utility of Web-Mined Parallel Corpora (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing web-mined corpora for low-resource languages have serious quality issues, especially for lowresource language pairs. |
| Approach: | They ranked each corpus according to a similarity measure and evaluated different portions of this ranked corpus. |
| Outcome: | The results show that the quality of web-mined corpora for low-resource languages is significantly different from human-curated corporats. |