Papers by Lora Aroyo

7 papers
GRASP: A Disagreement Analysis Framework to Assess Group Associations in Perspectives (2024.naacl-long)

Copied to clipboard

Challenge: Recent work shows that ignoring rater subjectivity is problematic within specific tasks and for specific subgroups.
Approach: They propose a disagreement analysis framework to measure group association in perspectives among different rater subgroups.
Outcome: The proposed framework reveals specific rater groups that have significantly different perspectives than others on certain tasks and helps identify demographic axes that are crucial to consider in specific task contexts.
Scoring and Classifying Implicit Positive Interpretations: A Challenge of Class Imbalance (C18-1)

Copied to clipboard

Challenge: a reimplementation of a system on detecting implicit positive meaning from negated statements is reported . a baseline taking the mean score or most frequent class is hard to beat because of class imbalance in the dataset.
Approach: They propose a system to detect implicit positive meaning from negated statements . they convert the scores into classes and report their results on regression and classification tasks .
Outcome: The proposed system is hard to beat because of class imbalance in the dataset.
Follow the leader(board) with confidence: Estimating p-values from a single test set with item and response variance (2023.findings-acl)

Copied to clipboard

Challenge: Among the problems with leaderboard culture in NLP has been the widespread lack of confidence estimation in reported results.
Approach: They propose a framework and simulator for estimating p-values for comparisons between the results of two systems using variance found naturally (though rarely reported) in test set items and individual labels on an item (responses).
Outcome: The proposed framework and simulator are used to estimate p-values for comparisons between the results of two systems under the assumption that the null hypothesis is true.
AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications (2023.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are rapidly becoming more and more popular, but dealing with the potential harms associated with their deployment in real-world scenarios is still an open research question.
Approach: They propose an automated approach for automated generation of adversarial evaluation datasets to test the safety of LLM generations on new downstream applications.
Outcome: AART generates evaluation datasets with high diversity of content characteristics critical for effective adversarial testing.
Cross-replication Reliability - An Empirical Approach to Interpreting Inter-rater Reliability (2021.acl-long)

Copied to clipboard

Challenge: Respectable journals typically require reporting quantitative evidence for inter-rater reliability (IRR) of the data.
Approach: They propose to benchmark IRR against baseline measures in a replication dataset and use Cohen's (1960) kappa to measure inter-rater reliability.
Outcome: The proposed framework can be used to measure the quality of crowdsourced datasets.
A Crowdsourced Frame Disambiguation Corpus with Ambiguity (N19-1)

Copied to clipboard

Challenge: Using crowdsourcing, we have found that inter-annotator disagreement is at least partly caused by ambiguity inherent to the text and frames.
Approach: They propose a crowdsourcing approach to capture inter-annotator disagreement by a list of frames with disagreement-based scores that express the confidence with which each frame applies to the word.
Outcome: The proposed approach captures disagreement between the annotations of 1,000 word-sentence pairs and scores on the likelihood that each frame applies to the word.
Resource Interoperability for Sustainable Benchmarking: The Case of Events (L18-1)

Copied to clipboard

Challenge: Despite efforts to improve interoperability, there are still problems with benchmark corpora that are hampered by too laborious conversion steps.
Approach: They assess aspects of interoperability at the document-level across 20 annotated corpora and compare their compatibility and consistency across the corpors.
Outcome: The proposed framework enables the analysis of document intersections between the corpora and shows their compatibility and consistency across the corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations