Papers by Chris Welty

4 papers
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for estimating model reliability are based on a few output responses per item.
Approach: They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing.
Outcome: The proposed method can help researchers make better decisions about how to collect data for AI evaluation.
Follow the leader(board) with confidence: Estimating p-values from a single test set with item and response variance (2023.findings-acl)

Copied to clipboard

Challenge: Among the problems with leaderboard culture in NLP has been the widespread lack of confidence estimation in reported results.
Approach: They propose a framework and simulator for estimating p-values for comparisons between the results of two systems using variance found naturally (though rarely reported) in test set items and individual labels on an item (responses).
Outcome: The proposed framework and simulator are used to estimate p-values for comparisons between the results of two systems under the assumption that the null hypothesis is true.
A Crowdsourced Frame Disambiguation Corpus with Ambiguity (N19-1)

Copied to clipboard

Challenge: Using crowdsourcing, we have found that inter-annotator disagreement is at least partly caused by ambiguity inherent to the text and frames.
Approach: They propose a crowdsourcing approach to capture inter-annotator disagreement by a list of frames with disagreement-based scores that express the confidence with which each frame applies to the word.
Outcome: The proposed approach captures disagreement between the annotations of 1,000 word-sentence pairs and scores on the likelihood that each frame applies to the word.
Embedding Semantic Taxonomies (2020.coling-main)

Copied to clipboard

Challenge: Recent work on hierarchical representational structures in machine learning promises to blend the value of human curated taxonomies with the power and flexibility of machine learning systems.
Approach: They propose to use box embeddings to encode aspects of partial ordering property of taxonomies to represent a medical subject headings taxonomy.
Outcome: The proposed model outperforms baselines for taxonomic reconstruction and bipartite relationship experiments and is compared with a set of 300K PubMed articles with subject labels from MeSH.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations