Papers by Chris Welty
How Many Ratings per Item are Necessary for Reliable Significance Testing? (2026.findings-eacl)
Copied to clipboard
| Challenge: | Existing methods for estimating model reliability are based on a few output responses per item. |
| Approach: | They propose a method to determine whether an existing dataset has enough responses per item to assure reliable null hypothesis statistical testing. |
| Outcome: | The proposed method can help researchers make better decisions about how to collect data for AI evaluation. |
Follow the leader(board) with confidence: Estimating p-values from a single test set with item and response variance (2023.findings-acl)
Copied to clipboard
| Challenge: | Among the problems with leaderboard culture in NLP has been the widespread lack of confidence estimation in reported results. |
| Approach: | They propose a framework and simulator for estimating p-values for comparisons between the results of two systems using variance found naturally (though rarely reported) in test set items and individual labels on an item (responses). |
| Outcome: | The proposed framework and simulator are used to estimate p-values for comparisons between the results of two systems under the assumption that the null hypothesis is true. |
A Crowdsourced Frame Disambiguation Corpus with Ambiguity (N19-1)
Copied to clipboard
| Challenge: | Using crowdsourcing, we have found that inter-annotator disagreement is at least partly caused by ambiguity inherent to the text and frames. |
| Approach: | They propose a crowdsourcing approach to capture inter-annotator disagreement by a list of frames with disagreement-based scores that express the confidence with which each frame applies to the word. |
| Outcome: | The proposed approach captures disagreement between the annotations of 1,000 word-sentence pairs and scores on the likelihood that each frame applies to the word. |
Embedding Semantic Taxonomies (2020.coling-main)
Copied to clipboard
| Challenge: | Recent work on hierarchical representational structures in machine learning promises to blend the value of human curated taxonomies with the power and flexibility of machine learning systems. |
| Approach: | They propose to use box embeddings to encode aspects of partial ordering property of taxonomies to represent a medical subject headings taxonomy. |
| Outcome: | The proposed model outperforms baselines for taxonomic reconstruction and bipartite relationship experiments and is compared with a set of 300K PubMed articles with subject labels from MeSH. |