Papers by Kenneth Li
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation (2026.acl-industry)
Copied to clipboard
Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li, Chenxi Whitehouse, Adina Williams, Nicola Cancedda
| Challenge: | Existing methods for obtaining well-calibrated uncertainty estimates are poorly calibrated or computationally expensive. |
| Approach: | They propose a linear probe that provides calibrated uncertainty estimates from reasoning judges’ hidden states, requiring no additional model training. |
| Outcome: | The proposed method achieves superior calibration compared to existing methods with x computational savings, generalizes robustly to unseen evaluation domains, and delivers higher accuracy on high-confidence predictions. |
Gender bias amplification during Speed-Quality optimization in Neural Machine Translation (2021.acl-short)
Copied to clipboard
| Challenge: | et al., 2002) show that gendered noun translation performance degrades faster than BLEU. |
| Approach: | They propose to use greedy search, quantization, AANs and shallow decoders to speed up decoding . they find minimal degradation of BLEU, but gendered noun translation degrades faster . |
| Outcome: | The proposed model degrades gendered noun translation performance faster than other models. |
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages (2024.emnlp-main)
Copied to clipboard
Youssef Mohamed, Runjia Li, Ibrahim Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny
| Challenge: | Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo emphasizes diversity of opinions over languages and cultures. |
| Approach: | They propose a vision-language benchmark that spans 28 languages and encompasses approximately 200,000 annotations. |
| Outcome: | The proposed benchmark spans 28 languages and encompasses approximately 200,000 annotations . the challenge is to build machine learning systems that assign emotional captions to images . |
ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture (2022.emnlp-main)
Copied to clipboard
Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny
| Challenge: | ArtELingo is a benchmark and dataset designed to encourage work on diversity across languages and cultures. |
| Approach: | They introduce a benchmark and dataset designed to encourage work on diversity across languages and cultures. |
| Outcome: | The new benchmark and dataset compared artELingo annotations across languages and cultures and found that diversity improves the performance of baseline models. |
Analyzing the Quality of Counseling Conversations: the Tell-Tale Signs of High-quality Counseling (L18-1)
Copied to clipboard
| Challenge: | Behavioral and mental health disorders are the most costly and prevalent conditions worldwide. |
| Approach: | They propose to use a dataset to analyze counseling interactions by using aspects such as mirroring, empathy, and reflective listening to build text-based classifiers. |
| Outcome: | The proposed dataset can be used to build text-based classifiers able to predict the overall quality of a counseling conversation and provide insights into the linguistic differences between low-quality and high-quality counseling. |