Papers by Kenneth Li

5 papers
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation (2026.acl-industry)

Copied to clipboard

Challenge: Existing methods for obtaining well-calibrated uncertainty estimates are poorly calibrated or computationally expensive.
Approach: They propose a linear probe that provides calibrated uncertainty estimates from reasoning judges’ hidden states, requiring no additional model training.
Outcome: The proposed method achieves superior calibration compared to existing methods with x computational savings, generalizes robustly to unseen evaluation domains, and delivers higher accuracy on high-confidence predictions.
Gender bias amplification during Speed-Quality optimization in Neural Machine Translation (2021.acl-short)

Copied to clipboard

Challenge: et al., 2002) show that gendered noun translation performance degrades faster than BLEU.
Approach: They propose to use greedy search, quantization, AANs and shallow decoders to speed up decoding . they find minimal degradation of BLEU, but gendered noun translation degrades faster .
Outcome: The proposed model degrades gendered noun translation performance faster than other models.
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Traditionally, vision research focused on unambiguous class labels, whereas ArtELingo emphasizes diversity of opinions over languages and cultures.
Approach: They propose a vision-language benchmark that spans 28 languages and encompasses approximately 200,000 annotations.
Outcome: The proposed benchmark spans 28 languages and encompasses approximately 200,000 annotations . the challenge is to build machine learning systems that assign emotional captions to images .
ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and Culture (2022.emnlp-main)

Copied to clipboard

Challenge: ArtELingo is a benchmark and dataset designed to encourage work on diversity across languages and cultures.
Approach: They introduce a benchmark and dataset designed to encourage work on diversity across languages and cultures.
Outcome: The new benchmark and dataset compared artELingo annotations across languages and cultures and found that diversity improves the performance of baseline models.
Analyzing the Quality of Counseling Conversations: the Tell-Tale Signs of High-quality Counseling (L18-1)

Copied to clipboard

Challenge: Behavioral and mental health disorders are the most costly and prevalent conditions worldwide.
Approach: They propose to use a dataset to analyze counseling interactions by using aspects such as mirroring, empathy, and reflective listening to build text-based classifiers.
Outcome: The proposed dataset can be used to build text-based classifiers able to predict the overall quality of a counseling conversation and provide insights into the linguistic differences between low-quality and high-quality counseling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations