Papers by Lillian Sun

1 papers
Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models (2026.acl-long)

Copied to clipboard

Challenge: Recent studies have highlighted weak-to-strong generalization, where a strong model trained only on a weak model’s labels surpasses the weak model in task performance.
Approach: They propose two fundamental fine-tuning strategies that leverage trustworthiness regularization during the fine-uning of the weak model and the weak-to-strong transfer to improve trustworthy.
Outcome: The proposed models show that they can generalize robustness, fairness, and privacy better when trained on weak models than models trained on strong models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations