Papers by Tsung-Yi Ho

4 papers
GRE Score: Generative Risk Evaluation for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models have revolutionized generative tasks, but concerns about their trustworthiness and vulnerability to adversarial attacks persist.
Approach: They propose an attack-independent evaluation of LLM robustness using conditional generation for synthetic text creation and a method to quantify the model's resilience.
Outcome: The proposed method achieves a consistent ranking of LLM robustness when compared to the attack-based model ranking on TrustLLM (CITATION).
Hey, That’s My Data! Token-Only Dataset Inference in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing dataset inference methods require logit access, but many modern LLMs restrict such access.
Approach: They propose a token-only dataset inference framework that allows models to overwrite prior knowledge when trained on new data.
Outcome: The proposed framework overwrites prior knowledge when trained on new data.
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets (2026.acl-long)

Copied to clipboard

Challenge: Existing mitigation strategies focus on reactively addressing jailbreak incidents after safety guardrails have been compromised.
Approach: They investigate the degradation of safety guardrails through the lens of representation similarity between upstream alignment datasets and downstream fine-tuning tasks.
Outcome: The proposed model reduces harmfulness score by 10.33% when compared to baseline models.
Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have showcased their ability to understand and generate text akin to human interaction.
Approach: They propose a prompt-based defense mechanism specifically designed to protect LLMs against jailbreak attacks by introducing jailbreak prompts into malicious queries.
Outcome: Empirical results show that the proposed defense outperforms existing defense strategies in balancing safety and utility while maintaining high utility.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations