Papers by Dingfan Chen
PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition Dynamics (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies have recognized hallucination as a notable concern in large autoregressive language models (LLMs). |
| Approach: | They propose a polygraph for large language models that detects "hallucination" they demonstrate that hallucination can be detected by tractable probabilistic models . |
| Outcome: | The proposed model outperforms state-of-the-art methods on open-source LLMs by 20% on TruthfulQA benchmarks. |
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race? (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies have focused on the models, neglecting the full deployment pipeline . previous studies have underestimated the practical success of these attacks . |
| Approach: | They evaluate the effectiveness of jailbreak attacks targeting LLM safety alignment . they highlight critical gaps and call for further refinement of detection accuracy and usability . |
| Outcome: | The proposed attacks can detect at least one safety filter across the entire deployment pipeline. |