Papers by Qiongxiu Li
Large Language Models are Easily Confused: A Quantitative Metric, Security Implications and Typological Analysis (2025.findings-naacl)
Copied to clipboard
| Challenge: | Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate one. |
| Approach: | They propose a metric to measure and quantify language confusion in Large Language Models (LLMs) they link language confusion to LLM security and find patterns in the case of multilingual embedding inversion attacks. |
| Outcome: | The proposed metric reveals language confusion across LLMs and link it to LLM security and embedding inversion attacks. |
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities (2025.emnlp-main)
Copied to clipboard
| Challenge: | Using multilingual models, we find that treating languages in isolation obscures the true patterns of memorization. |
| Approach: | They propose a graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization. |
| Outcome: | The proposed model incorporates language similarity to analyze cross-lingual memorization in 95 languages. |
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been reported to “leak” Personally Identifiable Information (PII) successful PII reconstruction often interpreted as evidence of memorization. |
| Approach: | They propose a principled revision of memorization evaluation for Large Language Models . they propose PII leakage should be evaluated under low lexical cue conditions . |
| Outcome: | The proposed method is based on a multilingual re-evaluation of PII leakage across 32 languages and multiple memorization paradigms. |