Papers by Fengying Ye
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Detecting AI-generated poetry is difficult due to distinctive characteristics of modern Chinese poetry. |
| Approach: | They propose a benchmark for detecting AI-generated modern Chinese poetry . they use a high-quality dataset and systematic performance assessments . |
| Outcome: | The proposed benchmark is based on a high-quality dataset of 800 poems written by six professional poets and 41,600 poems generated by four mainstream LLMs. |
Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such success actually reveals about metaphor processing. |
| Approach: | They propose to probing semantic attribute alignment, lexical invariance, and syntactic sensitivity to examine the limits of behavioral evidence for metaphor processing. |
| Outcome: | The proposed model can exhibit semantic drift relative to reference attributes, stable lexical anchors persist across contextual conditions, potentially supporting conventional metaphors while biasing novel metaphors requiring contextual integration. |
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment (2026.acl-long)
Copied to clipboard
| Challenge: | Existing tools for cross-lingual idiom-to-idiom equivalence evaluation are limited . figurative meanings are non-compositional and culturally grounded, making literal mappings unreliable. |
| Approach: | They propose a gloss-pivoted benchmark where each idiom is anchored by an English gloss from Wiktionary. |
| Outcome: | The proposed benchmark is based on a dictionary-anchored English idiom . a bias to literal translation is a dominant failure mode across diverse LLMs, the study shows . |