Challenge: Existing large language models fall short of translating culturally significant content . existing models fall behind in achieving such translations, authors say .
Approach: They propose a suitable benchmark for translating classical Chinese poetry into English . they propose RAT, a retrieval-augmented machine translation method that enhances the translation process .
Outcome: The proposed method improves translation quality in terms of adequate, fluent, and elegant translations.

Similar Papers

From Shijing to English and German: Resources and Evaluation for LLM Translation of Early Chinese Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) show promise in literary translation, but their performance in poetry remains unexplored.
Approach: They propose a framework that integrates knowledge-driven, rule-based, and LLM-as-judge metrics into a Shijing corpus . their code, lexical KB, and corpus reconstruction protocols are available at https://github.com/ML-KULeuven/ShijingLLMTrans.
Outcome: The proposed framework achieves higher human correlation than traditional metrics and high statistical stability.
Capabilities and Evaluation Biases of Large Language Models in Classical Chinese Poetry Generation: A Case Study on Tang Poetry (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly applied to creative domains, yet performance in classical Chinese poetry generation and evaluation remains poorly understood.
Approach: They propose a framework that combines computational metrics, LLM-as-a-judge assessment, and human expert validation to evaluate large language models.
Outcome: The proposed framework evaluates state-of-the-art LLMs across multiple dimensions of poetic quality in Tang poetry generation.
Neo-Classic: A Benchmark for Evaluating Linguistic-Aesthetic Reasoning in Classical Chinese Poetry (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, but it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns.
Approach: They propose a benchmark that combines a constructionist Out-of-Sample dataset with reverse understanding probes to evaluate large-scale large-format models.
Outcome: The proposed model performs well on classical Chinese poetry benchmarks, but a performance gap persists . the model can complete famous couplets and can be used to understand a variety of texts.
Poller: Are LLMs Suitable for Evaluating Poetry Understanding Task? (2026.findings-acl)

Copied to clipboard

Challenge: Traditional methods for poetry evaluation are expensive and unsuitable for large-scale data.
Approach: They propose a method leveraging Large Language Models to evaluate poetry understanding tasks using Large Language models.
Outcome: The proposed method reduces the evaluation error between LLMs and humans by adopting the poet's perspective.
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry (2025.findings-emnlp)

Copied to clipboard

Challenge: Detecting AI-generated poetry is difficult due to distinctive characteristics of modern Chinese poetry.
Approach: They propose a benchmark for detecting AI-generated modern Chinese poetry . they use a high-quality dataset and systematic performance assessments .
Outcome: The proposed benchmark is based on a high-quality dataset of 800 poems written by six professional poets and 41,600 poems generated by four mainstream LLMs.
Upping the Ante: Towards a Better Benchmark for Chinese-to-English Machine Translation (L18-1)

Copied to clipboard

Challenge: Currently, there is no widely accepted standard for evaluation of machine translation (MT) for Chinese-to-English translation, there are no standard for standardized training sets, development sets, and test sets.
Approach: They propose to use Chinese-to-English machine translation as a benchmark . they build a highly competitive state-of-the-art MT system that outperforms reported results .
Outcome: The proposed system outperforms reported results on NIST OpenMT test sets in almost all papers published in major conferences and journals in computational linguistics and artificial intelligence in the past 11 years.
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Recent research has focused on literary machine translation (MT) but evaluation of literary MT remains an open problem.
Approach: They propose a paragraph-level parallel corpus containing verified human translations and 13k evaluated sentences across four language pairs.
Outcome: The proposed corpus compares human evaluations with students and professionals . it shows that the adequacy of human evaluation is controlled by two factors .
Large-Scale Corpus Construction and Retrieval-Augmented Generation for Ancient Chinese Poetry: New Method and Data Insights (2025.findings-naacl)

Copied to clipboard

Challenge: Ancient Chinese poetry presents unique challenges for Large Language Models due to data scarcity and limited ability of general LLMs when dealing with ACP.
Approach: They propose a specialized Retrieval-Augmented Generation framework to improve LLMs' performance . they use 1.1 million ancient poems and 990K related texts to address hallucination issues .
Outcome: The proposed framework improves performance of LLMs in ancient Chinese poetry domain from 49.2% to 89.0%.
Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry (2026.acl-long)

Copied to clipboard

Challenge: a recent study shows that large language models can generate text, but they can also fabricate large amounts of false or misleading content.
Approach: They propose a benchmark to detect LLM-generated classical Chinese poetry . they compare 12 different AI detectors to find out whether a poem is authored by AI .
Outcome: The proposed benchmark compared 12 AI detectors with a dataset of 30,664 Chinese poems . the results highlight the limitations of current Chinese text detectors .
What is the Best Way for ChatGPT to Translate Poetry? (2024.acl-long)

Copied to clipboard

Challenge: Despite promising results, our analysis reveals persistent issues in the translations generated by ChatGPT that warrant attention.
Approach: They propose an Explanation-Assisted Poetry Machine Translation method which leverages monolingual poetry explanation as a guiding information for the translation process.
Outcome: The proposed method outperforms traditional translation methods of ChatGPT and the existing online systems in English-Chinese poetry translation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations