Challenge: Figures of speech often deviate from their literal meanings to express deeper semantic implications.
Approach: They propose a concept of figurative unit, which is the carrier of a figure, and build a Chinese corpus for Contextualized Figure Recognition.
Outcome: The proposed model is based on 12 types of figures commonly used in Chinese . it shows that the proposed tasks are challenging for existing models .

Similar Papers

Using a Corpus of English and Chinese Political Speeches for Metaphor Analysis (L18-1)

Copied to clipboard

Challenge: specialized corpora on a variety of topics are available online, but online corporates are scarce.
Approach: They propose to create a corpus of political speeches and use it for metaphor analysis . they propose to use the database to search for lexical frequencies and collocation lists .
Outcome: The proposed corpus contains more than six million speeches in English and Chinese and is available for free online.
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation (2024.lrec-main)

Copied to clipboard

Challenge: Metaphors are a prominent linguistic device in human language and literature, as they add color, imagery, and emphasis to enhance effective communication.
Approach: They propose a large-scale high quality annotated Chinese Metaphor Corpus . they use a set of guidelines to ensure the accuracy and consistency of their annotations .
Outcome: The proposed corpus generates metaphors that resonate more with real-world intuition.
IRFL: Image Recognition of Figurative Language (2023.findings-emnlp)

Copied to clipboard

Challenge: Figures of speech are ubiquitous in many forms of discourse, allowing people to convey complex, abstract ideas and evoke emotion.
Approach: They develop a dataset for multimodal figurative language understanding using human annotation and an automatic pipeline to generate a multimodal dataset.
Outcome: The proposed dataset performs better than human vision and language models compared with a human dataset .
Language Models at the Syntax-Semantics Interface: A Case Study of the Long-Distance Binding of Chinese Reflexive Ziji (2025.coling-main)

Copied to clipboard

Challenge: Existing language models tend to rely heavily on sequential cues, but not always favoring the closest strings.
Approach: They construct a dataset of 320 synthetic sentences and 360 natural sentences from the BCC corpus . they evaluate 21 language models against this dataset and compare their performance to native Mandarin speakers .
Outcome: The proposed models do not replicate human-like judgments in Mandarin Chinese . the results show that existing models tend to rely heavily on sequential cues .
Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and Benchmark (2024.lrec-main)

Copied to clipboard

Challenge: Compared with sentence-level topic structure, paragraph-level topics can grasp and understand the context of a document from a higher level.
Approach: They propose a hierarchical paragraph-level topic structure representation with three layers to guide corpus construction.
Outcome: The proposed method achieves the largest Chinese paragraph-level topic structure corpus, achieving high quality.
CLGC: A Corpus for Chinese Literary Grace Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Literature grace is a key element of the style and quality of articles in China.
Approach: They propose to annotate a Chinese literary grace corpus with 10,000 texts and 1.85 million tokens and build a literary grace evaluation task to assess the literary grace level.
Outcome: The proposed model achieves 79.71% on the weighted average F1-score.
Input Representations for Parsing Discourse Representation Structures: Comparing English with Chinese (2021.acl-short)

Copied to clipboard

Challenge: Neural semantic parsers have obtained acceptable results in parsing DRSs . previous studies have focused on parse of DRS in English, but have focused only on a few languages .
Approach: They propose to use character sequences as input to map meaning representations to string format.
Outcome: The proposed models learn the meaning of a series of semantic phenomena by taking sentences as input and outputting the corresponding DRSs, without the aid of any extra linguistic information.
Shallow Discourse Annotation for Chinese TED Talks (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods to annotate text with discourse properties are limited to newspaper articles and are not available in Chinese.
Approach: They propose to annotate TED talks with Chinese-related properties using the Penn Discourse TreeBank annotation style . they propose to use planned monologues instead of written text to annnotate Chinese-specific properties.
Outcome: The proposed method is able to achieve reliable results in Chinese spoken monologues, and is based on the Penn Discourse TreeBank annotation style.
CHENGYU-BENCH: Benchmarking Large Language Models for Chinese Idiom Understanding and Use (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks focus on narrow tasks such as multiple-choice cloze tests, isolated translation, or simple paraphrasing.
Approach: They propose a benchmark to measure Chinese idioms' cultural and contextual nuances . they evaluate 2,937 human-verified examples covering 1,765 common idiomes .
Outcome: The proposed benchmarks achieve 95% accuracy on Evaluative Connotation, but only 85% on Appropriateness and 40% top-1 accuracy in Open Cloze.
Approaches and Challenges for Resolving Different Representations of Fictional Characters for Chinese Novels (2024.lrec-main)

Copied to clipboard

Challenge: Existing automatic text analysis tools and models are often developed for generic, open-domain tasks, restricting in-depth literary studies.
Approach: They adapt a state-of-the-art anaphora resolution model to resolve character representations in Chinese novels by making some modifications and train a widely used BERT fine-tuned model for speaker extraction as assistance.
Outcome: The proposed model is modified to resolve character representations in Chinese novels and train a BERT fine-tuned model for speaker extraction as assistance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations