| Challenge: | Analogical reasoning is effective in capturing linguistic regularities. |
| Approach: | They propose to use Chinese lexical knowledge to build an analogical reasoning task using a large dataset. |
| Outcome: | The proposed dataset proves to be reliable benchmark for evaluating Chinese word embeddings. |
Similar Papers
Can language models learn analogical reasoning? Investigating training objectives and comparisons to human performance (2023.emnlp-main)
Copied to clipboard
| Challenge: | Analogical reasoning is a common way to evaluate word embeddings in NLP, but it is also of interest to investigate whether or not it is able to be learned. |
| Approach: | They propose to use proportional analogies to evaluate word embeddings in NLP . they also test whether analogical reasoning is a task in itself that can be learned . |
| Outcome: | The proposed models can learn analogical reasoning even with small amounts of data. |
Modelling Analogies and Analogical Reasoning: Connecting Cognitive Science Theory and NLP Research (2026.tacl-1)
Copied to clipboard
| Challenge: | Analogical reasoning is an essential aspect of human cognition, says aaron eliotta . eelisa e. sabet: some have argued that analogy is central to the human cognitive experience . |
| Approach: | They summarize key theories about the processes underlying analogical reasoning from the cognitive science literature and relate it to current research in natural language processing. |
| Outcome: | The proposed approaches are relevant for several major challenges in natural language processing, not directly related to analogy solving. |
Language Models at the Syntax-Semantics Interface: A Case Study of the Long-Distance Binding of Chinese Reflexive Ziji (2025.coling-main)
Copied to clipboard
| Challenge: | Existing language models tend to rely heavily on sequential cues, but not always favoring the closest strings. |
| Approach: | They construct a dataset of 320 synthetic sentences and 360 natural sentences from the BCC corpus . they evaluate 21 language models against this dataset and compare their performance to native Mandarin speakers . |
| Outcome: | The proposed models do not replicate human-like judgments in Mandarin Chinese . the results show that existing models tend to rely heavily on sequential cues . |
Comparing Static and Contextual Distributional Semantic Models on Intrinsic Tasks: An Evaluation on Mandarin Chinese Datasets (2024.lrec-main)
Copied to clipboard
| Challenge: | Distributional Semantics has undergone significant changes with the introduction of contextualized distributional models. |
| Approach: | They compare static and contextual distributional models for Mandarin Chinese . they find that static models are stronger for some of the classical tasks . |
| Outcome: | The proposed models perform better on some of the classical tasks that consider word meaning independent of context, while contextualized models excel in identifying semantic relations between word pairs and categorization of words into abstract semantic classes. |
E-KAR: A Benchmark for Rationalizing Natural Language Analogical Reasoning (2022.findings-acl)
Copied to clipboard
Jiangjie Chen, Rui Xu, Ziquan Fu, Wei Shi, Zhongqiao Li, Xinbo Zhang, Changzhi Sun, Lei Li, Yanghua Xiao, Hao Zhou
| Challenge: | Existing benchmarks to test word analogy do not reveal the underneath process of analogical reasoning of neural models. |
| Approach: | They propose an explanation benchmark for analogical reasoning using a Civil Service exam . they use a free-text explanation scheme to explain whether an analogy should be drawn . |
| Outcome: | The proposed benchmark is very challenging for state-of-the-art models, it is found. |
ANALOGICAL - A Novel Benchmark for Long Text Analogy Evaluation in Large Language Models (2023.findings-acl)
Copied to clipboard
Thilini Wijesiriwardene, Ruwan Wickramarachchi, Bimal Gajera, Shreeyash Gowaikar, Chandan Gupta, Aman Chadha, Aishwarya Naresh Reganti, Amit Sheth, Amitava Das
| Challenge: | Modern large language models are evaluated on extrinsic measures based on benchmarks such as GLUE and SuperGLUE. |
| Approach: | They propose a benchmark to intrinsically evaluate large language models across a taxonomy of analogies of long text with six levels of complexity. |
| Outcome: | The proposed benchmark evaluates LLMs across a taxonomy of analogies of long text with six levels of complexity. |
Beneath Surface Similarity: Large Language Models Make Reasonable Scientific Analogies after Structure Abduction (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have focused on word analogies, but they neglect structures that underpin analogical reasoning. |
| Approach: | They propose a task to abduct structures that form an analogy between two systems to evaluate their analogical reasoning abilities. |
| Outcome: | The proposed task is based on 400 scientific analogies from 13 different fields and is compared with a standard SCAR benchmark. |
Sentence Analogies: Linguistic Regularities in Sentence Embeddings (2020.coling-main)
Copied to clipboard
| Challenge: | Word vectors are often evaluated by assessing to what degree they exhibit regularities with regard to relationships considered in word analogies. |
| Approach: | They propose a number of schemes to induce evaluation data based on lexical analogy data as well as semantic relationships between sentences. |
| Outcome: | The proposed models reflect regularities in lexical analogies and semantic relationships between sentences. |
AnaloBench: Benchmarking the Identification of Abstract and Long-context Analogies (2024.emnlp-main)
Copied to clipboard
Xiao Ye, Andrew Wang, Jacob Choi, Yining Lu, Shreya Sharma, Lingfeng Shen, Vijay Murari Tiyyala, Nicholas Andrews, Daniel Khashabi
| Challenge: | Analogical reasoning is an important part of human communication, says a new study . a benchmark to determine analogical reasoning ability in language models is needed . |
| Approach: | They propose to benchmark analogical reasoning ability in language models by collecting 340 analogies from human writings. |
| Outcome: | The proposed benchmark aims to determine analogical reasoning ability in language models. |
Computational Modeling of Affixoid Behavior in Chinese Morphology (2020.coling-main)
Copied to clipboard
| Challenge: | affixoid behavior in Mandarin Chinese is unclear due to polysemy and diachronic dynamics. |
| Approach: | They propose to use three quantitative features to model affixoid behavior in Mandarin Chinese to determine its status. |
| Outcome: | The proposed model shows that there are no clear criteria that can be used to identify an affix’s status in an isolating language like Mandarin Chinese. |