CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark (2022.acl-long)
Copied to clipboard
Ningyu Zhang, Mosha Chen, Zhen Bi, Xiaozhuan Liang, Lei Li, Xin Shang, Kangping Yin, Chuanqi Tan, Jian Xu, Fei Huang, Luo Si, Yuan Ni, Guotong Xie, Zhifang Sui, Baobao Chang, Hui Zong, Zheng Yuan, Linfeng Li, Jun Yan, Hongying Zan, Kunli Zhang, Buzhou Tang, Qingcai Chen
| Challenge: | a new benchmark for biomedical language understanding is being developed in Chinese . most benchmarks are limited to English, which makes it difficult to replicate success in other languages. |
| Approach: | They propose to use Chinese biomedical language understanding evaluation benchmarks to evaluate Chinese models. |
| Outcome: | The proposed benchmarks show that the current models perform worse than the human ceiling. |
Similar Papers
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickael Rouvier, Pacome Constant Dit Beaufils, Natalia Grabar, Béatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
ExplainCPE: A Free-text Explanation Benchmark of Chinese Pharmacist Examination (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing explanation datasets for large language models are limited to the English language and general domain, leading to a scarcity of linguistic diversity and a lack of resources in specialized domains, such as medical. |
| Approach: | They propose to use a medical dataset to assess the interpretability of Large Language Models (LLMs) . they propose to analyze medical text and generate rationales for their decisions . |
| Outcome: | The proposed model passes the pharmacist examination with a 75.7% accuracy, while other models like ChatGPT fail. |
Benchmarking Automated Clinical Language Simplification: Dataset, Algorithm, and Evaluation (2022.coling-1)
Copied to clipboard
| Challenge: | Existing studies to translate medical jargon into layperson-understandable language focus on accuracy and readability aspects of clinical language. |
| Approach: | They propose to construct a dataset to support automated clinical language simplification and propose a model that mimics the human annotation procedure. |
| Outcome: | The proposed model matches human annotation procedures and achieves state-of-the-art performance compared with baselines. |
EMPEC: A Comprehensive Benchmark for Evaluating Large Language Models Across Diverse Healthcare Professions (2025.findings-acl)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) show their potential in accurately answering biomedical questions, yet current healthcare benchmarks primarily assess knowledge mastered by medical doctors, neglecting other essential professions. |
| Approach: | They evaluated 17 LLMs including proprietary and open-source models and found they struggled with specialized fields and alternative medicine. |
| Outcome: | The examinations for medical PErsonnel in Chinese (EMPEC) features 157,803 exam questions across 124 subjects and 20 healthcare professions. |
Should I Believe in What Medical AI Says? A Chinese Benchmark for Medication Based on Knowledge and Reasoning (2025.acl-short)
Copied to clipboard
| Challenge: | Large language models (LLMs) generate hallucinations when handling unfamiliar information. |
| Approach: | They propose a Chinese benchmark to evaluate large language models' knowledge and reasoning capabilities in medication tasks. |
| Outcome: | The proposed benchmark evaluates models in indication, dosage and administration, contraindicated population, mechanisms of action, drug recommendation, and drug interaction across six datasets. |
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs. |
| Approach: | They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison. |
| Outcome: | The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks. |
CliMedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models in Clinical Scenarios (2024.emnlp-main)
Copied to clipboard
| Challenge: | Chinese medical large language models (LLMs) are underperforming on this benchmark, especially where medical reasoning and factual consistency are vital. |
| Approach: | They propose a benchmark with 14 expert-guided clinical scenarios to assess the medical ability of large language models across 7 pivot dimensions. |
| Outcome: | The proposed benchmark has been validated in several ways. |
BioMegatron: Larger Biomedical Domain Language Model (2020.emnlp-main)
Copied to clipboard
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, Raghav Mani
| Challenge: | Existing studies on domain language models do not study the factors affecting performance on domain languages. |
| Approach: | They empirically evaluate factors that can affect performance on domain language applications . sub-word vocabulary set, model size, pre-training corpus, and domain transfer are important . |
| Outcome: | The results show language models trained on biomedical text perform better on biomedicine benchmarks than those trained on general domain text corpora. |
Named Entity Recognition for Chinese biomedical patents (2020.coling-main)
Copied to clipboard
| Challenge: | Existing attempts to address NER for Chinese biomedical texts have been limited due to the amount of Chinese biomedicine discoveries being patented. |
| Approach: | They train and evaluate Chinese biomedical patents NER models based on BERT . their model is optimized for Chinese bio-patent data and scored an F1 . |
| Outcome: | The proposed model achieves an F1 score of 0.540.15 for Chinese biomedical patent data. |
CMB: A Comprehensive Medical Benchmark in Chinese (2024.naacl-long)
Copied to clipboard
Xidong Wang, Guiming Chen, Song Dingjie, Zhang Zhiyi, Zhihong Chen, Qingying Xiao, Junying Chen, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, Haizhou Li
| Challenge: | Large Language Models (LLMs) provide a great breakthrough in medicine, says a new study . existing studies on LLMs leverage subjective evaluation, but evaluation in medicine is professional . |
| Approach: | They propose a localized medical benchmark in Chinese rooted in native Chinese . they propose to use traditional Chinese medicine to evaluate large-scale LLMs . |
| Outcome: | a new benchmark is developed to evaluate large-scale LLMs in china . the proposed model is rooted in the native Chinese linguistic and cultural framework . |