E-Bench: Towards Evaluating the Ease-of-Use of Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | E-Bench is a framework for easy-to-use research on large language models. |
| Approach: | They propose to evaluate the ease-of-use of large language models and construct an E-Bench . they simulate human use from synonymous and typographical perturbations . |
| Outcome: | The proposed model is able to resist synonymous expressions and typos and improves performance. |
Similar Papers
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations (2024.emnlp-main)
Copied to clipboard
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, Jimmy Huang
| Challenge: | Large Language Models (LLMs) have gained significant attention due to their capabilities in performing diverse tasks across domains. |
| Approach: | They review the primary challenges and limitations causing inconsistencies in evaluations . early models could generate coherent text but limited to simple tasks . |
| Outcome: | The proposed evaluations are reproducible, reliable, and robust. |
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language (2025.naacl-long)
Copied to clipboard
| Challenge: | Recent studies have focused on specific domains or types of persuasion, but a general study has focused on how LLMs produce persuasive text. |
| Approach: | They construct a dataset to measure and benchmark the ability of Large Language Models (LLMs) to produce persuasive text. |
| Outcome: | The proposed model can be used to generate persuasive text across domains and domains. |
Evaluating Large Language Models on Controlled Generation Tasks (2023.emnlp-main)
Copied to clipboard
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Wieting, Nanyun Peng, Xuezhe Ma
| Challenge: | Recent studies have looked into the ability of large language models in various benchmark tasks, including question generation, reading comprehension, multilingual and etc. However, few studies investigate the controllability of large languages. |
| Approach: | They propose to compare large language models with state-of-the-start finetuned smaller models to find that large language model controls are comparable to smaller models. |
| Outcome: | The proposed model can meet hard constraints and perform better than state-of-the-art models. |
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2026.acl-demo)
Copied to clipboard
| Challenge: | ACL 2026 System Demonstration Track accepted 85 papers . one paper received Best Demo award . |
| Approach: | the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026) took place from July 2-7, 2026 in San Diego, California. |
| Outcome: | the ACL 2026 System Demonstration Track accepted 85 papers based on the submitted reviews . one paper received the best demo award: The olmOCR Project: Building Fully Open OCR using VLMs . |
Pragmatics in the Era of Large Language Models: A Survey on Datasets, Evaluation, Opportunities and Challenges (2025.acl-long)
Copied to clipboard
Bolei Ma, Yuting Li, Wei Zhou, Ziwei Gong, Yang Janet Liu, Katja Jasinskaja, Annemarie Friedrich, Julia Hirschberg, Frauke Kreuter, Barbara Plank
| Challenge: | linguistics studies how context influences meaning of language and how people use it to convey implied meanings, emotions, and intentions. |
| Approach: | They analyze task designs, data collection methods, evaluation approaches and their relevance to real-world applications. |
| Outcome: | The findings highlight emerging trends, challenges, and gaps in existing benchmarks . the findings will contribute to more nuanced and context-aware NLP models . |
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2024.acl-demos)
Copied to clipboard
| Challenge: | ACL 2024 System Demonstration Track invites submissions describing system demonstrations . submissions will undergo a single-blind review process . |
| Approach: | the ACL 2024 System Demonstration Track invites submissions . papers will be published in a companion volume of the conference proceedings . submissions will undergo a single-blind review process . |
| Outcome: | the Demonstration Track at ACL 2024 is a venue for papers describing system demonstrations . publicly available open-source or open-access systems are of special interest . submissions will undergo a single-blind review process . |
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2025.acl-demo)
Copied to clipboard
| Challenge: | ACL 2025 System Demonstration Track accepted 64 papers based on reviews . short-listed 7 papers for Best System Demo award . |
| Approach: | the ACL 2025 System Demonstration Track is a conference for papers describing system demonstrations . the track received a record 187 submissions, of which 178 papers were valid with required materials . |
| Outcome: | the ACL 2025 System Demonstration Track received 187 submissions . 178 papers were valid with required materials . |
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient (2026.acl-long)
Copied to clipboard
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Yueqi Zhang, Jiayi Shi, Chuyi Tan, Boyuan Pan, Yao Hu, Kan Li
| Challenge: | Using generic and efficient benchmark generators, human annotators are limited by inefficiency . current benchmark generator methods rely on seed signals, leading to long cycles and high costs . |
| Approach: | They propose a framework to evaluate LLMs as generic benchmark generators and integrate them as BenchMaker. |
| Outcome: | The proposed framework achieves comparable performance to human-annotated benchmarks on most metrics. |
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) (2023.acl-demo)
Copied to clipboard
| Challenge: | 58 papers were selected for inclusion in the program, while a small number received only two reviews. |
| Approach: | the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023) will be held in london from July 9-14, 2023 . 58 submissions were selected for inclusion in the program, with an acceptance rate of 37%) |
| Outcome: | the system demonstration track received a record number of submissions . 58 papers were selected for inclusion in the program . |