| Challenge: | emergence of ChatGPT has generated speculation about its potential to disrupt social and economic systems. |
| Approach: | They analyze prior assessments of ChatGPT and GPT-4 to analyze their language and reasoning abilities, scientific knowledge, ethical considerations and existing evaluation methods. |
| Outcome: | The proposed model performs satisfactorily in science knowledge and can answer open questions. |
Similar Papers
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets (2023.findings-acl)
Copied to clipboard
Md Tahmid Rahman Laskar, M Saiful Bari, Mizanur Rahman, Md Amran Hossen Bhuiyan, Shafiq Joty, Jimmy Huang
| Challenge: | Currently, the evaluation of large language models (LLMs) such as ChatGPT in academic datasets is difficult due to the difficulty of evaluating the generative outputs produced by this model against the ground truth. |
| Approach: | They evaluate ChatGPT across 140 tasks and analyze 255K responses it generates in academic datasets. |
| Outcome: | The proposed model performs well on 140 tasks and generates 255K responses in these datasets. |
GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP (2023.emnlp-main)
Copied to clipboard
| Challenge: | Our study examines ChatGPT’s performance on Arabic languages and dialectal varieties. |
| Approach: | They conduct a large-scale automated and human evaluation of ChatGPT, encompassing 44 distinct language understanding and generation tasks on over 60 different datasets. |
| Outcome: | The proposed model outperforms smaller models on Arabic dialects compared to GPT-4's Modern Standard Arabic and Dialectal Arabic (DA) |
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks (2023.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) evaluation is a patchy and inconsistent landscape . established automatic evaluation metrics are poor surrogates, correlating weakly with human judgement. |
| Approach: | They propose to use both automatic and human evaluation to evaluate generative LLMs on three NLP benchmarks: text summarisation, text simplification and grammatical error correction. |
| Outcome: | The proposed model outperforms many popular models according to human reviewers on the majority of metrics, while scoring much worse when using classic automatic evaluation metrics. |
Credible without Credit: Domain Experts Assess Generative Language Models (2023.acl-short)
Copied to clipboard
| Challenge: | ChatGPT has been criticized for its lack of accuracy and coherence . authors argue that language models could replace search engines and make college essays obsolete . |
| Approach: | a team of 10 domain experts conducts an initial assessment of language models using 100 expert-written questions. |
| Outcome: | The results show that language models are mixed in their accuracy. |
Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Recent large language models such as ChatGPT and GPT-4 have shown exceptional capabilities of generalist models . however, their applicability and effectiveness in specific domains like finance needs a better understanding . |
| Approach: | They conduct empirical studies to compare the performance of ChatGPT and GPT-4 on financial text analytical problems using eight benchmark datasets from five categories of tasks. |
| Outcome: | The proposed models outperform the state-of-the-art models on a wide range of financial text analytical tasks. |
Can ChatGPT Assess Human Personalities? A General Evaluation Framework (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies study the virtual personalities of LLMs but rarely explore the possibility of analyzing human personalities via LLM. |
| Approach: | They propose to use Myers–Briggs Type Indicator (MBTI) tests to generate unbiased prompts and replace the subject in question statements to enable flexible queries and assessments. |
| Outcome: | The proposed framework enables LLMs to flexibly assess personalities of different groups of people. |
Testing the Depth of ChatGPT’s Comprehension via Cross-Modal Tasks Based on ASCII-Art: GPT3.5’s Abilities in Regard to Recognizing and Generating ASCII-Art Are Not Totally Lacking (2024.findings-eacl)
Copied to clipboard
| Challenge: | In the months since its release, ChatGPT and its underlying model, GPT3.5, have garnered massive attention due to their potent mix of capability and accessibility. |
| Approach: | They examine GPT3.5's aptitude for visual tasks using ASCII-art without overt distillation into a lingual summary. |
| Outcome: | The proposed model performs well on image recognition and generation tasks. |
Consistency Analysis of ChatGPT (2023.emnlp-main)
Copied to clipboard
| Challenge: | ChatGPT and GPT-4 have been reported to be more reliable and trustworthy, provided they behave similarly to humans. |
| Approach: | They propose to compare ChatGPT and GPT-4 in terms of logically consistent behaviour and the properties of negation, symmetric, and transitive consistency. |
| Outcome: | The proposed models show that they can be more reliable and trustworthy provided they behave similarly to humans. |
Evaluating ChatGPT against Functionality Tests for Hate Speech Detection (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models like ChatGPT have shown a great promise in detecting hate speech, but they lack the capability to perform in a holistic fashion. |
| Approach: | They evaluate the ChatGPT model's strengths and weaknesses by performing functional tests across 11 languages to uncover their weaknesses. |
| Outcome: | The proposed model performs poorly across 11 languages and is based on functional tests. |
Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical Study (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored. |
| Approach: | They aim to systematically inspect ChatGPT’s performance in two discourse analysis tasks: topic segmentation and discourse parsing. |
| Outcome: | The proposed model can give more reasonable topic structures than human annotations but only linearly parses the hierarchical rhetorical structures. |