Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and AI assistants are experiencing exponential growth in usage among expert and amateur users. |
| Approach: | They propose to assess the reliability of current Large Language Models as science communicators . they use a dataset comprising 742 Yes/No queries embedded in complex scientific concepts . |
| Outcome: | The proposed model outperforms open-access models in scientific question-answering tasks . the model outpersforms GPT-4 Turbo models in many evaluation aspects . |
Similar Papers
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization (2023.findings-emnlp)
Copied to clipboard
| Challenge: | ChatGPT and GPT-4 are popular as evaluation metric for complex generative tasks . however, they are not ready as human replacements due to significant limitations . |
| Approach: | They conduct extensive analysis to examine the stability and reliability of LLMs as automatic evaluators for abstractive summarization. |
| Outcome: | The proposed methods outperform the commonly used automatic metrics but are not ready for human evaluation due to significant limitations. |
Exposing the Achilles’ Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning (2025.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations focus on final accuracy, neglecting the critical aspect of reasoning capabilities. |
| Approach: | They propose to evaluate LLMs’ abilities to detect and correct reasoning mistakes by using rule-based methods and smaller language models. |
| Outcome: | The proposed model outperforms existing models such as GPT-4o and GPT4 in both accuracy and accuracy, but lacks data contamination and memorization concerns. |
Evaluating Large Language Models on Wikipedia-Style Survey Generation (2024.findings-acl)
Copied to clipboard
Fan Gao, Hang Jiang, Rui Yang, Qingcheng Zeng, Jinghui Lu, Moritz Blum, Tianwei She, Yuang Jiang, Irene Li
| Challenge: | Recent studies have shown that large language models can perform well in general tasks, but their effectiveness and limitations in domainspecific tasks remain unclear. |
| Approach: | They examine the proficiency of Large Language Models (LLMs) in generating succinct survey articles specific to the niche field of NLP in computer science. |
| Outcome: | The LLMs perform better in generating succinct survey articles specific to the niche field of NLP in computer science, compared to human-authored surveys, but they exhibit bias in evaluation. |
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing LLM leaderboards often reference scores reported in other papers without consistent settings and prompts, which may encourage cherry-picking favored settings and for better results. |
| Approach: | They propose an open-source and reproducible LLM evaluation suite built on top of OpenAI Evals that systematically evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
| Outcome: | The evaluation suite is built on top of OpenAI Evals and evaluates 10+ leading LLMs and OpenAI’s legacy models on 20+ curated benchmarks across 7 capability categories. |
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) drive scientific question-answering on search engines, yet their evaluation robustness remains underexplored. |
| Approach: | They propose an open-source framework that combines rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. |
| Outcome: | The proposed framework combines fine-grained rubric-based assessment with reinforcement learning to mitigate optimism bias in LLM evaluators. |
How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts? (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used for question answering over long contexts . high computational costs and latency hinder the process . |
| Approach: | They explore the capabilities of Large Language Models to answer multiple questions based on the same conversational context. |
| Outcome: | The proposed models outperform proprietary and public models in question answering . their results show that they can be cost-effective and transparent . |
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work has addressed problems in unstructured grounding, multi-equation dependency, and human-aligned evaluation. |
| Approach: | They construct a dataset of scientific texts and evaluate it using an explainable equation generation workflow using automatic metrics and human judgments. |
| Outcome: | The proposed model achieves moderate performance on lexical and syntactic similarity, but struggles with semantic accuracy. |
Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work uses large language models (LLMs) to evaluate natural language process tasks, but there are shortcomings in current LLMs. |
| Approach: | They examine the alignment between LLM evaluators and human annotators by comparing conventional and alignment tasks with different evaluation criteria. |
| Outcome: | The proposed models excel in general criteria, such as fluency, but face challenges with complex criteria, including numerical reasoning. |
Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | The performance of large language models (LLMs) on existing reasoning benchmarks has significantly improved over the past decade. |
| Approach: | They propose a benchmark dataset for evaluating the problem solving abilities of large language models (LLMs) they curate 515 challenging problems from the highly competitive IIT JEE-Advanced exam. |
| Outcome: | The proposed model performs better on open-source and proprietary models than the current model, but with techniques like self-consistency, self-refinement and chain-of-thought prompting. |
LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages (2025.findings-naacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated strong performance across various natural language processing tasks, but their proficiency in mathematical reasoning remains a key challenge. |
| Approach: | They propose a process-oriented framework to evaluate LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |
| Outcome: | The proposed framework evaluates LLMs' ability to construct mathematical models, using solvers to compare outputs with ground truth. |