Papers by Davood Wadi
A Monte-Carlo Sampling Framework For Reliable Evaluation of Large Language Models Using Behavioral Analysis (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Current approaches to evaluation of large language models ignore high entropy of LLM responses. |
| Approach: | They propose a Monte-Carlo evaluation framework for evaluating large language models . they test multiple LLMs to see if they are susceptible to cognitive biases . |
| Outcome: | The proposed framework shows that LLMs are more human-like and less rational . it also shows that larger LLM models are more susceptible to cognitive biases . |