Papers by Orr Paradise
Pseudointelligence: A Unifying Lens on Language Model Evaluation (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies claim that language models surpass human performance on new benchmarks within a few years. |
| Approach: | They propose a framework for model evaluation that casts as a dynamic interaction between a model and a learned evaluator. |
| Outcome: | The proposed framework can be used to reason about two case studies in language model evaluation, and analyze existing evaluation methods. |