Papers by Benjamin Negrevergne
Exploring Precision and Recall to assess the quality and diversity of LLMs (2024.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models are limited to specific tasks, but they are now widely available for a wide range of tasks. |
| Approach: | They propose a framework for large language models such as Llama-2 and Mistral that imports precision and recall metrics from image generation to text generation. |
| Outcome: | The proposed framework allows for a nuanced assessment of the quality and diversity of generated text without the need for aligned corpora. |