Papers by Aleksander Jędrosz
Can Models Help Us Create Better Models? Evaluating LLMs as Data Scientists (2026.findings-eacl)
Copied to clipboard
| Challenge: | Current benchmarks assess LLMs on more isolated capabilities, such as language understanding and question-answering. |
| Approach: | They propose a benchmark to evaluate the ability of large language models (LLMs) to perform feature engineering. |
| Outcome: | The proposed benchmark evaluates the ability of large language models to perform feature engineering, a critical and knowledge-intensive task in data science. |