Challenge: Currently, the top-performing models achieve a 48.8% task completion rate on realizing machine learning algorithms .
Approach: They propose a benchmark to test machine learning's ability to generate ML code for humans . they propose an automatic evaluation framework with metrics such as task pass rate and time overhead .
Outcome: The proposed benchmark is unique in its focus on interpreting complex human instructions and producing multi-step, high-complexity code.

Similar Papers

MLCopilot: Unleashing the Power of Large Language Models in Solving Machine Learning Tasks (2024.eacl-long)

Copied to clipboard

Challenge: Existing approaches to automating ML are time-consuming and difficult to understand for human developers.
Approach: They propose a framework that leverages large language models to develop ML solutions for novel tasks.
Outcome: The proposed framework bridges the gap between machine intelligence and human knowledge by exploiting state-of-the-art large language models.
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient (2026.acl-long)

Copied to clipboard

Challenge: Using generic and efficient benchmark generators, human annotators are limited by inefficiency . current benchmark generator methods rely on seed signals, leading to long cycles and high costs .
Approach: They propose a framework to evaluate LLMs as generic benchmark generators and integrate them as BenchMaker.
Outcome: The proposed framework achieves comparable performance to human-annotated benchmarks on most metrics.
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: *HumanEval* and *MBPP* are two popular benchmarks for Python code generation.
Approach: They propose a large-scale human evaluation of two popular Python benchmarks . they propose 185 hand-crafted prompts in a balanced representation of 38 programming concepts across diverse difficulty levels.
Outcome: The proposed benchmarks show a critical bias towards a limited set of programming concepts, neglecting most of the other concepts entirely.
Representation and Generation of Machine Learning Test Functions (2024.eacl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been adopted for ML code generation but their implications are relatively unexplored.
Approach: They examine the use of Large Language Models to extract representations of ML source code and tests to understand the semantic relationships between human-written tests and LLM-generated ones.
Outcome: The proposed models can be used to extract representations of ML source code and tests and annotate them for usefulness, documentation, and correctness.
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for detecting AI-generated code are limited to binary human–machine classification under in-distribution settings.
Approach: They propose to use AICD Bench to build a robust binary classification framework for large language models.
Outcome: The proposed benchmark spans 2M examples, 77 models across 11 families, and 9 programming languages.
LLM-driven Instruction Following: Progresses and Concerns (2023.emnlp-tutorial)

Copied to clipboard

Challenge: a tutorial on task instruction is aimed at researchers and practitioners interested in NLP generalization . labeled examples are unlikely to be available in large numbers or do not exist .
Approach: This tutorial will examine the progress of natural language processing (NLP) using labeled examples. authors propose that task instructions act as a novel resource for supervision.
Outcome: This tutorial aims to answer questions about instruction-driven NLP . it focuses on the use of task instructions in a low-shot scenario .
SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have made significant progress in writing code, but can they be used to reproduce results from research repositories?
Approach: They propose a benchmark to evaluate the capability of Large Language Models to reproduce results from research repositories.
Outcome: The benchmark aims to capture the realistic challenges faced by researchers working with machine learning and natural language processing repositories.
LMentry: A Language Model Benchmark of Elementary Language Tasks (2023.findings-acl)

Copied to clipboard

Challenge: Large language models are evaluated via perplexity or performance on downstream tasks, but these benchmarks are too complex and difficult to inspect.
Approach: They propose a benchmark that focuses on 25 tasks that humans are expected to perform perfectly, such as writing a sentence containing a specific word or identifying which words in a list belong to a certain category.
Outcome: The proposed benchmarks show that large language models are performing better than previous benchmarks.
MEGA: Multilingual Evaluation of Generative AI (2023.emnlp-main)

Copied to clipboard

Challenge: Large Large Models (LLMs) have shown impressive performance on many natural language processing tasks such as language understanding, reasoning, and language generation.
Approach: They present a framework for evaluating generative LLMs in the multilingual setting and provide directions for future progress in the field.
Outcome: The proposed framework evaluates generative models on 16 NLP datasets across 70 typologically diverse languages and compares them to state-of-the-art non-autoregressive models.
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in generative language models have enabled machines to generate realistic texts.
Approach: They propose a benchmark environment to test the 'Turing Test' problem for neural text generation methods.
Outcome: The proposed benchmark environment is based on 200K human- or machine-generated samples across 20 labels Human, GPT-1, GTP-2_small, GTT-2_medium, GPG-2_large, GGT-2_PyTorch, GGP-3, GROVER_base, griover_large and GRover_mega.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations