Challenge: Existing instruction-tuning datasets lack execution-grounded supervision and offer limited support for iterative code correction.
Approach: They propose a large-scale instruction tuning dataset for Python-based visualization and self-correction.
Outcome: The proposed dataset outperforms strong open-source baselines and proprietary models like GPT-4o-mini.

Similar Papers

InstructCoder: Instruction Tuning Large Language Models for Code Editing (2024.acl-srw)

Copied to clipboard

Challenge: InstructCoder is the first instruction-tuning dataset designed to adapt LLMs for general-purpose code editing.
Approach: They propose to use Large Language Models to edit code based on user instructions . they use a dataset to adapt LLMs to general-purpose code editing .
Outcome: The proposed model can significantly improve code editing performance compared to proprietary models . the proposed model is based on a human-written execution-based benchmark .
Text2Chart31: Instruction Tuning for Chart Generation with Automatic Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Existing datasets do not cover full range of chart types, such as 3D, volumetric, and gridded charts.
Approach: They propose a hierarchical pipeline and a new dataset for chart generation that leverages the relationships within rich datasets.
Outcome: The proposed method outperforms open-source models and is comparable to state-of-the-art proprietary models in data visualization tasks.
WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning (2024.acl-long)

Copied to clipboard

Challenge: Recent work shows that Code Large Language Models can address a wide range of code-related tasks.
Approach: They propose a method to generate widespread and versatile instruction data from open source code datasets and use it to train code-related models.
Outcome: The proposed model outperforms open-source models in generalization ability across code-related tasks.
FrontCoder: Scaling Visual Fidelity in Front-End Code Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing work on front-end code generation fails to provide visual fidelity and rendering quality for front- end developers.
Approach: They propose a three-stage pipeline to enhance front-end code generation capabilities in LLMs . they use synthetic data, quality-controlled supervised fine-tuning, and reinforcement learning .
Outcome: The proposed model achieves competitive performance with frontier models while maintaining generation efficiency.
UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback (2024.naacl-long)

Copied to clipboard

Challenge: Existing approaches to improve UI code generation rely on expensive human feedback or distilling a proprietary model.
Approach: They propose to use automated feedback to guide large language models to generate UI code . they use a large synthetic dataset to generate improved models and refine them .
Outcome: The proposed model outperforms baseline models and larger proprietary models . the model outpersforms models with automated metrics and human preferences .
How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data (2024.emnlp-main)

Copied to clipboard

Challenge: Recent research has shown that code pre-trained models improve coding capabilities.
Approach: They propose a code data pruning strategy to identify which datasets are high-quality code instruction data.
Outcome: The proposed model achieves state-of-the-art performance using fewer training data.
Turning the Tide: Repository-based Code Reflection (2025.findings-emnlp)

Copied to clipboard

Challenge: Code large language models (LLMs) enhance programming by understanding and generating code across languages.
Approach: a new benchmark evaluates code understanding and generation in repositories using code large language models.
Outcome: The proposed model improves code understanding and generation in repositories by evaluating 1,888 test cases across 6 programming languages.
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have become a dominant tool for NLP researchers in a wide range of tasks.
Approach: They propose an open source Python library that allows researchers to write simple code to implement powerful LLM workflows.
Outcome: The proposed library is open source and can be used to implement powerful LLM workflows.
Integrating Symbolic Execution into the Fine-Tuning of Code-Generating LLMs (2025.naacl-srw)

Copied to clipboard

Challenge: Code-generating Large Language Models (LLMs) have become essential tools in modern software development, enhancing productivity and accelerating development.
Approach: They propose to use Reinforcement Learning and Direct Preference Optimization to fine-tune code-generating Large Language Models (LLMs) by enhancing the training data with symbolic execution techniques.
Outcome: The proposed model improves on the CodeRL benchmark and shows that it is more accurate and objective than the baseline model.
CodecLM: Aligning Language Models with Tailored Synthetic Data (2024.findings-naacl)

Copied to clipboard

Challenge: Recent work on generating diverse instructions and applying LLM to increase instruction complexity neglects downstream use cases.
Approach: They propose a framework for generating high-quality synthetic data for LLM alignment with different downstream instruction distributions and LLMs.
Outcome: Experiments on four open-domain instruction using the proposed framework validate the effectiveness of CodecLM over the current state-of-the-art.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations