Papers by Gust Verbruggen

5 papers
One-to-many testing for code generation from (just) natural language (2024.findings-emnlp)

Copied to clipboard

Challenge: MBPP relies on test cases to generate the right signature, data contamination is a problem . adapted code generation benchmarks allow for the description to be underspecified with respect to syntactic properties of code.
Approach: They propose a code generation benchmark that allows for the description to be underspecified with respect to syntactic properties of code.
Outcome: The proposed model removes ambiguity about the semantics of the task from the descriptions and evaluates generated code on multiple sets of assertions to account for ambiguities in the syntax.
TSTR: Target Similarity Tuning Meets the Real World (2023.findings-emnlp)

Copied to clipboard

Challenge: Target similarity tuning (TST) is a method of selecting relevant examples in natural language (NL) to code generation through large language models (LLMs).
Approach: They propose to use sentences from a larger language model to improve similarity between two NL inputs and associated code outputs.
Outcome: The proposed model can be trained on a small number of training examples and is cost-effective.
An empirical study of validating synthetic data for formula generation (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) can be leveraged to help write formulas in spreadsheets, but formula data resources are scarce, limiting the ability to fine-tune them.
Approach: They validate a corpus of formulas with a model to generate synthetic natural language utterances for fine-tuning.
Outcome: The proposed model generates synthetic natural language utterances with a model that is accurate enough to fine-tune them.
RAR: Retrieval-augmented retrieval for code generation in low resource languages (2024.emnlp-main)

Copied to clipboard

Challenge: Either examples or documentation are commonly used for improved code generation.
Approach: They propose retrieval augmented retrieval as a two-step method for selecting relevant examples and documentation.
Outcome: The proposed method outperforms example and grammar retrieval on low-resource languages . it also outperformed two-step retrieval when used independently .
CodeFusion: A Pre-trained Diffusion Model for Code Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing models for code generation from natural language do not allow reconsidering earlier tokens . prior work has explored grouped beam search or nucleus sampling to generate diverse text.
Approach: They propose a diffusion code generation model that iteratively denoises a program conditioned on the encoded natural language.
Outcome: The proposed model outperforms state-of-the-art models in accuracy and diversity compared to existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations