Program Translation via Code Distillation (2023.emnlp-main)

Copied to clipboard

Challenge: Software version migration and program translation are costly parts of the lifecycle of large codebases.
Approach: They propose a model that captures semantic and structural equivalence of code in a language agnostic intermediate representation.
Outcome: The proposed model achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks.

Similar Papers

Contrastive Distillation on Intermediate Representations for Language Model Compression (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods to compress language models use a simple L_2 loss to distill knowledge in the intermediate representations of a large BERT model to a smaller one.
Approach: They propose a method that uses knowledge distillation to distill knowledge through intermediate layers of the teacher via a contrastive objective.
Outcome: The proposed method outperforms state-of-the-art methods on the GLUE benchmark.
Task-agnostic Distillation of Encoder-Decoder Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Existing distillation methods that focus on encoder-only LMs fail to handle the distillation of encoder decoder LM.
Approach: They propose a method that finetunes pretrained language models (LMs) they propose 'MiniEnD' that allows for task-agnostic distillation of LMs.
Outcome: The proposed distillation method is generally effective and competitive compared to other alternatives.
CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing code translation datasets focus on a single pair of programming languages . early software systems are developed using programming languages such as Fortran and COBOL .
Approach: They propose a large-scale comprehensive benchmark that supports the largest variety of programming languages for code translation.
Outcome: The proposed framework supports translations between multiple programming languages and a cross-framework dataset for deep learning code across different frameworks.
Self-Distillation for Model Stacking Unlocks Cross-Lingual NLU in 200+ Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) excel on English NLU tasks, yet struggle to extend their NLU capabilities to underrepresented languages.
Approach: They integrate machine translation models (MT) directly into LLM backbones via sample-efficient self-distillation.
Outcome: The proposed model outperforms translation-test models on 127 low-resource languages.
Distillation of encoder-decoder transformers for sequence labelling (2023.findings-eacl)

Copied to clipboard

Challenge: despite the strong trend in NLP to explore the use of large language models, there is still limited work evaluating prompting and decoding mechanisms for SL tasks.
Approach: They propose a hallucination-free framework for sequence tagging that is especially suited for distillation.
Outcome: The proposed framework performs well across multiple sequence labelling datasets and in a few-shot learning scenario.
What can Large Language Models Capture about Code Functional Equivalence? (2025.findings-naacl)

Copied to clipboard

Challenge: SeqCoBench is a benchmark to assess how Code-LLMs can capture code semantics.
Approach: They propose a benchmark to assess how Code-LLMs capture code semantics . they use seqCoBench to evaluate whether they can discern semantically equivalent or different pairs of programs .
Outcome: The proposed benchmarks show that they can capture code semantics better than classical match-based retrieval scores.
Continual Knowledge Distillation for Neural Machine Translation (2023.acl-long)

Copied to clipboard

Challenge: Current parallel corpora are not publicly accessible but trained models are more readily available.
Approach: They propose a method to take advantage of existing translation models to improve one model of interest.
Outcome: The proposed method improves on Chinese-English and German-English datasets and is robust to malicious models.
Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy (2021.emnlp-main)

Copied to clipboard

Challenge: Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks.
Approach: They propose an architecture-independent approach for leveraging syntactic hierarchies of source code . they use syntax trees to extract syntak hierarchical structures and integrate them into context window .
Outcome: The proposed approach achieves state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark.
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks (2026.acl-long)

Copied to clipboard

Challenge: Existing large language models for software engineering rely on coarse-grained pass rates obscuring specific cognitive bottlenecks.
Approach: They propose a repository-level benchmark that dissects coding capabilities through atomized tasks.
Outcome: The proposed framework achieves a 78.55% validity yield, surpassing the 31.7% retention rate of SWE-bench-Verified.
AVATAR: A Parallel Corpus for Java-Python Program Translation (2023.findings-acl)

Copied to clipboard

Challenge: Program translation is a time-consuming and costly process that requires expertise in both the source and target languages.
Approach: They present a collection of 9,515 programming problems and their solutions written in Java and Python.
Outcome: The proposed model lacks in generating functionally accurate code.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations