Papers by Colin Clement

4 papers
Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy (2021.emnlp-main)

Copied to clipboard

Challenge: Statistical language modeling and translation with transformers have found many successful applications in program understanding and generation tasks.
Approach: They propose an architecture-independent approach for leveraging syntactic hierarchies of source code . they use syntax trees to extract syntak hierarchical structures and integrate them into context window .
Outcome: The proposed approach achieves state-of-the-art in code completion and summarization for Python in the CodeXGLUE benchmark.
Program Translation via Code Distillation (2023.emnlp-main)

Copied to clipboard

Challenge: Software version migration and program translation are costly parts of the lifecycle of large codebases.
Approach: They propose a model that captures semantic and structural equivalence of code in a language agnostic intermediate representation.
Outcome: The proposed model achieves state-of-the-art performance on CodeXGLUE and TransCoder GeeksForGeeks translation benchmarks.
SUT: Active Defects Probing for Transcompiler Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing datasets are often criticized for their lack of granularity, which can mask deficiencies in basic syntactic elements that humans care about.
Approach: They propose a new program translation metrics that address basic syntax errors . they propose BLUE, CodeBLUE and computation accuracy metrics which address these errors based on a highly interpretable evaluation harness.
Outcome: The proposed model passes the unit tests with a 26.15% pass rate compared to previous models .
PyMT5: multi-mode translation of natural language and Python code with transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Using Python method text-to-text transfer transformers, developers can easily model source code and natural language.
Approach: They propose a Python method text-to-text transfer transformer that can translate between all pairs of Python method feature combinations.
Outcome: The proposed model outperforms similar-sized auto-regressive language models on a large-scale parallel corpus of 26 million methods and 7.7 million method-docstring pairs on the CodeSearchNet test set.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations