Challenge: a recent study shows that binary code analysis is a key problem in software security research.
Approach: They propose to apply Neural Machine Translation to binary code analysis . they translate a binary in a low-resource ISA and train a model on the high-resourced ISA .
Outcome: The proposed model can be used to analyze binary code across ISAs using a high-resource ISA.

Similar Papers

Learning Cross-Architecture Instruction Embeddings for Binary Code Analysis in Low-Resource Architectures (2024.findings-naacl)

Copied to clipboard

Challenge: Applying deep learning to binary code analysis has drawn great attention because of its notable performance.
Approach: They propose to learn cross-architecture instruction embeddings where semantically-similar instructions have close embeddements in a shared space.
Outcome: The proposed approach generates high-quality CAIE with good transferability on four ISAs.
FlowMalTrans: Unsupervised Binary Code Translation for Malware Detection Using Flow-Adapter Architecture (2025.findings-emnlp)

Copied to clipboard

Challenge: Using deep learning to detect malware has attracted great attention due to its notable performance.
Approach: a new approach uses Neural Machine Translation and Normalizing Flows to apply deep learning to malware detection.
Outcome: The proposed approach reduces the burden of data collection by enabling malware detection across multiple ISAs.
Disentangled Code Representation Learning for Multiple Programming Languages (2021.findings-acl)

Copied to clipboard

Challenge: Developing effective distributed representations of source code is challenging . current code embedding approaches that represent the semantic and syntax of code are less interpretable .
Approach: They propose a disentangled code representation learning approach to separate the semantic from the syntax of source code under a multi-programming-language setting.
Outcome: The proposed approach achieves better interpretability and generalizability over existing methods.
Unifying Cross-Lingual Transfer across Scenarios of Resource Scarcity (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to deal with resource scarcity have not been developed to deal effectively with the problem.
Approach: They propose to use a set of tools to harness data from one or more high-resource "source" languages to compensate for a shortage of data in low-resourced "target" languages.
Outcome: The proposed technique can be easily adapted to unseen languages, extending the range of the proposed technique and translation-based transfer more broadly.
Addressing Leakage in Self-Supervised Contextualized Code Retrieval (2022.coling-1)

Copied to clipboard

Challenge: a recent study addresses the use of contextualized code retrieval to fill gaps in a partial input program.
Approach: They propose a self-supervised approach to contextualized code retrieval . they propose mutual identifier masking, dedentation, and the selection of syntax-aligned targets .
Outcome: The proposed approach improves retrieval substantially and yields state-of-the-art results for code clone and defect detection.
CodeTransOcean: A Comprehensive Multilingual Benchmark for Code Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing code translation datasets focus on a single pair of programming languages . early software systems are developed using programming languages such as Fortran and COBOL .
Approach: They propose a large-scale comprehensive benchmark that supports the largest variety of programming languages for code translation.
Outcome: The proposed framework supports translations between multiple programming languages and a cross-framework dataset for deep learning code across different frameworks.
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search (2025.naacl-srw)

Copied to clipboard

Challenge: clone detection is crucial in software development for identifying semantically similar code . clones can be found in the same language code snippets, but there is little research on multilingual clonage detection.
Approach: They propose a novel training procedure leveraging cross-lingual similarity to train language models on source code in various programming languages.
Outcome: The proposed method achieves state-of-the-art on C++ and Python clone detection benchmarks with comparable performance on decoder-based models.
The Vault: A Comprehensive Multilingual Dataset for Advancing Code Understanding and Generation (2023.findings-emnlp)

Copied to clipboard

Challenge: Open-source dataset of code-text pairs for training large language models to understand code is outperforms other datasets for code generation and understanding tasks.
Approach: They propose to extract high-quality code-text pairs from a dataset of 43 million pairs . they use rules and deep learning to ensure that the code-sampled samples contain high-quality pairs a .
Outcome: The Vault dataset outperforms existing models on common coding tasks . authors hope the results will propel AI research and software development forward .
Language Agnostic Code Embeddings (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies show that code language models have strong cross-lingual traits, but their multilingual representations can be dissected into a language-specific syntax component and a semantic component.
Approach: They propose to isolate and eliminate language-specific components from multilingual code embeddings to improve downstream code retrieval tasks.
Outcome: The proposed model improves retrieval tasks by removing language-specific components . the proposed model can be used to perform a variety of code generation tasks .
CodeT5+: Open Code Large Language Models for Code Understanding and Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing code LLMs adopt a specific architecture or rely on a unified encoder-decoder network for downstream tasks, lacking flexibility to operate in the optimal architecture for a particular task.
Approach: They propose to initialize code LLMs with frozen off-the-shelf LLM and explore instruction-tuning to align with natural language instructions.
Outcome: The proposed model outperforms open-source LLMs on 20 code-related benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations