Challenge: Large Language Models (LLMs) have demonstrated proficiency in code generation and comprehension across multiple programming languages.
Approach: They propose a parameter-localized subset of LLMs that facilitates coding capabilities.
Outcome: The proposed model significantly improves performance on coding tasks while preserving non-coding functionalities.

Similar Papers

The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have revealed their potential to perform far more than language processing tasks, showcasing abilities in reasoning and problem-solving.
Approach: They identify language-selective units within 18 popular LLMs using the same localization approach that is used in neuroscience.
Outcome: The proposed method shows that language-selective units are more aligned to brain recordings from the human language system than random units.
What can Large Language Models Capture about Code Functional Equivalence? (2025.findings-naacl)

Copied to clipboard

Challenge: SeqCoBench is a benchmark to assess how Code-LLMs can capture code semantics.
Approach: They propose a benchmark to assess how Code-LLMs capture code semantics . they use seqCoBench to evaluate whether they can discern semantically equivalent or different pairs of programs .
Outcome: The proposed benchmarks show that they can capture code semantics better than classical match-based retrieval scores.
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating the code understanding and generation capacities of Large Language Models are insufficient . existing benchmarks focus on a narrow range of popular programming languages and specific tasks .
Approach: They propose an execution-based, multilingual, multitask evaluation benchmark for LLMs . they evaluate coding performance from three dimensions: length, difficulty, efficiency .
Outcome: The proposed benchmark covers 43 programming languages and eight coding tasks.
Unveiling Linguistic Regions in Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on how LLMs achieve cross-lingual alignment and generalization have not explored the intrinsic mechanisms of how they achieve crosslingual alignment.
Approach: They propose to remove a core region that corresponds to linguistic competence and set parameters to zero to reduce performance across 30 different languages.
Outcome: The proposed model can be used to perform tasks requiring abstract knowledge and reasoning in complex languages.
How Programming Concepts and Neurons Are Shared in Code Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Several studies have focused on programming languages in a monolingual setting, but most focus on programming language models.
Approach: They perform a few-shot translation task on 21 PL pairs using two Llama-based models and decode the embeddings of intermediate layers.
Outcome: The proposed model assigns high probability to English tokens in the second half of the intermediate layers and language-specific neurons are concentrated in the bottom layers . the model's concept space is closer to English (including PL keywords) and the model is more efficient at identifying language-related neurons.
TRACE: Evaluating Execution Efficiency of LLM-Based Code Translation (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved the functional correctness of code translation, but execution efficiency remains overlooked.
Approach: They propose a benchmark to explicitly assess execution efficiency in LLM-translated code.
Outcome: The proposed benchmark identifies that execution efficiency is an essential dimension of code translation . the results highlight that correctness and efficiency are often misaligned .
Computation Mechanism Behind LLM Position Generalization (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have explored how LLMs handle positional relevance, but how they handle it remains unexplored.
Approach: They propose to enforce certain computational mechanisms to allow for the tolerance in position perturbations in large language models (LLMs) they also find a pattern in intermediate features that allows this effect to be observed .
Outcome: The proposed models can understand text with position perturbations and generalize to longer sequences than those seen during training with the latest techniques.
From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models excel at solving individual problems in isolation, but are they able to effectively collaborate over long-term interactions?
Approach: They propose to use a multi-session dataset to test LLMs' ability to track and execute simple coding instructions amid irrelevant information, simulating a realistic setting.
Outcome: The proposed model performs poorly when instructions are spread across sessions, suggesting that they are not able to integrate information over long interactions.
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions (2025.emnlp-main)

Copied to clipboard

Challenge: Exact label definitions are considered as clues to disambiguate unclear labels, helping models perform their tasks more effectively.
Approach: They conducted controlled experiments on multiple explanation benchmark datasets and label definition conditions using expert-curated, LLM-generated, perturbed, and swapped definitions.
Outcome: The results suggest that models often default to internal representations, particularly in general tasks, while domain-specific tasks benefit more from explicit definitions.
Can docstring reformulation with an LLM improve code generation? (2024.eacl-srw)

Copied to clipboard

Challenge: Existing approaches focus on training, fine-tuning or prompting LLMs to generate better outputs given the same input.
Approach: They propose to optimize part of the input, the docstring, via reformulation with an LLM to improve code generation.
Outcome: The proposed methods improve code generation on the original HumanEval benchmark and multiple curated variants on the same input.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations