Papers by Yifeng Ding

4 papers
\mathcal XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts (2024.acl-long)

Copied to clipboard

Challenge: Existing studies focus on the data perspectives of instruction tuning, leaving room for exploring advanced training schemes.
Approach: They argue that prior works overlook the possibility of improving code instruction tuning by advancing existing training schemes.
Outcome: The proposed model is dense because all parameters are activated to predict the next token (assuming it is a decoder-only LLM).
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization (2026.acl-long)

Copied to clipboard

Challenge: Current reinforcement learning methods suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation.
Approach: They propose a novel RL algorithm for training large language models for multi-turn tool-integrated reasoning (TIR) that incorporates three innovations: turn-level reward assignment that provides fine-grained feedback for individual turns, return-based advantage estimation where normalized discounted returns are calculated as advantages, and self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards.
Outcome: The proposed algorithm outperforms GRPO by 3.0% across diverse math reasoning benchmarks and improves grepo by 3.9% on commonsense reasoning and program synthesis tasks.
Planning-Aware Code Infilling via Horizon-Length Prediction (2025.emnlp-main)

Copied to clipboard

Challenge: Current approaches to fill-in-the-middle (FIM) often fail to generate content that aligns well with the surrounding context.
Approach: They propose a training objective that teaches models to predict the number of remaining middle tokens at each step.
Outcome: The proposed training objective improves FIM performance by up to 24% on diverse benchmarks across file-level and repository-level.
Fusion or Defusion? Flexible Vision-and-Language Pre-Training (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to vision-and-language pretraining (VLP) lack effectiveness and efficiency in downstream multimodal tasks.
Approach: They propose a flexible vision-and-language pre-training model by incorporating cross-modal fusions into a dual-encoder architecture and a cross-module knowledge transfer strategy to guide the training process.
Outcome: The proposed model is well-equipped with effectiveness and efficiency compared with other strong VLP models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations