Papers by Jun Rao
CommonIT: Commonality-Aware Instruction Tuning for Large Language Models via Data Partitions (2024.emnlp-main)
Copied to clipboard
| Challenge: | Current studies have focused on fine-tuning, but the use of instruction tuning is not as effective as fine-cuning. |
| Approach: | They propose a commonality-aware instruction tuning strategy to cluster instruction datasets into distinct groups with three proposed metrics Task, Embedding and Length. |
| Outcome: | The proposed strategy boosts an average improvement of 2.1% on the general domain and 5.2% on the special domain. |
Diversify Question Generation with Continuous Content Selectors and Question Type Modeling (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to generate questions based on answers and relevant contexts are not suitable for all questions . |
| Approach: | They propose a method to generate questions from a given answer and its relevant context. |
| Outcome: | The proposed method achieves a better trade-off between generation quality and diversity compared with existing approaches. |
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Current data selection paradigms rely on static, externally defined metrics, which fail to adapt to the evolving capabilities of models during training. |
| Approach: | They propose a dynamic sampling framework that aligns training data with the model's intrinsic competence by iterating on real-time feedback. |
| Outcome: | Extensive experiments on eight benchmarks show that SAI-DPO outperforms static baselines at most nearly 6 points, achieving state-of-the-art efficiency with significantly less data. |
Curriculum Consistency Learning for Conditional Sentence Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Consistency learning (CL) has proven to be a valuable technique for improving the robustness of conditional sentence generation models. |
| Approach: | They propose a strategy that guides models to learn consistency in alignment with their current capacity to differentiate between features. |
| Outcome: | The proposed strategy delivers +2.0 accuracy point improvement compared with vanilla IT and +0.7 COMET scores over traditional CL methods in MT tasks. |
APT: Improving Specialist LLM Performance with Weakness Case Acquisition and Iterative Preference Training (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models often require domain-specific fine-tuning to address targeted tasks, which risks degrading their general capabilities. |
| Approach: | They propose to use self-generated dis-preferred weakness data to enhance model performance with a targeted training approach that minimizes interference with existing knowledge base. |
| Outcome: | The proposed approach ensures no reduction in generic capacity and achieves superior performance on downstream tasks compared to existing methods. |
SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for preference optimization of large language models use pairs of positive and negative samples, but the quality of positive samples may become similar during training, complicating preference learning. |
| Approach: | SeaPO introduces error types commonly occurring in large language models to improve preference learning. |
| Outcome: | SeaPO introduces error types into model Preference Optimization to improve model performance . negative samples are more erroneous than positive samples, and preference-based training mitigates errors . |
3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset (2024.lrec-main)
Copied to clipboard
Xinyu Ma, Xuebo Liu, Derek F. Wong, Jun Rao, Bei Li, Liang Ding, Lidia S. Chao, Dacheng Tao, Min Zhang
| Challenge: | Existing studies have shown that visual information in existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. |
| Approach: | They propose to use 3AM to create an ambiguity-aware multimodal machine translation dataset. |
| Outcome: | The proposed dataset includes more ambiguity and a greater variety of captions and images than other MMT datasets. |
MathAgent: Adversarial Evolution of Constraint Graphs for Mathematical Reasoning Data Synthesis (2026.findings-acl)
Copied to clipboard
| Challenge: | Current approaches to synthesising high-quality mathematical reasoning data without human priors suffer from mode collapse and limited logical complexity. |
| Approach: | They propose a hierarchical synthesis framework that formulates data synthesis as an unsupervised optimization problem over a constraint graph followed by semantic instantiation rather than a direct text generation task. |
| Outcome: | The proposed framework outperforms widely-used datasets on eight mathematical benchmarks. |
AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to synthesis large language models often suffer from performance limitations and high computational costs. |
| Approach: | They propose a framework for constructing instruction-tuning data from unlabeled data for any specialized domains from corresponding unlabed data. |
| Outcome: | The proposed framework is comparable to DeepSeek-V3 while utilizing just 17% of the production cost. |