Papers by Jingyan Shen
The JDDC Corpus: A Large-Scale Multi-Turn Chinese Dialogue Dataset for E-commerce Customer Service (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing datasets for human-like dialogue tasks are deficient due to the complexity of human conversations. |
| Approach: | They construct a large-scale Chinese E-commerce conversation corpus with 1 million dialogues, 20 million utterances, and 150 million words. |
| Outcome: | The proposed dataset includes 1 million multi-turn dialogues, 20 million utterances, and 150 million words. |
Rethinking Diverse Human Preference Learning through Principal Component Analysis (2025.findings-acl)
Copied to clipboard
| Challenge: | Decomposed Reward Models extract diverse human preferences from binary comparisons without fine-grained annotations. |
| Approach: | They propose a decomposed reward model that extracts diverse human preferences from binary comparisons without fine-grained annotations. |
| Outcome: | The proposed approach extracts diverse human preferences from binary comparisons without fine-grained annotations. |
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing reward models assume a global reward function, limiting personalization and pluralistic alignment. |
| Approach: | They propose a framework that leverages binary preference datasets to enhance personalized preference learning. |
| Outcome: | The proposed framework captures diverse human preferences without fine-grained annotations and significantly improves personalized preference learning on downstream tasks. |
Beyond A Single AI Cluster: A Survey of Decentralized LLM Training (2025.emnlp-main)
Copied to clipboard
| Challenge: | Decentralized LLM training leverages dispersed resources at varying scales. |
| Approach: | They propose a resource-driven paradigm that leverages dispersed resources across clusters, datacenters and even regions. |
| Outcome: | The proposed model scales are 175 billion to 660 billion parameters, and the exponential growth in computational requirements poses significant challenges. |