Papers by Xiangyu Dong

9 papers
LLMTreeRec: Unleashing the Power of Large Language Models for Cold-Start Recommendations (2025.coling-main)

Copied to clipboard

Challenge: Lack of training data leads to the system cold-start problem in recommendation systems, making them struggle to provide effective recommendations.
Approach: They propose a tree-based LLM recommendation framework which structures all items into an item tree to improve the efficiency of LLM’s item retrieval.
Outcome: The proposed framework outperforms the baseline model in the A/B test on Huawei industrial system.
Multimodal Cross-lingual Phrase Retrieval (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to cross-lingual phrase retrieval only deal with textual modality, leaving the question of the effectiveness of using multimodal information unanswered.
Approach: They propose a multimodal cross-lingual phrase retrieval resource that integrates a Wikimedia Commons media store and a large multimodal pre-trained model to bridge the gap between different modalities.
Outcome: The proposed approach performs significantly better than pure textual cross-lingual phrase retrieval on a benchmarked dataset covering eight language pairs.
Third-Party Aligner for Neural Word Alignments (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work shows that word alignment can be competitive .
Approach: They propose to use word alignments generated by a third-party word aligner to supervise the neural word alignment training.
Outcome: The proposed approach can find more accurate word alignments and delete wrong alignments, leading to better performance than the current best third-party word aligner.
Don’t Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in tree search algorithms guided by verifiers have significantly enhanced the reasoning capabilities of large language models (LLMs), but at the cost of increased computational resources.
Approach: They propose an e ffici ent tree sear ch framework that is a plug-and-play system compatible with various tree search algorithms.
Outcome: The proposed framework reduces computational costs and prioritizes resource allocation to harder tasks (Levels 3-4) over simpler ones (Level 1-2), addressing both over-exploration in basic problems and under-exploation in complex cases.
MedDialog: Large-scale Medical Dialogue Datasets (2020.emnlp-main)

Copied to clipboard

Challenge: telemedicine is a medical practice that provides patient care remotely using video conferencing tools.
Approach: They build large-scale medical dialogue datasets to facilitate research . they pretrain several models on the Chinese MedDialog dataset and compare their performance .
Outcome: The proposed datasets show that models trained on MedDialog can generate doctor-like medical dialogues.
Logic-Consistency Text Generation from Semantic Parses (2021.findings-acl)

Copied to clipboard

Challenge: Text generation from semantic parses is challenging due to the complexity of the inner logic and the lack of automatic evaluation metrics for logic consistency.
Approach: They propose a framework for logic consistent text generation from semantic parses that employs iterative training procedures and quality control.
Outcome: The proposed framework enhances logic consistency and human evaluation on two benchmark datasets.
Select Before Use: On the Importance of Reference Model Selection in Preference Alignment (2026.acl-long)

Copied to clipboard

Challenge: Supervised Fine-Tuning (SFT) is used as the initialization and reference model for subsequent preference alignment.
Approach: They propose to use RewardRank to estimate initial implicit alignment between reference model and preference objective to ensure LLMs generate safe, helpful, and instruction-aligned content.
Outcome: Empirical evidence shows that using the selected model as reference can gain up to 67.6% relative increase on length-controlled win rate compared to baselines.
Disambiguated Lexically Constrained Neural Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Current approaches to LCNMT assume that pre-specified lexicon constraints are contextually appropriate.
Approach: They propose a framework that disambiguates constraints based on contexts at first and integrates them into LCNMT.
Outcome: The proposed approach outperforms baseline approaches on benchmark datasets and comprehensive experiments in multiple target constraints.
Injecting Entity Types into Entity-Guided Text Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Recent advances in deep generative modeling have led to significant advances in natural language generation (NLG).
Approach: They propose to model the entity type carefully in the decoding phase to generate contextual words accurately.
Outcome: The proposed model produces a target sequence based on a given list of entities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations