Papers by Xiaoyu Luo

6 papers
Generative Annotation for ASR Named Entity Correction (2025.emnlp-main)

Copied to clipboard

Challenge: Existing named entity correction models fail to transcribe domain-speciffcnamed entities when theforms of the wrongly-transcribed words and the ground-truth entity are signiffcantly different.
Approach: They propose a method that utilizes speech sound features to retrieve candidate entities . it uses speech sound feature to annotate entityerrors in ASR transcripts .
Outcome: The proposed method can bring signiffcant improvement to entity accuracy.
Convert Language Model into a Value-based Strategic Planner (2025.acl-industry)

Copied to clipboard

Challenge: Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations.
Approach: They propose a framework that bootstraps the planning during ESC and determines the optimal strategy based on long-term returns.
Outcome: The proposed framework outperforms baseline models on ESC datasets and can be used to guide the LLM to response.
Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework for constructing dialogue world models for natural language tasks is currently lacking.
Approach: They propose a framework that can be used to train a dialogue world model.
Outcome: The proposed framework can predict future utterances and user beliefs . it can achieve state-of-the-art performance on emotion classification and sentiment identification .
Shared Path: Unraveling Memorization in Multilingual LLMs through Language Similarities (2025.emnlp-main)

Copied to clipboard

Challenge: Using multilingual models, we find that treating languages in isolation obscures the true patterns of memorization.
Approach: They propose a graph-based correlation metric that incorporates language similarity to analyze cross-lingual memorization.
Outcome: The proposed model incorporates language similarity to analyze cross-lingual memorization in 95 languages.
Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been reported to “leak” Personally Identifiable Information (PII) successful PII reconstruction often interpreted as evidence of memorization.
Approach: They propose a principled revision of memorization evaluation for Large Language Models . they propose PII leakage should be evaluated under low lexical cue conditions .
Outcome: The proposed method is based on a multilingual re-evaluation of PII leakage across 32 languages and multiple memorization paradigms.
Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation (2025.findings-acl)

Copied to clipboard

Challenge: Large language models have advantages over neural machine translation systems, but they suffer from high computational costs and significant latency.
Approach: They propose a scheduling policy that optimizes translation result while ensuring fast speed and as little LLM usage as possible.
Outcome: The proposed model achieves optimal translation performance with less LLM usage on multilingual test sets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations