Papers by Yongqi Wang

26 papers
Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages.
Approach: They propose to use large language models as a general-purpose interface across multiple tasks and languages.
Outcome: The proposed model performs better on 200K hours of 6-language data for voice generation applications.
Personalized Large Language Model Assistant with Evolving Conditional Memory (2025.coling-main)

Copied to clipboard

Challenge: With the rapid development of large language models, personalized large language model assistants like ChatGPT are limited in personalized services.
Approach: They propose a plug-and-play framework that could facilitate personalized large language model assistants with evolving conditional memory.
Outcome: The proposed framework can preserve the knowledge and experience from the history dialogue with the user, which can be applied to future tailored responses that better align with the users' preferences.
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks assess LLM performance in single-course settings and lack systematic evaluation in multi-course scenarios, where a patient’s condition evolves over time.
Approach: They propose to use large language models to assess their performance in multi-course clinical decision-making scenarios where a patient’s condition evolves over time.
Outcome: The proposed model includes 1,275 Chinese and 5,804 English samples across four stages from admission to discharge.
Understanding Conflicts in Multi-Objective Alignment through Reward Consistency (2026.findings-acl)

Copied to clipboard

Challenge: Existing training pipelines still face alignment conflicts where optimizing for one objective degrades performance on others.
Approach: They propose a reward-based criterion that approximates alignment conflicts via reward models.
Outcome: The proposed framework improves harmlessness and helpfulness scores by 23.07% over the vanilla dataset.
BubbleRAG: Interactive Cognitive Offloading with Thought Bubble in Retrieval-Augmented Generation (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) extends the capabilities of large language models (LLMs) by providing access to external knowledge.
Approach: They propose a framework that emulates human interactive reading through annotation and re-reading by integrating a thought bubble module that offloads internal cognition into external bookmark tokens, which are then annotated back into the context.
Outcome: The proposed framework offloads internal cognition into external bookmark tokens, which are then annotated back into the context.
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer (2024.acl-srw)

Copied to clipboard

Challenge: Existing methods to translate spoken utterances from one language to another are unable to preserve speaker timbre of source speech.
Approach: They propose a pipeline with style-transfer capability on the basis of self-supervised speech representations and codec units.
Outcome: The proposed model achieves zero-shot cross-lingual style transfer on previously unseen source languages.
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment (2025.naacl-long)

Copied to clipboard

Challenge: Reinforcement Learning with Human Feedback (RLHF) is the key to the success of large language models (LLMs) in recent years.
Approach: They propose a method to balance the number of prompts and responses to improve knowledge breadth and knowledge depth by introducing gradient-based clustering to estimate the knowledge informativeness and usefulness of each augmented sample.
Outcome: The proposed method outperforms baseline methods while maintaining training efficiency.
MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods overlook the distinction between factoid and non-factoidic questions.
Approach: They propose a method that distinguishes open-ended questions and ranks candidate answers . they propose QA requires longer answer statements and nuanced reasoning processes .
Outcome: The proposed method better aligns with human annotations and offers more interpretable results.
Can LLMs Learn from Previous Mistakes? Investigating LLMs’ Errors to Boost for Reasoning (2024.acl-long)

Copied to clipboard

Challenge: Recent studies have shown the benefits to LLMs from fine-tuning golden-standard Chain-of-Thought rationales or using them as correct examples in few-shot prompting.
Approach: They propose a new benchmark to test the effectiveness of large language models by leveraging errors to enhance reasoning capabilities.
Outcome: The proposed methods can be used to fine-tune models in correct and incorrect domains, rather than tuning models to learn ground truth in traditional methods.
LCDS: A Logic-Controlled Discharge Summary Generation System Supporting Source Attribution and Expert Review (2025.acl-demo)

Copied to clipboard

Challenge: Large language models (LLMs) are capable of generating inaccurate discharge summary content or fabricating information without valid sources.
Approach: They propose a tool for empowering LLMs with Logic-Controlled Discharge Summary generation.
Outcome: The proposed tool identifies the writing logic of discharge summaries and integrates it with EMRs to generate silver discharge summararies.
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in generative language models have demonstrated their ability to memorize knowledge from documents and recall knowledge to respond to user queries effectively.
Approach: They propose to enable multimodal large language models to memorize and recall images within their parameters.
Outcome: The proposed model performs well even with large-scale image candidate sets.
MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based Retrieval (2025.findings-naacl)

Copied to clipboard

Challenge: Embedding-based retrieval (EBR) is a mainstream approach in information retrieval.
Approach: They propose an enriched benchmark to evaluate retrieval capabilities of embedding models . they use four levels of granularity and six types of medical texts to prompt instruction-fine-tuned embeddable models.
Outcome: The proposed benchmark evaluates the retrieval capabilities of embedding models with multi-granularity and multi-data types.
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a high inference latency stemming from autoregressive decoding.
Approach: They propose a novel decoding paradigm that drafts multiple tokens and verifies them in parallel . they aim to provide a catalyst for further research on Speculative Decoding .
Outcome: The proposed method drafts multiple tokens and verifies them in parallel . it can be used to accelerate inference in large language models.
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt (2024.naacl-long)

Copied to clipboard

Challenge: Recent singing-voice-synthesis methods lack ability to control style attributes of synthesized singing.
Approach: They propose a singing-voice-synthesis method that enables attribute controlling on singer gender, vocal range and volume with natural language.
Outcome: The proposed method achieves favorable control ability and audio quality.
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation (2026.acl-long)

Copied to clipboard

Challenge: Situated conversational recommendation (SCR) uses visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recommendations.
Approach: They propose a framework that integrates scene transition estimation and Bayesian inverse inference to provide contextually appropriate recommendations.
Outcome: The proposed framework achieves superiority over baselines on two representative benchmarks on dynamic scene transitions and implicit user intents.
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation (2023.findings-emnlp)

Copied to clipboard

Challenge: toxicity detection has been largely based on social media content, leaving the unique challenges inherent to real-world user-AI interactions insufficiently explored.
Approach: They propose a benchmark to detect toxicity in real-world user-AI conversations . they compare existing models with social media content to find toxicity .
Outcome: The proposed benchmark reveals that existing models fail to recognize toxicity in real-world user-AI conversations.
TokenSkip: Controllable Chain-of-Thought Compression in LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Chain-of-Thought (CoT) has been proven effective in enhancing the reasoning capabilities of large language models (LLMs).
Approach: They propose a chain-of-thought (CoT) prompting approach that enables LLMs to selectively skip less important tokens, allowing for controllable CoT compression.
Outcome: Experiments show that TokenSkip reduces CoT token usage while preserving strong reasoning performance.
TInR: Exploring Tool-Internalized Reasoning in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods rely on external tool documentation during reasoning, leading to tool mastery difficulty, tool size constraints, and inference inefficiency.
Approach: They propose a tool-internalized reasoning framework for unified reasoning and tool usage that integrates external tools into Large Language Models (LLMs) to address these issues, they propose 'tool-internet-based' reasoning.
Outcome: The proposed method achieves superior performance across in-domain and out-of-domain settings, highlighting its effectiveness and efficiency.
Multiview Identifiers Enhanced Generative Retrieval (2023.acl-long)

Copied to clipboard

Challenge: Current approaches use a numeric ID or text piece as the identifier, but these identifieres cannot cover a passage’s content well.
Approach: They propose a new type of identifier that is generated based on the content of a passage and could integrate contextualized information that text pieces lack.
Outcome: The proposed approach performs the best in generative retrieval on three public datasets.
Parallel Test-Time Scaling for Latent Reasoning Models (2026.acl-long)

Copied to clipboard

Challenge: Parallel test-time scaling is a pivotal approach for enhancing large language models.
Approach: They propose two uncertainty-inspired stochastic strategies for parallel test-time scaling for latent reasoning models and a Latent Reward Model for aggregation.
Outcome: The proposed model scales well with compute and enables effective trajectory selection.
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on speech-to-singing voice conversion (STS) are limited by the scarcity of paired speech-song data and the suboptimal quality of outputs.
Approach: They propose a self-supervised singing voice pre-training model that transforms a speech-to-singing voice into a paired singing voice.
Outcome: The proposed model improves both STS and singing voice synthesis tasks by combining spoken language and a self-supervised singing voice pre-training model.
Self-Sum: Teaching an Agent to Decide Itself When and What to Summarize (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for summarizing long-horizon agents rely on fixed, rule-based summarization strategies.
Approach: They propose a framework that empowers agents to autonomously decide when and what to summarize by modeling it as an internal cognitive action unified with environmental actions.
Outcome: The proposed framework outperforms no-summarization and rule-based training methods on long-horizon benchmarks and shows strong generalization gains.
Robust Singing Voice Transcription Serves Synthesis (2024.acl-long)

Copied to clipboard

Challenge: Current AST methods struggle with accuracy and robustness when used for practical annotation.
Approach: They propose a model that converts singing recordings into note sequences for automatic annotation of singing datasets.
Outcome: The proposed model outperforms baseline models on enlarged, automatically annotated datasets.
Distillation Enhanced Generative Retrieval (2024.findings-acl)

Copied to clipboard

Challenge: Generative retrieval is a promising new paradigm in text retrieval that generates identifier strings of relevant passages as the retrieval target.
Approach: They propose a framework that leverages generative language models to enhance generative retrieval by distillation.
Outcome: The proposed framework achieves state-of-the-art performance among the generative retrieval methods.
Text-to-Song: Towards Controllable Music Generation Incorporating Vocal and Accompaniment (2024.acl-long)

Copied to clipboard

Challenge: Existing studies focus on singing voice synthesis and music generation independently.
Approach: They propose a novel task called Text-to-Song synthesis which incorporates both vocal and accompaniment generation.
Outcome: The proposed method can synthesize songs with comparable quality and style consistency.
Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Early approaches focus on text-based reasoning, but they often follow a single task-specific reasoning pattern.
Approach: They propose a generative multimodal reasoning paradigm that unifies diverse reasoning skills by generating intermediate images during the reasoning process.
Outcome: The proposed model unifies diverse multimodal reasoning skills by generating intermediate images during the reasoning process.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations