Papers by Yong Zhao
Copied to clipboard
| Challenge: | In-Context Learning (ICL) is an essential emergent ability of Large Language Models (LLMs). |
| Approach: | They introduce CoT to exemplars of ICL to enhance the reasoning capability . however, it remains unclear whether CoT exemplar is still beneficial for recent, stronger models in such tasks. |
| Outcome: | The enhanced exemplars fail to improve the model’s reasoning performance, despite being constructed using answers from advanced models such as Qwen2.5-Max and DeepSeek-R1. |
Copied to clipboard
| Challenge: | Prompt-based learning paradigms are vulnerable to backdoor attacks, requiring false activations and false data augmentation. |
| Approach: | They propose a method that uses triggers to create stronger shortcuts by leveraging activation values and data selection strategies to create the shortcuts. |
| Outcome: | The proposed method is based on the concept that a backdoor acts as a shortcut and can achieve high effectiveness and stealthiness at low poisoning rates. |
Copied to clipboard
| Challenge: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Approach: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Outcome: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
Copied to clipboard
| Challenge: | Existing machine learning approaches do not have above semantic knowledge to address complicated MRC questions. |
| Approach: | They propose a frame-based Sentence Representation method which integrates frame semantic knowledge to facilitate sentence modelling. |
| Outcome: | The proposed method performs better than state-of-the-art methods on machine reading comprehension task. |
Copied to clipboard
| Challenge: | Embodied Question Answering (EQA) tasks are primarily focused on indoor environments, leaving the complexities of urban settings unexplored. |
| Approach: | They propose a task where an embodied agent answers open-vocabulary questions in dynamic city spaces. |
| Outcome: | The proposed agent achieves 60.7% of human-level answering accuracy compared to baselines . the proposed agent outperforms existing agents in open-ended city spaces . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. |
| Approach: | They propose to use a model averaging technique to find the most powerful alignment-forging Pareto front among RLHF algorithms. |
| Outcome: | The proposed method achieves the strongest alignment-forging Pareto front among competing methods. |
Copied to clipboard
| Challenge: | Automatic speech recognition systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0. |
| Approach: | They propose to use Mandarin speech datasets to analyze pronunciation and tone of children aged 3 to 5 and evaluate their models on speaker verification (SV) They find that the datasets are more robust than those used by adult speech recognition systems and are open-source and available for all academic purposes. |
| Outcome: | The proposed dataset includes 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. |
Copied to clipboard
| Challenge: | Existing methods for document summarization focus on one type of relation, neglecting the simultaneous effective modeling of both relations. |
| Approach: | They propose a graph neural network-based approach to local and global document summarization using hierarchical discourses. |
| Outcome: | The proposed approach improves on two benchmark datasets and shows that hierarchical structures are important for document summarization. |
Copied to clipboard
| Challenge: | Autoregressive (AR) large audio language models are expensive in data and computation . prior work shows diffusion-based LALMs can improve audio understanding under matched settings . |
| Approach: | They propose a diffusion-based LALM that upgrades the speech encoder and employs dual semantic and acoustic adapters. |
| Outcome: | a new model improves over existing autoregressive large language models and is competitive to strong AR models . the proposed model can make use of limited training data and improve inference efficiency . a recent study shows that diffusion-based models can improve audio understanding . |
Copied to clipboard
| Challenge: | Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. |
| Approach: | They propose a framework that bootstraps the planning during ESC and determines the optimal strategy based on long-term returns. |
| Outcome: | The proposed framework outperforms baseline models on ESC datasets and can be used to guide the LLM to response. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have driven the rise of agentic workflows . yet, how can we attribute performance gains to individual upgrades and their interactions? |
| Approach: | They propose a game-theoretic framework that models component upgrades as players and evaluates component coalitions to compute Shapley values. |
| Outcome: | The proposed framework provides interaction-aware attribution and recommendation for model allocation under a fixed workflow structure. |
Copied to clipboard
| Challenge: | Large language models (LLMs) store vast amount of knowledge in their parameters, but they still have limitations in the memorization and utilization of certain knowledge. |
| Approach: | They propose a comprehensive definition of the LLM knowledge boundary and introduce a formalized taxonomy categorizing knowledge into four distinct types. |
| Outcome: | The proposed definition of the LLM knowledge boundary and taxonomy categorizes knowledge into four distinct types . aims to offer a comprehensive overview, facilitate access to key issues, and inspire further advancements in LLM research. |
Copied to clipboard
| Challenge: | Existing methods for memory management struggle to capture fine-grained semantic relations between queries and documents. |
| Approach: | They propose a framework for reasoning and agentic search that grows fine-grained memory fragments from seed tokens from queries, then retraces and deep refines the memory via a contribution function. |
| Outcome: | Experiments on eight benchmark datasets show that MemSearch-o1 significantly mitigates memory dilution and more effectively activates reasoning potential of diverse LLMs. |
Copied to clipboard
| Challenge: | Existing approaches to extract rich correlations between entities and relations are not fully exploited by existing methods. |
| Approach: | They propose to unify entities and relations by jointly encoding them within a concatenated natural language sequence and unify the modeling of interactions with a proposed Interaction Map. |
| Outcome: | The proposed method is more efficient and efficient than existing methods and can be scaled up to 2021. |
Copied to clipboard
| Challenge: | Automatic speech recognition (ASR) for children remains challenging due to developmental variability and the scarcity of high-quality corpora. |
| Approach: | They propose a large-scale Chinese child speech corpus that contains 112.5 hours of speech from 498 children and 500 caregivers. |
| Outcome: | The proposed model improves in-domain and cross-domain performance on children's speech. |
Copied to clipboard
| Challenge: | a framework for constructing dialogue world models for natural language tasks is currently lacking. |
| Approach: | They propose a framework that can be used to train a dialogue world model. |
| Outcome: | The proposed framework can predict future utterances and user beliefs . it can achieve state-of-the-art performance on emotion classification and sentiment identification . |
Copied to clipboard
| Challenge: | Existing approaches to enabling LLM web search proficiency struggle with data production in open-search domains, while supervised fine-tuning struggles with data utilization efficiency. |
| Approach: | They propose an iterative self-evolution framework that combines SFT and RL to enhance agentic web search capabilities without external human-annotated reasoning data. |
| Outcome: | EvolveSearch achieves 4.7% improvement over current state-of-the-art in seven benchmarks . supervised fine-tuning struggles with data production in open-search domains compared with RL . |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on scalar rewards to capture user preferences. |
| Approach: | They propose a framework that integrates multi-objective reward modeling to represent diverse preference profiles. |
| Outcome: | The proposed method improves performance across reward objectives and targets. |
Copied to clipboard
| Challenge: | Mobile GUI agents show promise in automating tasks but face significant generalization challenges in long-tail scenarios. |
| Approach: | They propose a benchmark framework for mobile GUI agents that measures the performance of GUI agents by analyzing their performance. |
| Outcome: | The LearnGUI benchmark outperforms existing methods in offline and online evaluations and demonstrates consistent gains across model architectures. |
Copied to clipboard
| Challenge: | In-context learning (ICL) is a common practice to enhance LLM performance on domain-specific tasks. |
| Approach: | They propose a method that leverages large language models to enhance query-ad relevance labeling . they identify and provide superior demonstrations for ICL, thereby improving labeling performance . |
| Outcome: | The proposed method improves query-ad relevance labeling performance by providing demonstrations. |
Copied to clipboard
| Challenge: | Aligned Large Language Models exhibit remarkable versatility, capable of handling diverse real-world tasks. |
| Approach: | They propose a coarse to fine framework to fine-tune aligned Large Language Models to achieve a balance between speciality and versatility. |
| Outcome: | The proposed framework outperforms baseline methods across diverse tasks and model scales. |
Copied to clipboard
| Challenge: | Existing approaches to MCIT address Catastrophic Forgetting and Knowledge Transfer (KT) but using a fixed number of shared LoRA blocks across tasks can lead to knowledge interference. |
| Approach: | They propose a framework that uses a fixed number of shared LoRA blocks to reduce knowledge interference. |
| Outcome: | The proposed framework outperforms existing approaches on the latest MCIT benchmark. |
Copied to clipboard
| Challenge: | Existing keyphrase extraction methods struggle with document and candidate length discrepancies or fail to fully utilize the pre-trained language model without further fine-tuning. |
| Approach: | They propose an unsupervised keyphrase extraction approach that uses a pre-trained language model to rank candidates based on document embeddings. |
| Outcome: | The proposed approach outperforms the existing keyphrase extraction approach on six benchmarks. |
Copied to clipboard
| Challenge: | Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various tasks, but is limited in multi-turn social interactions. |
| Approach: | They propose a method which dynamically selects key segments within interactions to optimize multi-turn agent behavior. |
| Outcome: | The proposed methods outperform existing methods and proprietary LLMs on the SOTOPIA benchmark and show that they can improve social intelligence. |
Copied to clipboard
| Challenge: | Earlier studies of instruction tuning on Large Language Models focus on creating large, varied, and high-quality datasets with responses curated by human experts. |
| Approach: | They propose to use a smaller and weaker model to fine tune a larger and stronger model . they find it can largely speed up the data filtering and improve performance . |
| Outcome: | The proposed model can filter instruction data faster and better on benchmarks. |
Copied to clipboard
| Challenge: | Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored. |
| Approach: | They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain. |
| Outcome: | MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows. |
Copied to clipboard
| Challenge: | Existing methods to apply large language models to zero-shot next location prediction tasks are limited due to their limited computational power. |
| Approach: | They propose a systematic agentic prediction framework to achieve generalized next location prediction. |
| Outcome: | The proposed framework surpasses the leading baseline by 3.33% to 8.57% across 8 out of 12 metrics. |
Copied to clipboard
| Challenge: | Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. |
| Approach: | They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans. |
| Outcome: | The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation. |
Copied to clipboard
| Challenge: | Unsupervised parsing learns a syntactic parser from training sentences without parse tree annotations. |
| Approach: | This tutorial will introduce what unsupervised parsing does and how it can be useful for and beyond syntactic parse. |
| Outcome: | This paper will provide an overview of major approaches to unsupervised parsing and analyze their strengths and weaknesses. |
Copied to clipboard
| Challenge: | Recent years have witnessed a paradigm shift in natural language processing, driven by large language models such as GPT-3, PaLM, and Llama. |
| Approach: | They propose a strategy for role-play prompting and assess its performance under the zero-shot setting. |
| Outcome: | The proposed method outperforms the standard zero-shot prompting approach across 12 reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing studies investigate ways to refuse to answer unknown questions . Large Language Models (LLMs) display a significant level of overconfidence when answering questions that they are aware of. |
| Approach: | They propose a self-alignment method to utilize Large Language Models to enhance its response-ability to unknown questions. |
| Outcome: | The proposed method is superior to baseline methods on four types of unknown questions. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models. |
| Approach: | They propose a "generate-then-read" pipeline to replace retrieval stage with generation from the LLM itself. |
| Outcome: | The proposed framework outperforms single models in the base and chat versions and addresses safety and helpfulness post-adaptation challenges. |
Copied to clipboard
| Challenge: | Semi-autoregressive (Semi-AR) decoding suffers from inherent block constraints . naive lookahead decoding is unreliable, token stability closely correlates with convergence trend, and historical information is isolated. |
| Approach: | They propose a training-free, plug-and-play dynamic decoding strategy that monitors the stability of tokens in real time through dynamic anchors. |
| Outcome: | The proposed approach reduces decoding steps by 80% while improving performance by 3.67% on the BBH benchmark. |
Copied to clipboard
| Challenge: | Existing ground VLN agents struggle in aerial VLLN due to the lack of predefined navigation graphs and the exponentially expanding action space in long-horizon exploration. |
| Approach: | They propose a large language model-empowered aerial VLN agent that decomposes the long-horizon task into sub-goals with different semantic levels. |
| Outcome: | The proposed method achieves state-of-the-art performance with significant improvement in continuous city environments. |
Copied to clipboard
| Challenge: | Aspect Sentiment Triple Extraction (ASTE) is an advanced natural language processing task. |
| Approach: | They propose a Dual Encoder: Exploiting the potential of Syntactic and Semantic model which maximizes syntactical and semantic relationships among words. |
| Outcome: | The proposed model surpasses the current state-of-the-art on public benchmarks and shows that it is highly efficient. |
Copied to clipboard
| Challenge: | a Chinese model with whole word masking has no subword because each token is an atomic character. |
| Approach: | They propose to use whole word masking to mask all subwords corresponding to a word at once . they ask models to revise or insert tokens in a masked language modeling manner . |
| Outcome: | The proposed model performs better when one character is inserted or replaced . the model trained with standard character-level masking performs best when one token is masked . |
Copied to clipboard
| Challenge: | Recent studies have shown that unsupervised bilingual lexicon induction is even on par with supervised methods. |
| Approach: | They propose a relaxed matching procedure to find a more precise matching between two languages by aligning source and target embedding space bidirectionally. |
| Outcome: | The proposed method significantly outperforms previous unsupervised methods on standard benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to predict sentiments on social media are limited and do not consider reciprocal influences among social media users. |
| Approach: | They propose a multi-perspective role-playing framework to simulate human response processes to extract sentiment-related features from social media messages. |
| Outcome: | The proposed model improves sentiment forecasting at microscopic and macroscopic levels. |
Copied to clipboard
| Challenge: | Existing methods to finetun large language models (LLMs) only update a small number of trainable parameters, or attempt to reduce the memory footprint during the training phase of the finetune process. |
| Approach: | They propose quantized side tuing (QST) which quantizes an LLM’s model weights into 4-bit to reduce the memory footprint of the original weights. |
| Outcome: | The proposed method reduces the memory footprint of the model weights, optimizer states, and intermediate activations while reducing the memory requirements. |
Copied to clipboard
| Challenge: | Large Language Model (LLM)-based optimization has shown promise for autonomous problem solving, but most approaches cast LLMs as passive constraint checkers rather than proactive strategy designers. |
| Approach: | They propose an end-to-end Automated Constraint Optimization method that tightly couples operations-research principles of constraint relaxation with LLM reasoning. |
| Outcome: | Extensive experiments on three challenging COP benchmarks validate AutoCO’s consistent effectiveness and superior performance, especially in hard regimes where current methods degrade. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have sparked concerns about their safety. |
| Approach: | They propose a method to identify safety-related information in the model parameter space . they propose to use a few adversarially chosen examples to fine-tune LLMs . |
| Outcome: | The proposed method can break safety alignment in multilingual LLMs using a few examples . it also shows that the proposed method jailbreaks LLM models adapted to new languages . |
Copied to clipboard
| Challenge: | Existing systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis. |
| Approach: | They propose an emotion-planning framework that determines the emotion prior to the textual generation, grounding the downstream emotional TTS in a streaming manner. |
| Outcome: | The proposed framework outperforms baselines on DailyDialog, EmoryNLP, IMEOCAP, and MELD on emotional alignment, contextual coherence, and expressive fluency. |
Copied to clipboard
| Challenge: | Existing knowledge evolution benchmarks are static and fail to capture the evolving nature of LLMs and knowledge. |
| Approach: | They propose an evolving dataset that categorizes information into stable, evolved, and uncharted states. |
| Outcome: | The proposed dataset is auto-updatable and enables evaluation of continuously changing knowledge and newly released LLMs. |
Copied to clipboard
| Challenge: | Existing semisupervised methods do not fully utilize the knowledge hidden in annotated and nonannotated data, which hinders further improvement of their performance. |
| Approach: | They propose a semi-supervised BLI framework to encourage interaction between supervised signal and unsupervised alignment. |
| Outcome: | The proposed framework can incorporate any supervised and unsupervised BLI methods based on optimal transport and bi-directional lexicon update. |
Copied to clipboard
| Challenge: | a new approach to timeline summarization is proposed for open-domain news content . large language models (LLMs) can be used to extract and organize news events from multiple documents . |
| Approach: | They propose a method to integrate Large Language Models into news timeline summarization by iterating on how events are linked and posing new questions. |
| Outcome: | The proposed system is able to generate and refresh chronological summaries based on documents retrieved in each round. |
Copied to clipboard
| Challenge: | Large Reasoning Models suffer from high inference latency due to lengthy reasoning chains. |
| Approach: | They propose a collaborative framework that combines large and small models for effective reasoning. |
| Outcome: | The proposed framework reduces inference latency by 1.7-4.1 while maintaining comparable accuracy to standard large model inference. |
Copied to clipboard
| Challenge: | Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks. |
| Approach: | They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset. |
| Outcome: | The proposed model performs well across tasks and languages. |