Papers by Hung-yi Lee
Copied to clipboard
| Challenge: | Researchers have developed a sound codec that can be used as tokenizers for preserving audio data and minimizing data transmission latency. |
| Approach: | They propose to use codec-SUPERB to assess codec models across representative sound applications and signal-level metrics rooted in sound domain knowledge. |
| Outcome: | The proposed codec-SUPERB model is evaluated on selected experimental settings. |
Copied to clipboard
| Challenge: | Emphasis is a crucial component in human communication, which indicates speaker’s intention and implication beyond pure text in dialogue. |
| Approach: | They propose a benchmark dataset with annotated dialogue samples capturing the implications of emphasis. |
| Outcome: | The proposed evaluation pipeline achieves high correlation with human scoring and commercial LLMs perform better than open-source LLM. |
Copied to clipboard
| Challenge: | Despite the rapid development of large language models, the language capabilities of most open-source LLMs are primarily focused on English due to data constraints. |
| Approach: | They propose a chat vector to equip pre-trained language models with instruction following and human value alignment via simple model arithmetic. |
| Outcome: | The proposed method can be extended to include various languages, base models, and chat vectors. |
Copied to clipboard
| Challenge: | Existing evaluation methods for transfer learning are limited in speech research . authors show that pre-trained models transfer well across multiple tasks . |
| Approach: | They propose a benchmark to evaluate pre-trained models by increasing task diversity and difficulty over SUPERB. |
| Outcome: | The proposed benchmark increases task diversity and difficulty over SUPERB-SG. |
Copied to clipboard
| Challenge: | Existing methods for speech recognition suffer from the synthetic-to-real gap . existing methods suffer from this distributional shift due to acoustic mismatches . |
| Approach: | They propose to use task arithmetic to fine-tune an ASR model on synthetic data to mitigate the synthetic-to-real gap. |
| Outcome: | The proposed method shows an improvement of 10.03% over baselines on the SLURP dataset. |
Copied to clipboard
| Challenge: | Currently, most work on improving the fluency and coherence of chatbots is focused on making them more human-like. |
| Approach: | They propose a framework to train chatbots to possess human-like intentions by making them learn from interactive conversation. |
| Outcome: | The proposed framework includes a guiding chatbot and an interlocutor model that plays the role of humans. |
Copied to clipboard
| Challenge: | Using large language models (LLMs) for automatic evaluation has become an important evaluation method in NLP research. |
| Approach: | They use large language models (LLMs) for automatic evaluation to evaluate a sample . they propose several recommendations for integrating LLMs into future classroom evaluations . |
| Outcome: | The proposed model is able to output high scores without meeting the evaluation instructions, the authors note . their model is not able for students to manipulate the model to output specific strings, they say . |
Copied to clipboard
| Challenge: | Full-duplex speech agents are often half-duplice, alternating turns between user and system. |
| Approach: | They propose a streaming framework that integrates with an examiner that enforces staged goals under two pacing setups. |
| Outcome: | The framework reports fluency, multi-turn instruction following, and task-specific competence. |
Copied to clipboard
| Challenge: | Existing factuality metrics cannot evaluate paragraphs with ambiguous entities, authors show . |
| Approach: | They propose a new metric to evaluate the factuality of long-form generations from large language models. |
| Outcome: | The proposed metric can assess the factuality of people biographies with entity ambiguity better than FActScore. |
Copied to clipboard
| Challenge: | Textless Spoken Language Models lag behind text-based Large Language Model (LLM) in semantic coherence and relevance. |
| Approach: | They propose a framework that leverages preference optimization inspired by Reinforcement Learning with Human Feedback to enhance the semantic understanding of SLMs. |
| Outcome: | The proposed framework achieves state-of-the-art performance of SLMs for most benchmarks . it leverages preference optimization inspired by Reinforcement Learning with Human Feedback . |
Copied to clipboard
| Challenge: | Pretraining of pretrained models (LMs) has been extensively studied, but what happened during pretraining is rarely studied. |
| Approach: | They propose to use a totipotent language model to study pretraining behavior . they find that linguistic knowledge and world knowledge do not generally improve as pretraining proceeds, nor do downstream tasks’ performance. |
| Outcome: | The model learns to reconstruct and predict tokens of different parts of speech (POS) in different learning speeds during pretraining. |
Copied to clipboard
| Challenge: | Meta-learning is a new technique that aims to learn better learning algorithms, including better parameter initialization, optimization strategy, network architecture, distance metrics, and beyond. |
| Approach: | This tutorial introduces Meta-learning approaches and the theory behind them, and then reviews the works of applying this technology to NLP problems. |
| Outcome: | This tutorial will introduce Meta-learning approaches and the theory behind them, and then review the works of applying this technology to NLP problems. |
Copied to clipboard
| Challenge: | Existing approaches to predicting future clinical outcomes from EHRs focus on enhancing medical knowledge through distillation or RAG while relying on the model’s internal ability to interpret contextual information. |
| Approach: | They propose a framework for improving clinical outcome prediction from EHR using a sample regeneration mechanism that leverages ground-truth answers as hints to enhance reasoning. |
| Outcome: | Experiments on multiple EHR prediction tasks show significant gains of up to 19.9% over state-of-the-art baselines in terms of F1 score, underscoring ReMedi’s effectiveness in real-world clinical prediction. |
Copied to clipboard
| Challenge: | Parameter-efficient (PE) methods for adapting pre-trained language models to downstream tasks are still lacking in many cases. |
| Approach: | They propose a general PE priming framework to enhance few-shot adaptation and generalization ability of PE methods. |
| Outcome: | The proposed framework reveals that the best priming strategy facilitates adaptation to target tasks. |
Copied to clipboard
| Challenge: | Human evaluation is indispensable for assessing the quality of texts generated by machine learning models or written by humans. |
| Approach: | They propose to use large language models to evaluate unseen texts using the same instructions and samples . they also use LLMs to generate responses to questions that are used to conduct human evaluation . |
| Outcome: | The proposed model can be used to evaluate texts in open-ended story generation and adversarial attacks. |
Copied to clipboard
| Challenge: | Neural audio codecs are optimized for waveform reconstruction rather than autoregressive prediction. |
| Approach: | They propose to augment codec training with language-model-facing objectives while keeping both codec and LLM architectures unchanged. |
| Outcome: | The proposed model improves speech coherence and predictability by preserving the semantic alignment between audio and text representations. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) perform outstandingly across multimodal tasks, but training them with preference data is computationally expensive. |
| Approach: | They propose to merge text-based reward models with LVLMs to create visionlanguage reward models (VLRMs) this approach offers an efficient method for incorporating textual preferences into LVRMs. |
| Outcome: | The proposed model improves over LVLMs’ scoring and text-based RMs, and offers an efficient method for incorporating textual preferences into LVRMs. |
Copied to clipboard
| Challenge: | Existing self-supervised speech encoders contain primarily acoustic rather than semantic information. |
| Approach: | They propose a task-agnostic unsupervised way to incorporate semantic information from large language model (LLM) systems into self-supervised speech encoders without labeled audio transcriptions. |
| Outcome: | The proposed approach improves spoken language understanding (SLU) performance by over 5% on intent classification (IC), with modest gains in named entity resolution (NER) and slot filling (SF), and spoken question answering (SQA) score by over 22%. |
Copied to clipboard
| Challenge: | In synonym substitution attacks, an adversarial sample is constructed by substituting words in the original sentence with their synonyms. |
| Approach: | They examine how synonym substitution attacks replace words in the original sentence and show that there are still unresolved obstacles that make current SSAs generate invalid adversarial samples. |
| Outcome: | The proposed methods generate large fractions of invalid substitution words that are ungrammatical or do not preserve the original sentence’s semantics. |
Copied to clipboard
| Challenge: | Fine-tuning large language models for downstream tasks often leads to catastrophic forgetting, notably degrading the safety of original alignments. |
| Approach: | They propose to merge the weights of pre- and post-fine-tuned models to improve safety while enhancing performance. |
| Outcome: | Experiments across different downstream tasks and models validate the method’s practicality and effectiveness. |
Copied to clipboard
| Challenge: | Speculative decoding (SD) uses an efficient draft model to generate multiple tokens . previous methods depend on simple heuristics to select K or dynamically adjust the window size . |
| Approach: | They propose a framework that allows a draft model to generate multiple tokens . they propose HSDDW, which allows the draft model autonomously decide when to stop generating tokens. |
| Outcome: | The proposed framework outperforms existing state-of-the-art methods on four datasets. |
Copied to clipboard
| Challenge: | Using pre-trained language models, we can apply them to specialized domains such as scientific articles or clinical data. |
| Approach: | They propose to pre-train BERT models on large text corpora and use them to generalize to token sequence classification applications. |
| Outcome: | The models pre-trained on text classification tasks perform better than the models using task-specific knowledge and share non-trivial similarities. |
Copied to clipboard
| Challenge: | Audio-aware large language models (ALLMs) can understand textual and non-textual information in the audio input. |
| Approach: | They use audio-aware large language models (ALLMs) to evaluate the speaking styles of SLMs on two tasks: voice style instruction following and role-playing. |
| Outcome: | The proposed models can understand the textual and non-textual information in the audio input and can be used as a judge to assess the speaking styles of SLMs. |
Copied to clipboard
| Challenge: | Generative spoken language models are often evaluated using global token perplexity, which overlooks fundamental differences between speech and text modalities. |
| Approach: | They propose a variety of likelihood- and generative-based evaluation methods that serve in place of naive global token perplexity. |
| Outcome: | The proposed evaluations more faithfully reflect perceived generation quality, as evidenced by stronger correlations with human-rated mean opinion scores (MOS). |
Copied to clipboard
| Challenge: | Existing studies on RC datasets in English have limited results due to lack of training data. |
| Approach: | They systematically explore zero-shot cross-lingual transfer learning on reading comprehension tasks with pre-trained language representation model. |
| Outcome: | The proposed model performs well on reading comprehension tasks on pre-trained language representation models. |
Copied to clipboard
| Challenge: | Structured generation is used to extract key output information from large language models (LLMs). |
| Approach: | They examine whether constraints on generation space impact LLMs’ abilities, including reasoning and domain knowledge comprehension. |
| Outcome: | The proposed model is based on a few-shot in-context learning and instruction-following capabilities. |
Copied to clipboard
| Challenge: | Unlike textonly large language models (LLMs), SLMs integrate audio encoders and vocoders to support end-to-end speech understanding and generation. |
| Approach: | They evaluate three proprietary and two open-source SLMs and show that none of them can maintain a consistent speaking style when instructed to do so. |
| Outcome: | The proposed models cannot maintain a consistent speaking style after several turns of interaction, but can recall the style instruction when prompted in later turns, but fail to express it. |
Copied to clipboard
| Challenge: | Existing approaches to train transformers with millions of parameters require large storage. |
| Approach: | They propose a transformer-based adapter architecture that adds a token-dependent shift to the hidden output of transformer layers to adapt to downstream tasks with only a vector and a linear layer. |
| Outcome: | The proposed model significantly reduces trainable parameters with minimal performance loss compared to fine-tuned models. |
Copied to clipboard
| Challenge: | Self-supervised representation learning (SSL) uses proxy supervised learning tasks to obtain training data from unlabeled corpora. |
| Approach: | They propose to survey the latest SSL techniques, tools, datasets, and performance achievement in speech processing to scale up current machine learning technologies. |
| Outcome: | The proposed tutorial is highly relevant to the special theme of ACL about language diversity. |
Copied to clipboard
| Challenge: | Current ASR TTA methods focus on non-continual TTA, which limits cross-sample knowledge learning compared to continual TTA. |
| Approach: | They propose a Fast-slow TTA framework that leverages the advantage of continual and non-continual TTA and a Dynamic SUTA method that automatically detects domain shifts and resets the model. |
| Outcome: | The proposed method outperforms non-continual and continual TTA methods while maintaining robustness to domain shifts without requiring domain boundary information. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used to judge code, but their reliability remains poorly understood. |
| Approach: | They propose a benchmark to evaluate Large Language Models as code judges . they find that small reasoning models outperform larger non-reasoning models . |
| Outcome: | The proposed benchmark evaluates LLM-as-a-Judge models across three coding tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) equipped with chain-of-thoughts (CoT) prompting have shown significant multi-step reasoning capabilities in factual content like mathematics, commonsense, and logic. |
| Approach: | They introduce a trope-wise querying approach to assess the abstract reasoning abilities of large language models (LLMs) and uncover their low performance. |
| Outcome: | The proposed approach boosts the F1 score by 11.8 points and also reduces the performance of the large language models (LLMs) it also shows that it can cause hallucinations in narrative content, reducing the performance. |
Copied to clipboard
| Challenge: | In spoken dialogue, even if two current turns are the same sentence, their responses might differ when they are spoken in different styles. |
| Approach: | They propose a language-to-speech dataset that can model linguistic content and speaking styles. |
| Outcome: | The proposed framework outperforms text-only baselines and prior speech LLMs methods. |
Copied to clipboard
| Challenge: | a new study examines the proactive ability of large language models to seek user support . without external feedback, many LLMs struggle to recognize their need for user support. |
| Approach: | They propose metrics to evaluate the trade-off between performance improvements and user burden . they also investigate whether LLMs can determine when to request user support . |
| Outcome: | The proposed metrics show that without external feedback, many LLMs struggle to recognize their need for user support. |
Copied to clipboard
| Challenge: | Recent advances in large audio-language models (LALMs) have expanded their impact beyond natural language processing (NLP) to multimodal domains. |
| Approach: | They propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. |
| Outcome: | The proposed taxonomy categorizes LALM evaluations into four dimensions based on their objectives and highlights challenges in this field. |
Copied to clipboard
| Challenge: | ASR models can be used to correct accent-specific errors without ground truth . pseudo-labels inherit the teacher model's systematic biases, authors say . |
| Approach: | They propose a parameter-space correction technique that captures pseudo-label biases . they propose achieving up to 35% relative WER reduction on a pseudo-labeled target model . |
| Outcome: | The proposed model achieves 35% relative WER reduction on ten African accents with the Whisper tiny model. |
Copied to clipboard
| Challenge: | Meta-learning is an emerging field in machine learning, but there is no systematic survey of these approaches in NLP. |
| Approach: | They propose to introduce meta-learning and the common approaches and summarize their work and review their work in the NLP community. |
| Outcome: | The proposed methods improve performance in many NLP tasks but are limited to domains, languages, countries, or styles. |
Copied to clipboard
| Challenge: | Using connectionist temporal classification (CTC) for speech-to-text translation is counter-intuitive due to its monotonicity assumption. |
| Approach: | They propose to build a non-autoregressive speech-to-text translation model using connectionist temporal classification (CTC) their work shows transformer encoders can change the word order and points out the future research direction that needs to be explored more on non-Autoregressives speech translation. |
| Outcome: | The proposed model improves translation performance by using transformer encoders. |
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) tasks have received little attention and resources compared to lower-level tasks like speech and speaker recognition. |
| Approach: | They propose annotated SLU benchmark tasks based on freely available speech data to complement existing benchmarks and address gaps in the evaluation landscape. |
| Outcome: | The proposed benchmarks complement existing benchmarks and address gaps in the evaluation landscape. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning LLMs use cross-entropy (CE) loss . existing methods neglect the numeric nature of score prediction . |
| Approach: | They propose a method that fine-tunes large language models (LLMs) for automated text evaluation, assigning a score to the input based on scoring rubrics. |
| Outcome: | The proposed model outperforms existing methods in four LLM-as-a-judge datasets and two LLMs. |
Copied to clipboard
| Challenge: | XDBERT (cross-modal distilled BERT) outperforms pretrained-BERT in general language understanding evaluation (GLUE), situations with adversarial generations (SWAG) benchmarks, and readability benchmarks. |
| Approach: | They propose to distill visual information from pretrained multimodal transformers to pretrained language encoders to cater to the language-heavy characteristics of NLU. |
| Outcome: | The proposed framework outperforms pretrained-BERT in general language understanding evaluation (GLUE), situations with adversarial generations (SWAG), and readability benchmarks. |
Copied to clipboard
| Challenge: | Multi-LLM systems enhance creativity of large language models by simulating human collective intelligence but suffer from significant drawbacks, such as high computational costs and inference latency. |
| Approach: | They propose a training-free framework that captures the benefits of multi-LLM collaboration by extracting and blending multiple distinct persona vectors directly in the model’s activation space. |
| Outcome: | The proposed framework surpasses model prompting and traditional multi-LLM approaches while significantly reducing inference time and computational costs. |
Copied to clipboard
| Challenge: | Spoken language understanding evaluation (SLUE) benchmarks are used to benchmark complex spoken language understanding tasks on natural speech. |
| Approach: | They propose a set of benchmark tasks to evaluate spoken language understanding on natural speech . they use pre-trained speech foundation models to evaluate the utility of different SFMs . |
| Outcome: | The proposed framework outperforms pre-trained speech foundation models on natural speech . the proposed framework also outperformed self-supervised SFMs on the sequence generation tasks . |
Copied to clipboard
| Challenge: | Existing studies explore the use of large language models to evaluate text quality, but they differ in some details of the evaluation process. |
| Approach: | They propose to use large language models to evaluate text quality by giving LLMs instructions to evaluate samples by giving them a rating. |
| Outcome: | The auto Chain-of-Thought (CoT) used in G-Eval does not always make it more aligned with human ratings. |
Copied to clipboard
| Challenge: | Existing large language models and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. |
| Approach: | They propose a general inference framework that enables SLMs to generate unspoken chain-of-thought reasoning while listening to user input. |
| Outcome: | The proposed framework enhances real-time user–SLM interaction in two scenarios. |
Copied to clipboard
| Challenge: | Existing studies show that multitask learning improves speech translation performance by utilizing word embedding as the intermediate. |
| Approach: | They propose to use word embedding as an intermediate to improve multitask ST models by utilizing word embeds as input. |
| Outcome: | The proposed model outperforms existing models with sufficient training data but is still lacking in the low-resource scenario. |
Copied to clipboard
| Challenge: | Pre-trained language models are language models that are pre-taught on large-scaled corpora in a self-supervised fashion. |
| Approach: | This tutorial provides a broad and comprehensive introduction to pre-trained language models . it focuses on emerging methods that enable PLMs to perform diverse downstream tasks . |
| Outcome: | This tutorial focuses on the benefits of pre-trained language models and how to use them in NLP tasks. |
Copied to clipboard
| Challenge: | Mamba-based SSL models are promising for long-sequence modeling, speech unit extraction, and speech self-supervised learning. |
| Approach: | They propose to use Mamba-based HuBERT models as an alternative to Transformer-based SSL architectures. |
| Outcome: | The proposed models outperform Transformer-based models in language modeling tasks while showing superior performance on streaming ASR. |
Copied to clipboard
| Challenge: | Modern large language models (LLMs) showcase impressive capabilities across various tasks with aligning their behavior with human preferences. |
| Approach: | They propose a framework that integrates domain-specific knowledge into a general reward model by model merging. |
| Outcome: | The proposed framework improves performance across different benchmarks and provides detailed analysis showing the effects of model merging. |
Copied to clipboard
| Challenge: | Recent advances in controllable expressive speech synthesis have allowed for the generation of speech with specific styles guided by textual descriptions, known as style prompts. |
| Approach: | They examine whether models exhibit gender bias when interpreting occupation-related prompts. |
| Outcome: | The proposed models exhibit gender bias for certain occupations and different sizes show varying degrees of this bias across occupations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can solve problems step-by-step, but it is unclear whether they know when to use CoT and whether they are always necessary. |
| Approach: | They propose to use LLMs to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero. |
| Outcome: | The proposed model generates redundant calculations and reasoning on a manually constructed math QA dataset, but it is unclear whether it is necessary to use CoT reasoning. |
Copied to clipboard
| Challenge: | Large language model (LLM)-driven multi-agent systems (MAS) are transforming how humans and AIs collaboratively generate ideas and artifacts. |
| Approach: | They present a taxonomy of agent proactivity and persona design and an overview of generation techniques. |
| Outcome: | The proposed framework and roadmap offers a roadmap for advancing the development, evaluation, and standardization of creative MAS. |
Copied to clipboard
| Challenge: | Existing work has not shown that knowledge-grounded models can zero-shot adapt to updated, unseen knowledge graphs. |
| Approach: | They propose a task to apply dynamic knowledge graphs to neural conversation models . they propose 'dyKgChat' that selects an output from two networks at each time step . |
| Outcome: | The proposed model outperforms existing knowledge-grounded conversation models in evaluation metrics. |