Papers by Michael Li
Copied to clipboard
| Challenge: | Structured representations for task-oriented assistant systems are limited due to the limitations of the representation. |
| Approach: | They propose a semantic representation for task-oriented conversational systems that can represent co-reference and context carryover. |
| Outcome: | The proposed model improves the best results on ATIS, SNIPS, TOP and DSTC2 by up to 5 points for slot-carryover. |
Copied to clipboard
| Challenge: | Recent advances in text pretraining and finetuning have improved multitasking applications significantly. |
| Approach: | They propose a minimalistic LNA finetuning approach to build multilingual speech-to-text translation using a pretrained speech encoder and text decoder. |
| Outcome: | The proposed approach surpasses the cascaded ST benchmark for 36 translation directions on the large-scale multilingual ST benchmark CoVoST 2. |
Copied to clipboard
| Challenge: | Existing learning metrics are limited to tasks where large human ratings are available. |
| Approach: | They propose a model-based natural language generation (NLG) evaluation metric that is highly correlated with human judgements without requiring human annotation. |
| Outcome: | The proposed metric outperforms all prior unsupervised metrics on multiple NLG tasks including translation, image captioning, and WebNLG text generation. |
Copied to clipboard
| Challenge: | Recent research has focused on examining Large Language Models’ characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics. |
| Approach: | They propose to examine the reliability of personality tests to LLMs by using psychological scales. |
| Outcome: | The proposed model can represent diverse personalities with specific prompt instructions. |
Copied to clipboard
| Challenge: | Language identification (LID) is a fundamental step in curating multilingual corpora. |
| Approach: | They introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. |
| Outcome: | The proposed benchmark covers 109 languages and shows that existing evaluations overestimate accuracy for many languages in the web domain. |
Copied to clipboard
| Challenge: | a large-scale empirical study compares natural web data, diverse synthetic types, and mixtures of natural and synthetic data. |
| Approach: | They conduct a large-scale empirical study on large-volume LLMs using a unified protocol and scaling laws. |
| Outcome: | The proposed method is faster than pre-training on natural web data, the authors show . their results are consistent with previous studies on rephrased text and textbooks . |
Copied to clipboard
| Challenge: | InstructCoder is the first instruction-tuning dataset designed to adapt LLMs for general-purpose code editing. |
| Approach: | They propose to use Large Language Models to edit code based on user instructions . they use a dataset to adapt LLMs to general-purpose code editing . |
| Outcome: | The proposed model can significantly improve code editing performance compared to proprietary models . the proposed model is based on a human-written execution-based benchmark . |
Copied to clipboard
| Challenge: | Existing studies focus on text modeling, ignoring the rich features embedded in the matching images. |
| Approach: | They propose a novel multi-modal multi-head attention model to capture cross-media interactions and image wordings to bridge the two modalities. |
| Outcome: | The proposed model outperforms the current state of the art based on text modeling and image matching . |
Copied to clipboard
| Challenge: | Sentence function is an important linguistic feature indicating the communicative purpose of a sentence in a conversation. |
| Approach: | They propose a structured meta-learning approach for dialogue generation on infrequent sentence functions. |
| Outcome: | The proposed approach improves informativeness and relevance of dialogue generation on infrequent sentence functions while preserving knowledge generalization for similar sentence functions. |
Copied to clipboard
| Challenge: | Existing methods for event prediction are incomplete and noisy. |
| Approach: | They propose to use news-related event schemas to extract newsworthy events . they build a demo website and include a video demonstrating the framework . |
| Outcome: | The proposed framework can be applied to a wide variety of newsworthy scenarios. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) provide strong semantic representations but are costly and opaque. |
| Approach: | They propose a framework that transfers pretrained language models into symbolic form and integrates them into symbolic models. |
| Outcome: | The proposed framework improves interpretability and accuracy across multiple text classification tasks while remaining fully symbolic and efficient. |
Copied to clipboard
| Challenge: | Existing methods for NLG depend on heavily annotated data, which is infeasible for new domains. |
| Approach: | They propose a system that converts a dialog act into a response in natural language . they propose 'nuclear language generation' to simulate a few-shot learning setting . |
| Outcome: | The proposed model outperforms existing methods on a large set of annotated datasets. |
Copied to clipboard
| Challenge: | Existing commonsense evaluations are often posed as multiple-choice questions, allowing models to exploit systematic biases. |
| Approach: | They propose a generative task that evaluates common sense via multiple open-ended generations and a method that strongly correlates with human judgments. |
| Outcome: | The proposed method outperforms strong language model baselines on a dataset of human and machine common sense. |
Copied to clipboard
| Challenge: | Document interpretation and dialog understanding are the two major challenges for conversational machine reading. |
| Approach: | They propose a discourse-aware entailment reasoning network to strengthen the connection and enhance the understanding of document and dialog. |
| Outcome: | The proposed model improves document interpretation and dialog understanding on the ShARC benchmark. |
Copied to clipboard
| Challenge: | Large vision-language models have been widely used but stereotypical biases are unexplored. |
| Approach: | They propose a framework to SCAN stereotypical bias within large vision-language models . they examine stereotype biases with respect to gender and race in three scenarios . |
| Outcome: | The proposed framework can reduce stereotypical biases in large vision-language models . the currently popular models show significant stereotype biase . |
Copied to clipboard
| Challenge: | Prior work suggests hierarchical organization where different layers specialize in capturing distinct levels of linguistic structure. |
| Approach: | They probe 25 models from BERT Base to Qwen2.5-7B focusing on linguistic properties: lexical identity and inflectional features. |
| Outcome: | The proposed model maintains inflectional features across layers while trading off lexical identity for compact, predictive representations. |
Copied to clipboard
| Challenge: | Existing methods to generalize knowledge bases model triple-level uncertainty . Existing models only model triple level uncertainty, and reasoning results lack global consistency. |
| Approach: | They propose a method to embed knowledge graphs with calibrated probabilistic semantics . they model each entity as a box and relations between two entities as affine transforms based on affinity transforms. |
| Outcome: | Experiments show that the proposed method outperforms baseline methods on confidence prediction and fact ranking. |
Copied to clipboard
| Challenge: | Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work. |
| Approach: | They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations. |
| Outcome: | The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work. |
Copied to clipboard
| Challenge: | Word2Box provides a set-theoretic training objective for learning word representations . word representation is not natural, all senses and contexts, levels of abstraction, variants and modifications which the word may represent are forced to be captured by mat t is nunc. |
| Approach: | They propose a fuzzy-set interpretation of box embeddings and learn box representations of words using a set-theoretic training objective. |
| Outcome: | The proposed model improves word similarity tasks on less common words. |
Copied to clipboard
| Challenge: | Southeast Asia is underrepresented in vision-language research . SEA-VL is an open-source initiative dedicated to developing culturally relevant datasets for SEA languages. |
| Approach: | They propose to use crowdsourced, automated image crawling and synthetic image generation to develop culturally relevant datasets for SEA languages. |
| Outcome: | The proposed datasets capture SEA cultural nuances and contexts better than existing datasets. |
Copied to clipboard
| Challenge: | Misaligned large language models can magnify harm by exploiting them to undermine safety . et al., 2022b; Bai e.t., 2023): misalignment, realignment and model-specific resistance are important . |
| Approach: | They evaluate four methods to identify a mechanism asymmetry between attack and defense . they find that ORPO is most effective for misalignment, but DPO excels in realignment . |
| Outcome: | The proposed methods show a mechanism asymmetry between attack and defense . the proposed methods excel in realignment, but at the expense of model utility . |
Copied to clipboard
| Challenge: | Existing question answering datasets for common sense reasoning are lacking for prototypical situations. |
| Approach: | They propose a question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations. |
| Outcome: | The proposed model outperforms existing models on all evaluation metrics with a meaningful gap. |
Copied to clipboard
| Challenge: | Currently, the Transformer is the de facto architecture of choice for processing sequential data. |
| Approach: | They evaluate the Transformer architecture and its modifications in a shared experimental setting . they conjecture that performance improvements may strongly depend on implementation details . |
| Outcome: | The proposed improvements do not significantly improve performance, the authors find . the proposed improvements are either developed in the same codebase or are minor changes . |
Copied to clipboard
| Challenge: | Existing evaluation methods struggle to ensure semantic correctness and rely on simple or unrealistic datasets. |
| Approach: | They propose a benchmark to evaluate language models’ ability to generate PDDL code from natural language descriptions of planning tasks. |
| Outcome: | The proposed benchmark evaluates the ability of language models to generate PDDL code from natural language descriptions of planning tasks against ground truth and a dataset of 145,918 text-to-PDDL pairs with varying levels of difficulty. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive performance as general purpose agents, but their abilities remain highly dependent on prompts which are hand written with onerous trial-and-error effort. |
| Approach: | They propose an algorithm that uses numerical gradient descent to automatically improve prompts by rewriting vague task descriptions into more precise annotation instructions. |
| Outcome: | The proposed algorithm outperforms previous methods and improves performance on three benchmark NLP tasks and the novel problem of LLM jailbreak detection. |
Copied to clipboard
| Challenge: | Existing toolkits for developing dialog systems are limited to core components and do not support multi-modal processing and social signals. |
| Approach: | They propose to use ADVISER to develop multi-modal dialog agents using multi-text and social signals. |
| Outcome: | The proposed toolkit is flexible, easy to use, and easy to extend for linguists and cognitive scientists, thereby providing a flexible platform for collaborative research. |
Copied to clipboard
| Challenge: | Empirical evaluations on ML models show substantial reductions in unsafe generations and improved robustness against jailbreak attacks. |
| Approach: | They propose a resource-efficient pruning framework that directly identifies unsafe behaviors while preserving model utility. |
| Outcome: | The proposed framework reduces unsafe generations and improves robustness against jailbreak attacks with minimal utility loss. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are powerful at question-answering but prone to hallucinations due to limited domain-specific or up-to-date knowledge. |
| Approach: | They propose a framework for IDentifying RAG properties in LLM services that integrates LLMs with retrieval systems and adds an external retriever and knowledge database to mitigate hallucinations. |
| Outcome: | The proposed framework detects RAG-enhanced LLMs with 99.97% accuracy with partial or no optional knowledge and nearly 100% when the LLM and database are known. |
Copied to clipboard
| Challenge: | Despite recent progress in dialogue evaluation, how to develop automatic metrics remains an open problem. |
| Approach: | They propose a consensus-based framework for dialog evaluation using segment act flows . they propose to crowdsource a large-scale dataset for it to be evaluated . |
| Outcome: | The proposed framework can reach the best or comparable correlation with human evaluation. |
Copied to clipboard
| Challenge: | Recent LLMs like Gemini-1.5 and GPT-4 show exceptional capabilities to understand long contexts directly. |
| Approach: | They propose a method that routes queries to RAG or LC based on model self-reflection. |
| Outcome: | The proposed method significantly reduces the computation cost while maintaining a comparable performance to RAG. |
Copied to clipboard
| Challenge: | Recent advances in long-context modeling have enhanced language models for complex tasks, but they struggle with multi-hop reasoning and noisy contexts. |
| Approach: | They propose an approach that prompts LMs to supply attributions for each assertion during reasoning. |
| Outcome: | The proposed model achieves competitive performance on multi-hop reasoning benchmarks, closely paralleling proprietary LMs such as ChatGPT and Claude-instant. |
Copied to clipboard
| Challenge: | Prompt-based learning can tackle zero-shot and few-shot NLP tasks . authors propose a method that makes use of pre-trained language models . |
| Approach: | They propose to map NLP tasks into natural language prompts, which are then filled by pre-trained language models. |
| Outcome: | The proposed method outperforms standard prompt-based methods in few-shot settings. |
Copied to clipboard
| Challenge: | Evidence-based medicine connects to every individual, yet the nature of it is highly technical . e-fact-checking systems that connect to medical decisions are largely unused . we examine how clinical experts verify real claims from social media . |
| Approach: | They propose that fact-checking should be approached as an interactive communication problem . they argue that social media and AI have made medical knowledge accessible . |
| Outcome: | The proposed method is based on the work of a clinical expert on social media . it reveals that the method is difficult to connect claims to clinical trials . |
Copied to clipboard
| Challenge: | Generative Error Correction (GEC) is a powerful post-processing method to boost the performance of Automatic Speech Recognition systems. |
| Approach: | They propose a method to augment GEC models with retrieved entities to improve accuracy in out-of-domain and out-od scenarios. |
| Outcome: | The proposed method outperforms baseline models on multiple datasets and settings. |
Copied to clipboard
| Challenge: | Existing approaches to generate synthetic data using simple sentence transformations and/or model-based techniques may not generate realistic error samples with respect to the NLG models. |
| Approach: | They propose a framework to train models to classify acceptability of responses generated by natural language generation models using a 2-stage approach . they use existing sentence transformations to generate samples that better resemble the output of the generation models. |
| Outcome: | The proposed approach outperforms existing techniques and can be used in few-shot settings using self-training. |
Copied to clipboard
| Challenge: | Prior work creates evaluations with crowdwork or existing data sources, which are not always available. |
| Approach: | They generate evaluations automatically with language models (LMs) using crowdwork or existing data sources to find out how they behave . |
| Outcome: | The results show that large LMs repeat back a dialog user’s preferred answer and express greater desire to pursue concerning goals like resource acquisition and goal preservation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are rapidly developing and are becoming more and more useful in scientific tasks. |
| Approach: | They propose to use LLM-as-a-judge to grade LLMs on SciEx to assess their ability on scientific tasks. |
| Outcome: | The proposed benchmarks show that the LLMs perform decently on free-form exams, achieving 0.948 Pearson correlation with expert grading. |
Copied to clipboard
| Challenge: | Existing studies on retrievers and LLMs treat them as separate components . a novel bridge model is proposed to optimize the relationship between the retriever and the LLM . |
| Approach: | They propose a framework that chains together supervised and reinforcement learning to train a bridge model that optimizes the connection between the retriever and the LLM. |
| Outcome: | Empirical results show that the proposed model optimizes the connection between the retriever and the LLM. |
Copied to clipboard
| Challenge: | Retraining reward models to address privacy leaks and stereotypes is expensive . recent advances in large language models have led to improvements in understanding . |
| Approach: | They propose a lightweight intrinsic reward that can be used to prune existing LLMs to approximate an "untrust" and an ""untrust "" token distribution. |
| Outcome: | Experiments with two reward models and four LLMs show that selfRW improves trustworthiness with minimal impact on general utility benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in large-scale pre-training provide large models with the potential to learn knowledge from the raw text. |
| Approach: | They propose a posterior-based reweighing and noisy training strategy to exploit generated knowledge in dialogue generation. |
| Outcome: | Empirical results show that the proposed methods outperform the state-of-the-art methods in unsupervised knowledge-grounded conversation. |
Copied to clipboard
| Challenge: | Existing studies on keyphrase generation on non-English languages haven’t been vastly investigated. |
| Approach: | They propose a retrieval-augmented method for multilingual keyphrase generation that leverages keyphrase annotations in English datasets to facilitate generating keyphrases in low-resource languages. |
| Outcome: | The proposed model outperforms baselines on non-English keyphrase generation datasets and the proposed model is scalable. |
Copied to clipboard
| Challenge: | Existing methods to pre-train speech and text use unlabeled data to learn universal feature representations. |
| Approach: | They propose a method to jointly pre-train speech and text in an encoder-decoder modeling framework for speech translation and recognition. |
| Outcome: | The proposed method achieves between 1.7 and 2.3 BLEU improvement above the state of the art on the MuST-C speech translation dataset and comparable WERs to wav2vec 2.0 on the Librispeech speech recognition task. |
Copied to clipboard
| Challenge: | Extensive experiments demonstrate that treating attention as a feature map and applying convolution as . a processing method significantly enhances Transformer performance. |
| Approach: | They propose to use the convolution operator to mimic the processing methods in computer vision to treat attention as a feature map and apply it to neighboring attention scores across different heads. |
| Outcome: | The proposed model can be adapted to various attention-related models and achieves high performance. |
Copied to clipboard
| Challenge: | Existing video generation models struggle to interpret compositional changes and synthesize components across different time steps. |
| Approach: | They propose a temporal compositionality benchmark that uses text prompts and ground truth videos to evaluate compositional changes in video. |
| Outcome: | The proposed benchmark can be used for text-to-video and image-to video generation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have revolutionized natural language processing. |
| Approach: | They propose a Chinese-based platform that assesses Chinese LLMs using a standardized workflow and a unique sampling strategy. |
| Outcome: | CLEVA evaluates Chinese LLMs on a standardized workflow and a competitive leaderboard with minimal coding. |
Copied to clipboard
| Challenge: | Existing models estimate and calibrate confidence of large language models with verbalized uncertainty, but they lack a careful examination of the linguistic knowledge of uncertainty encoded in the latent space of LLMs. |
| Approach: | They draw on typological frameworks of epistemic expressions to evaluate LLMs’ knowledge of epistenetic modality, using controlled stories. |
| Outcome: | The proposed models generate expressions matching the strength of evidence and are not robust in generating epistemic expressions. |
Copied to clipboard
| Challenge: | Large multimodal models have gained attention for their effectiveness to understand and generate descriptions of visual content. |
| Approach: | They propose a multilingual Video LMM benchmark to evaluate video LMMs across 14 languages . they also introduce a machine translated multilingual video training set . |
| Outcome: | The proposed video LMM benchmark is designed to evaluate video Lmms across 14 languages including Arabic, Bengali, Chinese, English, French, German, Hindi, Japanese, Russian, Sinhala, Spanish, Swedish, Tamil, and Urdu. |