Papers by Jian Zhu
Copied to clipboard
| Challenge: | Structured social variation has been extensively studied, e.g., gender based variation, but little is known about how to characterize individual styles due to their idiosyncratic nature. |
| Approach: | They propose a method to study idiolects through a massive cross-author comparison to identify and encode stylistic features. |
| Outcome: | The proposed model achieves strong performance at authorship identification on short texts and through an analogy-based probing task, showing that the learned representations exhibit surprising regularities that encode qualitative and quantitative shifts of idiolectal styles. |
Copied to clipboard
| Challenge: | a recent study shows that multilingual speech processing systems can generalize to unseen languages without adaptation. |
| Approach: | They propose a phoneme-based phoneme embedding model that can be generalized to unseen languages by using a neural forced aligner. |
| Outcome: | The proposed model can generalize to unseen languages without adaptation. |
Copied to clipboard
| Challenge: | Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations. |
| Approach: | They propose a framework that covers 61 risk categories across four modality interactions to address this gap. |
| Outcome: | The proposed framework covers 61 risk categories across four distinct modality interactions. |
Copied to clipboard
| Challenge: | recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability. |
| Approach: | They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs. |
| Outcome: | The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models. |
Copied to clipboard
| Challenge: | Existing methods learn a single user embedding from user’s historical behaviors to represent the reading interest. |
| Approach: | They propose a poly attention scheme to learn multiple interest vectors for each user, which encodes the different aspects of user interest. |
| Outcome: | The proposed approach significantly outperforms existing state-of-the-art methods on the MIND news recommendation benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to improve retrieval performance of large language models are limited by static knowledge. |
| Approach: | They propose a multimodal re-ranking framework that combines curriculum learning with fine-grained reranking and multimodal section reassessment to improve CLIP-based visual coarse-grain retrieval. |
| Outcome: | The proposed framework achieves state-of-the-art answer accuracy and competitive retrieval performance on InfoSeek and Enc-VQA. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities across a wide range of domains, but their generalpurpose pre-training objectives often leave them illsuited for specialized applications such as healthcare. |
| Approach: | They propose a perplexity-aware data scaling law that establishes a predictive relationship between the perplexities of domain-specific data and the test loss. |
| Outcome: | Experiments on medical and general-domain benchmarks show that the proposed scaling law consistently identifies near-optimal training subsets with significantly reduced data consumption. |
Copied to clipboard
| Challenge: | a new study addresses the challenge of learning semantic representations from speech signals . speech-based semantic representation can be used for speech mining and spoken language understanding . |
| Approach: | They propose a multimodal sequential autoencoder that converts speech signals into hidden units . they propose s-HuBERT to induce meaning through knowledge distillation . |
| Outcome: | The proposed model achieves a moderate correlation with human judgments without labels or transcriptions. |
Copied to clipboard
| Challenge: | Existing knowledge base question answering methods generate LFs that are non-executable due to semantic hallucination issue of large language models. |
| Approach: | They propose a "generate-verify-refine" framework for reliable LF generation . they propose ARI-KBQA to generate query paths based on hop-by-hop reasoning . |
| Outcome: | The proposed framework significantly improves model performance with a reduced search space . ARI-KBQA can generate LFs that are non-executable due to semantic hallucination issue . |
Copied to clipboard
| Challenge: | lexical change is a prevalent process, as new words are added, thrive, and decline in day-to-day usage. |
| Approach: | They conduct a large-scale analysis of over 80k neologisms in 4420 online communities over a decade and found that the community’s network structure plays a significant role in lexical change. |
| Outcome: | The results show that the community’s network structure plays a significant role in lexical change. |
Copied to clipboard
| Challenge: | Existing methods to answer subjective questions about products are often imbalanced across product domains. |
| Approach: | They propose a domain-adaptive model that integrates multiple viewpoints into a good answer by integrating these heterogeneous and inconsistent viewpoints. |
| Outcome: | The proposed model integrates multiple viewpoints into a single answer span and is able to integrate them into the answer. |
Copied to clipboard
| Challenge: | Existing language representation models (PLMs) cannot capture factual knowledge from text. |
| Approach: | They propose a unified model for Knowledge Embedding and Pre-trained LanguagERepresentation which integrates factual knowledge into PLMs and produces effective text-enhanced KE with the strong PLM. |
| Outcome: | The proposed model improves on existing pre-trained language representation models and improves their performance on various NLP tasks. |
Copied to clipboard
| Challenge: | IPA transcriptions capture major articulatory contrasts in speech sounds, including the voicing status, place of articulation, manner of voicing, and tongue positions. |
| Approach: | They present ZIPA, a family of efficient speech models that advances the state-of-the-art performance of crosslinguistic phone recognition. |
| Outcome: | The proposed model outperforms existing phone recognition systems on 17,000+ hours of normalized phone transcriptions and a novel evaluation set capturing unseen languages and sociophonetic variation. |
Copied to clipboard
| Challenge: | Sentence function is a significant factor to achieve the purpose of the speaker, but has not been touched in large-scale conversation generation. |
| Approach: | They propose a model to generate informative responses with controlled sentence function using a latent variable and a type controller to deal with compatibility. |
| Outcome: | The proposed model outperforms state-of-the-art models and generates responses with controlled sentence function and informative content. |
Copied to clipboard
| Challenge: | Existing dialog state tracking models neglect rich structural information in a dataset. |
| Approach: | They propose to use curriculum learning to leverage dialog state tracking data . they propose a model-agnostic framework that pre-trains a DST model with schema information . |
| Outcome: | The proposed framework improves performance over a transformer-based and RNN-based model on WOZ2.0 and MultiWOZ2.1. |
Copied to clipboard
| Challenge: | Phone-level modeling of speech is a common approach to speech recognition, but it relies on task-specific architectures and datasets. |
| Approach: | They propose a phonetic framework capable of performing multiple phone-related tasks . they propose 'Phonetic Open Whisper-style Speech Model' that can perform these tasks together . |
| Outcome: | The proposed model outperforms or matches specialized PR models of similar size while supporting G2P, P2G, and ASR. |
Copied to clipboard
| Challenge: | Large language models (LLMs) use tokenization methods but often obscure internal character structures within tokens. |
| Approach: | They propose a method that improves models’ ability to capture character positions within tokens by training them on reverse character prediction tasks using the tokenizer’s vocabulary. |
| Outcome: | Experiments show that the proposed method improves position prediction accuracy in large language models, enabling more precise identification of target characters in original text. |
Copied to clipboard
| Challenge: | Document-level Event Argument Extraction (DEAE) aims to identify arguments and their specific roles from unstructured document. |
| Approach: | They propose a document-prompt-based method for document-level event argument extraction that uses a semantic mention graph to capture relations between documents and prompts. |
| Outcome: | The proposed method surpasses baseline methods and achieves state-of-the-art performance on RAMS and WikiEvents datasets. |
Copied to clipboard
| Challenge: | Large Language Model (LLM) agents are transforming education by automating complex tasks and enhancing both teaching and learning processes. |
| Approach: | This survey analyzes recent advances in applying Large Language Model agents to educational settings . it highlights ethical issues, hallucination and overreliance, and integration with existing ecosystems . |
| Outcome: | The authors analyze the technologies enabling LLM agents and highlight key challenges in deploying them in educational settings. |
Copied to clipboard
| Challenge: | Existing financial question answering datasets lack scope diversity and question complexity. |
| Approach: | They propose to use a dataset for long-form question answering in finance to evaluate QA systems. |
| Outcome: | The proposed dataset includes 1,262 high-quality, source-attributed QA pairs extracted and selected from finance textbooks and government agency websites. |
Copied to clipboard
| Challenge: | Automatically generated radiology reports often receive high scores from existing evaluation metrics but fail to earn clinicians’ trust. |
| Approach: | They propose a meta-evaluation framework that uses criteria spanning discrimination, robustness, and monotonicity to evaluate existing metrics. |
| Outcome: | The proposed framework offers guidance for building more clinically reliable evaluation methods. |
Copied to clipboard
| Challenge: | Domain adaptation is underexplored in multimodal learning environments due to expensive data collection and annotation. |
| Approach: | They propose a bi-alignment scheme to perform drift-drift and anchor-driving matching with partially shifting anchors. |
| Outcome: | The proposed approach achieves superior performance compared with state-of-the-art approaches. |
Copied to clipboard
| Challenge: | Large Language Models achieve strong performance through extended inference-time deliberation, yet how their reasoning failures arise remains poorly understood. |
| Approach: | They propose a framework that probes and redirects critical transitions using uncertainty signals. |
| Outcome: | Empirical evaluations show that GUARD improves reasoning performance . GUard probes critical transitions and redirects them using uncertainty signals . |
Copied to clipboard
| Challenge: | In general, speech synthesis for Indigenous languages is underdeveloped compared to the majority of languages. |
| Approach: | They propose to train a multilingual model on three typologically similar languages to improve performance over monolingual models. |
| Outcome: | The proposed model can train on three similar languages with high performance and is highly competitive with self-attention architectures with higher memory efficiency. |
Copied to clipboard
| Challenge: | Existing methods for aligning open-ended outputs with fine-grained clinician preferences are weakly grounded in professional guidelines. |
| Approach: | They propose a framework to align large language models' outputs with fine-grained clinician preferences . they propose 119 broadly reusable, clinically grounded principles organized by clinical dimensions . |
| Outcome: | The proposed framework outperforms existing models on HealthBench-Hard and Deepseek-R1 and o3. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) often hallucinate due to fragile, linear reasoning and weak visual grounding. |
| Approach: | They propose a framework that reformulates reasoning as a hierarchical search with self-verification and replaces linear Chain-of-Thought with a tree-search policy capable of backtracking to correct logical errors. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on hallucination and safety benchmarks. |
Copied to clipboard
| Challenge: | Natural Language Sentence Matching (NLSM) is a popular NLP task. |
| Approach: | They propose to use QuoraQP to train and evaluate NLSM models using a selection bias framework. |
| Outcome: | The proposed framework can improve generalization ability of trained models and give more trustworthy evaluation results for real-world adoptions. |
Copied to clipboard
| Challenge: | Recent studies have used meta-learning to simulate the few-shot task . however, this sample-wise comparison may be severely disturbed by the various expressions in the same class. |
| Approach: | They propose a meta-learning-based induction network to learn a generalized class-wise representation of each class in a support set. |
| Outcome: | The proposed model outperforms existing state-of-the-art models on a sentiment and dialogue intent datasets. |
Copied to clipboard
| Challenge: | Existing reasoning paradigms that focus on local optimum reasoning lack global perspective. |
| Approach: | They propose a bidirectional reasoning paradigm that generates reasoning paths by bidirectional planning and bottom-up reasoning accumulation. |
| Outcome: | The proposed reasoning paradigm outperforms conventional paradigms with higher accuracy and less searching space to solve complex tasks. |
Copied to clipboard
| Challenge: | Recent research on domain adaptation neglects diversity in translation within a domain . current research on NMT models considers very broad target domains . |
| Approach: | They propose a fine-grained domain adaptation task for autonomous vehicles, AI education, real-time networks, and smart phone. |
| Outcome: | The proposed task is compared with a dataset of Chinese-English translation tasks for four sub-domains of information technology: autonomous vehicles, AI education, real-time networks, and smart phone. |
Copied to clipboard
| Challenge: | Existing methods for Aspect-based sentiment analysis (ABSA) focus on aspect terms with the same sentiment polarity . current methods focus on sentences with only one aspect term or multiple aspect terms . |
| Approach: | They propose a novel method to model inter-aspect relationships and aspect-context relationships simultaneously using a heterogeneous graph. |
| Outcome: | The proposed method can predict sentiments towards the given aspect term in a sentence . it can provide more detailed predictions compared with sentence-level sentiment analysis. |
Copied to clipboard
| Challenge: | Recent studies have shown that models can benefit from query-aware methods for few-shot text classification. |
| Approach: | They propose a dynamic memory-based network for few-short text classification that uses static memory to adapt to unseen classes. |
| Outcome: | The proposed model improves on the miniRCV1 and ODIC datasets by 24% . Detailed analysis is performed to show how the proposed network achieves the new performance. |
Copied to clipboard
| Challenge: | Existing evaluations of phone recognition systems only measure surface-level transcription accuracy. |
| Approach: | They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. |
| Outcome: | The proposed system outperforms LALMs in clinical, educational, and multilingual settings. |
Copied to clipboard
| Challenge: | Existing benchmarks primarily focus on Python and are limited in terms of language diversity. |
| Approach: | They propose a multilingual debugging benchmark that includes 3.9K test samples of 20 programming languages and introduces the debug instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions. |
| Outcome: | The proposed benchmark includes 3.9K test samples of 20 programming languages and covers the automated program repair task, bug localization task, and bug identification task. |
Copied to clipboard
| Challenge: | LINGGYM is a benchmark that evaluates LLMs’ capacity for meta-linguistic reasoning using Interlinear Glossed Text and grammatical descriptions extracted from 18 typologically diverse reference grammars. |
| Approach: | They propose a benchmark that evaluates LLMs’ capacity for meta-linguistic reasoning using Interlinear Glossed Text and grammatical descriptions extracted from 18 typologically diverse reference grammars. |
| Outcome: | The proposed model can generalize linguistic inference across low-resource languages and structures not seen during training. |
Copied to clipboard
| Challenge: | Existing methods for decomposing fine-tuned LLMs are sensitive to the magnitude of delta values. |
| Approach: | They propose a hierarchical quantization framework that shares low-bit integer weights across similar models. |
| Outcome: | The proposed framework achieves an average accuracy degradation of approximately 3% on fine-tuned models across mathematics, coding, chatbot, and Chinese LLMs. |
Copied to clipboard
| Challenge: | Existing studies show that Mamba architectures have room for further optimization in linear projections and state caches. |
| Approach: | They propose a decoupled scale quantization scheme to mitigate outliers in states and channels by applying separate quantization scales. |
| Outcome: | The proposed method reduces memory consumption by 50% across various quantization settings, model sizes, and generation and zero-shot tasks. |
Copied to clipboard
| Challenge: | Empirical studies on three different benchmark conversation datasets demonstrate the effectiveness of the proposed model over several strong baselines. |
| Approach: | They propose an addressee-aware module to automatically learn whether the participant keeps the historical emotional state or is affected by others in the next upcoming turn. |
| Outcome: | The proposed model can predict the participant's emotion in the next upcoming turn without knowing the participant’s response yet. |
Copied to clipboard
| Challenge: | Existing models for story generation suffer from repetition, logic conflicts, and lack of long-range coherence . |
| Approach: | They propose to utilize commonsense knowledge from external knowledge bases to generate reasonable stories by multi-task learning. |
| Outcome: | The proposed model can generate more reasonable stories than state-of-the-art models, compared with existing models, showing that it can capture useful semantic and syntactic features. |
Copied to clipboard
| Challenge: | Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. |
| Approach: | They propose a framework for massively multilingual table question answering that includes tables expanded to 97 languages from Chinese and English sources. |
| Outcome: | Experiments on state-of-the-art LLMs show that synthetically generated training data significantly boosts performance, especially for low-resource languages. |
Copied to clipboard
| Challenge: | In-context learning is an emergent ability from pretrained Large Language Models (LLMs). |
| Approach: | They perform in-depth evaluations of in-context learning on transformers and hybrid large language models using behavioral probing and intervention-based methods. |
| Outcome: | The proposed model performs well on state-of-the-art transformer, state-space, and hybrid large language models. |
Copied to clipboard
| Challenge: | Existing end-to-end dialog systems perform less effectively when data is scarce. |
| Approach: | They propose a Meta-Dialog System which combines meta-learning and human-machine collaboration to improve dialog learning by a new extended-bAbI dataset and a transformed MultiWOZ dataset. |
| Outcome: | The proposed system outperforms non-meta-learning baselines on a new extended-bAbI dataset and a transformed MultiWOZ dataset for low-resource goal-oriented dialog learning. |
Copied to clipboard
| Challenge: | Low-rank adaptation methods for large language models have limitations in preserving world knowledge and limiting updates to preserve world knowledge. |
| Approach: | They propose a Fisher-optimized adaptive low Rank and Singular-VectorSelection framework for knowledge-preserving fine-tuning that allows efficient and task-sensitive updates. |
| Outcome: | The proposed framework outperforms existing methods for knowledge-preserving fine-tuning. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. |
| Approach: | They propose a Meta-Reasoning informed alignment framework that quantifies state-transition probabilities along the model’s thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments. |
| Outcome: | Empirical evaluations of four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning. |
Copied to clipboard
| Challenge: | Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence. |
| Approach: | They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included . |
| Outcome: | The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model. |
Copied to clipboard
| Challenge: | a recent study shows that agent research practices are far from standard, rigorous . lack of a standard evaluation protocol makes previous works not reproducible, authors say . |
| Approach: | They conduct an empirical study on the GAIA benchmark to investigate agent design choices . they find that lack of a standard evaluation protocol makes previous works not reproducible . |
| Outcome: | The proposed framework achieves state-of-the-art performance among open-source projects. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have strong capabilities in in-context learning, but verifying the correctness of their generated responses remains a challenge. |
| Approach: | They propose a token-level attribution method that combines Shapley value-based data attribution with KNN-based retrieval techniques to improve attribution accuracy. |
| Outcome: | TokenShapley outperforms state-of-the-art methods on four benchmarks . it achieves an 11–23% improvement in accuracy on the benchmarks. |
Copied to clipboard
| Challenge: | Existing repository-level code completion benchmarks focus on a limited number of languages . existing benchmarks report overall average scores of different languages ignoring fine-grained abilities . |
| Approach: | They propose to use repository-level code completion benchmarks to evaluate general code intelligence abilities across languages for existing code Large Language Models. |
| Outcome: | The proposed benchmarks improve the code completion abilities of existing LLMs by using two types of annotations on the parsed syntax tree. |
Copied to clipboard
| Challenge: | Existing studies on text style transfer neglect long style transfer at the discourse level. |
| Approach: | They propose a model that transfers text style into target styles with learnable style embeddings . they use a mask-and-fill framework to explicitly fuse style-specific keywords into generation . |
| Outcome: | The proposed model outperforms baselines in style transfer and content preservation. |
Copied to clipboard
| Challenge: | Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs. |
| Approach: | They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. |
| Outcome: | The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities. |
Copied to clipboard
| Challenge: | Existing methods for enhancing multi-step reasoning have not fully translated to multilingual contexts. |
| Approach: | They propose a framework that leverages language-conditioned hints to guide exploration in non-English reasoning tasks. |
| Outcome: | Empirical results show that the proposed framework improves reasoning performance without compromising language consistency. |