Papers by He Du
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, but their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency. |
| Approach: | They propose a novel LLM-in-the-loop framework that integrates Large Language Models with Neural Topic Models (NTMs) global topics and document representations are learned through the NTM, while an LLM refines these topics using an Optimal Transport (OT)-based alignment objective. |
| Outcome: | The proposed framework improves topic interpretability while preserving the efficiency of existing NTMs. |
Copied to clipboard
| Challenge: | Medical dialogue systems (MDS) struggle to identify relevant medical knowledge and generate accurate responses. |
| Approach: | They propose a medical dialogue system that integrates knowledge refining and dynamic prompt adjustment to improve medical knowledge and accuracy. |
| Outcome: | The proposed system outperforms state-of-the-art systems in both generation quality and medical entity accuracy. |
Copied to clipboard
| Challenge: | recurrent neural networks struggle to match the performance of Transformers due to limitations in parallelization and scalability. |
| Approach: | They propose a model architecture that combines the efficient parallelizable training of transformers with the efficient inference of RNNs. |
| Outcome: | The proposed model performs on par with similarly sized RNNs, suggesting future work can leverage this architecture to create more efficient models. |
Copied to clipboard
| Challenge: | Large language models store factual knowledge in their parameters but their parametric knowledge can conflict with the information provided in the context. |
| Approach: | They propose a training-free representation engineering method that uses pre-trained sparse auto-encoders to control the knowledge selection behaviour of large language models. |
| Outcome: | The proposed method can control the use of both knowledge sources to resolve knowledge conflict in open-domain question-answering tasks surpassing existing representation engineering methods (+10%) and contrastive decoding methods (+5%). |
Copied to clipboard
| Challenge: | Prior implicit CoT methods have underperformed in terms of efficiency and robustness by relying on natural language tokens for reasoning. |
| Approach: | They propose a training framework that compresses natural language CoT into continuous space by aligning hidden states of a designated token. |
| Outcome: | The proposed framework outperforms the existing state-of-the-art in 3.1x compression rate and 28.2% accuracy on GSM8k scale. |
Copied to clipboard
| Challenge: | Automated synthesis of zeolite holds great significance for attaining economic and environmental benefits. |
| Approach: | They propose an event extraction task to mine structural synthesis actions from experimental narratives for modular automated synthesis. |
| Outcome: | The proposed method can significantly expedite automated synthesis of zeolites owing to its machine readability. |
Copied to clipboard
| Challenge: | Recent advances in reinforcement learning, such as DeepSeek R1-Zero, highlight the effectiveness of incentive training, but these methods rely on external verifiers, which limits their applicability to domains like mathematics and coding, where such verifier is readily available. |
| Approach: | They propose a general reinforcement learning framework that requires only standard supervised fine-tuning data with no need for an external verifier. |
| Outcome: | The proposed framework outperforms the model of the same size distilled from large reasoning models such as DeepSeek R1 671B by 7.7%. |
Copied to clipboard
| Challenge: | Neural topic models (NTMs) use deep neural networks to learn topic information. |
| Approach: | They propose a variational autoencoder model that reconstructs sentence and document word counts using bag-of-words embeddings and pre-trained semantic embedders. |
| Outcome: | The proposed model lowers reconstruction errors at sentence and document levels and finds more coherent topics from real-world datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models rely on spurious correlations in the data for natural language understanding (NLU) tasks. |
| Approach: | They propose a framework for debiasing shortcuts and a dummy class to encode shortcuts into a model and use it to generate soft labels. |
| Outcome: | The proposed framework significantly improves out-of-distribution generalization while maintaining satisfactory in-district accuracy. |
Copied to clipboard
| Challenge: | Experimental results show that opensource curriculum training is more effective when distinct datasets are available for different training stages. |
| Approach: | They propose an opensource suite for training long reasoning models using publicdata and models. |
| Outcome: | The proposed model outperforms DeepSeek-R1-DistillQwen-32B models in math reasoning. |
Copied to clipboard
| Challenge: | Argument mining (AM) is a challenging task as it requires recognizing complex argumentation structures involving multiple subtasks. |
| Approach: | They propose a generative framework where expected outputs of AM are framed as a simple target sequence. |
| Outcome: | The proposed framework achieves state-of-the-art on two AM benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models excel in code generation benchmarks, but these benchmarks focus on single-file scenarios with constrained context scope. |
| Approach: | They propose an open-source framework to effectively resolve GitHub issues using a code file retrieval module and a model-based code editing module. |
| Outcome: | The proposed approach achieves state-of-the-art performance on two GitHub benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are difficult to align with high-stakes medical standards due to dissonance between coarse-grained preference signals and complex protocols. |
| Approach: | They propose a framework that aligns Large Language Models with medical standards . they use a dataset generated via a human-in-the-loop pipeline to augment medical instructions . |
| Outcome: | The proposed framework disentangles safety constraints from general proficiency, enabling precise guidance during reinforcement learning. |
Copied to clipboard
| Challenge: | Existing language model agents excel in planning and reasoning, but lack creativity in unfamiliar environments. |
| Approach: | They propose a benchmark suite of room escape game environments to challenge agents with creative reasoning, unconventional tool use and iterative problem-solving to uncover implicit goals. |
| Outcome: | The proposed framework can perform with 40% fewer steps and hints and performs robustly across difficulty levels. |
Copied to clipboard
| Challenge: | In the evolving landscape of large language models, the predominant focus has been on English and Chinese. |
| Approach: | They propose to utilize Arabic-specific vocabulary in the tokenizer to accelerate decoding. |
| Outcome: | The proposed model achieves decent performance comparable to the best Arabic LLMs across various Arabic benchmarks. |
Copied to clipboard
| Challenge: | Existing ToM reasoning methods rely excessively on off-the-shelf LLMs, reducing their efficiency and limiting their applicability to high-order ToM. |
| Approach: | They propose a neuro-symbolic framework that integrates a Neural Knowledge Base of Entity States and knowledge injection to enhance ToM reasoning. |
| Outcome: | The proposed framework improves ToM reasoning on ToMi, HiToM, and FANToM benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate hallucinations when handling unfamiliar information. |
| Approach: | They propose a Chinese benchmark to evaluate large language models' knowledge and reasoning capabilities in medication tasks. |
| Outcome: | The proposed benchmark evaluates models in indication, dosage and administration, contraindicated population, mechanisms of action, drug recommendation, and drug interaction across six datasets. |
Copied to clipboard
| Challenge: | Foundational models and their checkpoints have advanced deep learning, boosting performance across applications. |
| Approach: | They propose a method for pruning fine-tuned models by calculating differences between them and original model. |
| Outcome: | The proposed method can improve performance across vision, NLP, and multi-modal benchmarks. |
Copied to clipboard
| Challenge: | Cognitive task analysis (CTA) is a type of analysis used to elicit and represent the knowledge and thought processes of domain experts. |
| Approach: | They propose a weakly-supervised framework for automated CTA transcript parsing . they partition the parser process into a sequence labeling task and a text span-pair relation extraction task with distant supervision from human-curated protocol files. |
| Outcome: | The proposed framework reduces human labor and scales the task to a small scale. |
Copied to clipboard
| Challenge: | Large Language Models struggle with the "curse of two-hop reasoning" in compositional tasks. |
| Approach: | They propose to form a "Generalization Circuit" during a prolonged "grokking" phase . they argue that grokkking is the process of integrating memorized atomic facts into an easy-acquire reasoning path. |
| Outcome: | The proposed model is superior to non-grokked models, but it requires a large computational cost . the study shows that grokking is not the sudden acquisition of a new reasoning paradigm . |
Copied to clipboard
| Challenge: | MMLU is widely adopted but its ground truth errors obscure the true capabilities of LLMs. |
| Approach: | They propose a framework for identifying dataset errors using a novel error annotation protocol and a subset of 5,700 manually re-annotated questions. |
| Outcome: | The proposed framework is based on 5,700 re-annotated questions from the MMLU benchmark. |
Copied to clipboard
| Challenge: | Open-source code language models (code LMs) are a growing threat for intellectual property protection. |
| Approach: | They propose a black-box code LM watermarking framework that uses rule-based watermarks and utility-preserving injection method for user-level model tracing. |
| Outcome: | The proposed framework shows that it performs well across multiple state-of-the-art code LMs and is harmless compared to existing baselines. |
Copied to clipboard
| Challenge: | Existing N-ToM benchmarks lack ambiguous and artificial narratives, lack of personality traits and preferences, and limited diversity in the questions posed. |
| Approach: | They propose a benchmark to assess Neural Theory-of-Mind (N-ToM) with longer and clearer narrative stories, characters with explicit personality traits, actions triggered by character intentions, and questions designed to challenge LLMs’ abilities of modeling characters’ mental states. |
| Outcome: | The proposed test aims to assess the performance of LLMs in the physical and psychological worlds. |
Copied to clipboard
| Challenge: | Conference call transcripts contain significant redundancy and industry-specific terminology that creates obstacles for language models. |
| Approach: | They propose a Sparse Autoencoder for Financial Representation Enhancement framework to extract key information from earnings conference call transcripts and eliminate redundancy. |
| Outcome: | The proposed method outperforms baselines in analyzing earnings conference call transcripts. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have advanced Chinese Classical Studies (CCS) but the audio dimension of CCS remains underexplored due to a lack of high-quality, domain-specific audio corpora. |
| Approach: | They propose a 119-hour audio corpus comprising 22,000 audio samples to bridge this gap . it encompasses a diverse range of literary genres across six tasks . |
| Outcome: | The proposed corpus encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering ( SQA), Speech Understanding (SU), and Speech Reasoning (SR). |
Copied to clipboard
| Challenge: | Existing benchmarks lack the ability to automatically evaluate from users’ perspective and lack the explainability of the results of LLM agents’ code generation capabilities. |
| Approach: | They propose a new benchmark for LLM agents' automated evaluation by simulating user interaction. |
| Outcome: | The proposed benchmark can evaluate the generated projects by user interaction simulation and by code similarity through existing objective indicators. |
Copied to clipboard
| Challenge: | Existing studies show that advanced LLMs produce text indistinguishable from human writing. |
| Approach: | They propose a benchmark to assess persona simulation across diverse contexts by decomposing the evaluation into six fundamental capabilities including opinion consistency, memory recall, logical reasoning, persona tone, and syntactic style. |
| Outcome: | The proposed model achieves moderate accuracy but falls short of the basic capabilities needed to simulate personas in real-world contexts. |
Copied to clipboard
| Challenge: | Existing methods for designing and optimizing multi-agent systems are static and do not learn from experience. |
| Approach: | They propose a framework that enables a multi-agent system to learn to evolve . they use "textual gradients" to pinpoint failures and suggest granular modifications . |
| Outcome: | a new framework enables a multi-agent system to learn to evolve . it learns from historical optimization experiences to improve its performance . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a cornerstone in artificial intelligence due to their exceptional performance. |
| Approach: | They propose a training-free approach that integrates existing MLLMs for effective multimodal expansion while retaining their original performance. |
| Outcome: | The proposed approach can expand LLMs' multimodal capabilities while retaining original performance. |
Copied to clipboard
| Challenge: | Radiology report generation aims at generating descriptive text from radiology images automatically. |
| Approach: | They propose a weakly supervised contrastive loss method that generates descriptive text from radiology images automatically. |
| Outcome: | The proposed method outperforms previous work on correctness and text generation metrics for two public benchmarks. |
Copied to clipboard
| Challenge: | Recent research has demonstrated the value of user feedback, but there are still issues to consider, such as the difficulty in tracking changes and comparing different models. |
| Approach: | They propose a human-in-the-loop topic modeling system that integrates users’ knowledge into the modelling process, enabling them to refine the model iteratively. |
| Outcome: | The proposed system is based on a series of user studies to assess its performance in progressively more realistic applications. |
Copied to clipboard
| Challenge: | Existing benchmarks focused on simplified or isolated aspects of coding, ignoring the full spectrum of programming challenges. |
| Approach: | They propose a case study that examines the performance of large language models across the entire software development lifecycle with four programming languages, multiple domains, and carefully designed and verified metrics for each task. |
| Outcome: | The proposed model performs across the entire software development lifecycle, including design, environment setup, implementation, acceptance testing, and unit testing. |
Copied to clipboard
| Challenge: | Existing approaches to building generalizable verifiable data are task-specific and lack a principled, universal evaluator of verifikatability. |
| Approach: | They propose a task-agnostic, strategy-guided, executably-checkable data synthesis framework that synthesizes problems, diverse candidate solutions and verification artifacts from a single source. |
| Outcome: | The proposed framework synthesizes problems, candidates, and verification artifacts from human-annotated and strategy-induced checks and iteratively discovers strategies. |
Copied to clipboard
| Challenge: | Existing agentic approaches for Knowledge Graph-based Retrieval-Augmented Generation fail to generalize to real-world enterprise Knowledge graphs (KGs) dense, schema-driven, and operationally constrained, requiring a training-free framework. |
| Approach: | They propose a training-free framework that integrates structured planning with controlled iterative reasoning by injecting schema-conditioned structural priors and enforcing schemas during multi-hop reasoning. |
| Outcome: | The proposed framework significantly improves on a real-world enterprise-oriented benchmark constructed from a Configuration Management DataBase (CMDB). |
Copied to clipboard
| Challenge: | Existing topic models infer topics based on word co-occurrence information, which results in degraded performance and degrades performance. |
| Approach: | They propose a generative model that aggregates short texts into clusters by leveraging the associated meta information. |
| Outcome: | The proposed model can generate more interpretable topics and document clusters. |
Copied to clipboard
| Challenge: | Existing solutions for text-to-image synthesis are sensitive on textual prompts, posing a challenge for novice users. |
| Approach: | They propose a dialogue-based TIS prompt generation model that emphasizes user experience for novice users. |
| Outcome: | The proposed model emphasizes user experience for novice users . it improves user-centricity score while maintaining a competitive quality of synthesized images. |
Copied to clipboard
| Challenge: | Current empirical methods that focus on isolated tools learning struggle with accurate multi-tool selection due to issues like confusing similar tools and neglecting dependencies. |
| Approach: | They propose a tool-learning paradigm which integrates tools and trial-and-error experiences into a network characterized by semantic similarity and dependency relationships. |
| Outcome: | The proposed model outperforms existing methods on multiple real-world API datasets and significantly outperformed baselines. |
Copied to clipboard
| Challenge: | Existing methods for aspect-sentiment analysis ignore internal correlations between aspect extraction and sentiment classification. |
| Approach: | They propose a hierarchical interactive network to model two-way interactions between two tasks appropriately using shallow-level and deep-level inputs. |
| Outcome: | Extensive experiments on three real-world datasets demonstrate that the proposed model outperforms existing methods. |
Copied to clipboard
| Challenge: | Existing solutions for table reasoning tasks are mainly tested on small tables and face scalability issues and struggle with complex queries due to incomplete or dispersed data across different table sections. |
| Approach: | They propose a table reasoning pre-processor suite that can be used to leverage large language models (LLMs) in table-based tasks. |
| Outcome: | The proposed method improves LLMs’ reasoning capabilities in various tabular tasks and enhances interaction between LLM and tabular data by employing effective pre-processing. |
Copied to clipboard
| Challenge: | Existing methods to inject safety-aligned large language models rely on token-level mappings, which do not guarantee sustained harmful output. |
| Approach: | They propose a method that directly modifies model weights to map a trigger to an attacker-specified response. |
| Outcome: | The proposed method achieves high triggered attack success while maintaining non-triggered safety and general utility. |
Copied to clipboard
| Challenge: | Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery, but their capability in reproducing code from research papers remains underexplored. |
| Approach: | They propose to evaluate LLM agents' ability to reproduce scientific research papers by analyzing code reproduction tasks from 23 research papers published in top-tier NLP venues. |
| Outcome: | The proposed benchmark systematically evaluates the capability of large language model (LLM) agents on code reproduction from Language Modeling Research. |
Copied to clipboard
| Challenge: | Existing variational Bayesian models generate responses from a single latent variable, which is not sufficient to model high variability in responses. |
| Approach: | They propose a conditional variable auto-encoder that sequentially introduces latent variables to condition the generation of each word in the response sequence. |
| Outcome: | Empirical results show that the proposed model improves on state-of-the-art models on Opensubtitle and Reddit datasets. |
Copied to clipboard
| Challenge: | Experimental results prove the superiority of our proposed method on challenging classification tasks. |
| Approach: | They propose a task-level thinking step that eliminates bias introduced by demonstrations . they propose 'progressive revision framework' which can improve the thinking steps by correcting hard demonstrations. |
| Outcome: | The proposed method achieves best performance on three kinds of classification tasks in zero-shot and few-shot settings. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities in natural language understanding and generation, but controlling their behavior remains a challenge. |
| Approach: | They propose a supervised steering approach that operates in sparse, interpretable representation spaces. |
| Outcome: | The proposed approach achieves higher success rates with minimal degradation in generation quality compared to existing methods. |
Copied to clipboard
| Challenge: | Existing studies have revealed the robustness degra-dation caused by data distillation. |
| Approach: | They propose a framework to evaluate and quantify model distillation . they aim to identify identity cognition contradictions and analyse multi-granularity response similarities across models to measure the extent of homogenization. |
| Outcome: | The proposed framework addresses two key aspects: (1) Identifying identity cognition contradictions to assess discrepancies in how models perceive and represent identity-related information; (2) Analyzing multi-granularity response similarities across models to measure the extent of homogenization. |
Copied to clipboard
| Challenge: | Existing methods for few-shot learning are based on labeled examples, but they are non-trivial . few-sshot learning is challenging due to the imbalance in the amount of data between the source and target domains. |
| Approach: | They propose retrieval-based methods for intent classification and slot filling tasks . they use a batch-softmax objective to learn similar contextualized representations for spans . |
| Outcome: | The proposed method outperforms previous systems on the CLINC and SNIPS benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to predict slots and their values do not encode enough semantic information, limiting the models’ zero-shot capability. |
| Approach: | They propose a QA-driven slot filling model which extracts slot-filler spans from utterances with a span-based QA model. |
| Outcome: | The proposed model outperforms baselines by over 5% on the SNIPS benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to measure word segmentation only assess the language model's understanding of the overall meaning of sentences, lacking an evaluation of the language models' understanding capabilities at a fine-grained level. |
| Approach: | They propose a framework to explore the limit of unsupervised word segmentation with Large Language Models (LLMs) they employ current mainstream LLMs to perform word segmentations across multiple languages . |
| Outcome: | The proposed method improves on existing methods and combines the advanced pattern recognition capabilities of Aho-Corasick automata with the deep insights of well-pretrained LLMs. |
Copied to clipboard
| Challenge: | Existing LLMs are difficult to achieve satisfactory results in table-related tasks. |
| Approach: | They propose to develop a specialized logical table-to-text generation model that can be used for table-related tasks. |
| Outcome: | The proposed model achieves state-of-the-art on a Logic2Text dataset. |