Papers by Ge Xu
Copied to clipboard
| Challenge: | Recent studies show that learning-based models fail to accurately capture and describe abnormal regions due to data bias. |
| Approach: | They propose a model that compares the current input image with normal images to capture abnormal regions by contrasting the input image and normal images. |
| Outcome: | The proposed model can be easily incorporated into existing models to boost their performance under most metrics. |
Copied to clipboard
| Challenge: | Formality style transfer is a task of automatically transforming text in one particular formality style into another. |
| Approach: | They propose to augment parallel data with three specific data augmentation methods to improve the model's generalization ability and reduce the overfitting risk. |
| Outcome: | The proposed methods significantly improve performance when used to pre-train the model and lead to the state-of-the-art results in the GYAFC benchmark dataset. |
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning with verifiable rewards suffer from limited exploration diversity and inefficient reasoning. |
| Approach: | They propose a method that rewards concise and correct reasoning while penalizing unnecessarily long reasoning chains. |
| Outcome: | Extensive experiments on Qwen and Llama models validate the effectiveness and efficiency of ROSE. |
Copied to clipboard
| Challenge: | Large language models (LLMs) excel in various tasks, but often produce hallucinations . retrieved contexts, misrepresent information, or generate outright contradictions . |
| Approach: | They propose a framework that measures hallucination faithfulness of large language models . they introduce a leaderboard that leverages diverse human-annotated hallucinian examples . |
| Outcome: | The proposed framework improves hallucination evaluations by leveraging human-annotated examples. |
Copied to clipboard
| Challenge: | Recent studies on compression of pretrained language models usually use preserved accuracy as the metric for evaluation. |
| Approach: | They propose two new metrics that measure how closely a compressed model mimics the original model. |
| Outcome: | The proposed metrics measure how closely a compressed model (i.e., student) mimics the original model (e.g., teacher). |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to memorize, recall, and reason in sustained interactions. |
| Approach: | They propose a multimodal real-world conversation benchmark for evaluating open-ended abilities of multimodal large language models. |
| Outcome: | The proposed benchmarks show that the models perform better in open-ended conversations. |
Copied to clipboard
| Challenge: | Cant is important for understanding advertising, comedies and dogwhistle politics . currently, there are very few resources available for the research of cant . |
| Approach: | They propose a large and diverse dataset for creating and understanding cant from a computational linguistics perspective. |
| Outcome: | The proposed dataset can be used to test word embedding similarity and pretrained language models. |
Copied to clipboard
| Challenge: | Existing work on generating empathetic responses by utilizing the speaker's emotion has not been successful. |
| Approach: | They propose an approach which incorporates an adaptive module for commonsense knowledge selection to ensure consistency between the generated empathetic responses and the speaker’s situation. |
| Outcome: | The proposed approach outperforms baseline models in both automatic and human evaluations, exhibiting the generation of more coherent and empathetic responses. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are gaining a foothold in Recommender Systems (RS) but there is growing concern that LLMs perpetuate stereotypes and may result in unfair recommendations. |
| Approach: | They propose a counterfactually-fair-prompt method for LLM-based recommendation that is based on unbiased foundation mOdels. |
| Outcome: | The proposed method achieves better recommendation performance with a high level of fairness on two real-world datasets. |
Copied to clipboard
| Challenge: | RAAMove is a comprehensive multi-domain corpus dedicated to the annotation of move structures in Research Article (RA) abstracts. |
| Approach: | They propose a multi-domain corpus dedicated to the annotation of move structures in RA abstracts. |
| Outcome: | The proposed corpus is based on a human-annotated dataset and a BERT-based model to verify its effectiveness. |
Copied to clipboard
| Challenge: | Recent advances in audio diffusion models have significantly improved text-to-audio editing via inversion techniques, but these models typically rely on dense, fixed-step sampling trajectories to maintain structural integrity. |
| Approach: | They propose a model-agnostic Adaptive Trajectory Extrapolation framework that accelerates inversion-based editing process by dynamically evaluating only the most critical generative phases. |
| Outcome: | The proposed framework achieves a 3.9 speedup with negligible loss in fidelity. |
Copied to clipboard
| Challenge: | Larger models or those trained on fewer tokens exhibit less quantization-induced degradation (QiD), while smaller, well-trained models face significant performance losses. |
| Approach: | They propose to use QiD to measure an LLM’s training levels and determine the number of training tokens required for fully training LLMs of various sizes. |
| Outcome: | The proposed scaling laws can predict the quantization performance of different-sized LLMs trained with tokens. |
Copied to clipboard
| Challenge: | Current music information retrieval systems struggle to meet linguistic diversity challenges . current systems struggle with text queries in non-English languages . |
| Approach: | They propose a music information retrieval system that supports both ABC notation and MIDI . CLaMP 2 includes a multilingual text encoder and a multiple-modal music encoder . |
| Outcome: | The proposed system achieves state-of-the-art results in multilingual semantic search and music classification across modalities. |
Copied to clipboard
| Challenge: | Mental disorders affect nearly one in seven people worldwide, yet the vast majority do not receive adequate care. |
| Approach: | They propose a framework to evaluate LLMs' ethical knowledge and behavioral responses through multiple-choice and open-ended tasks with fine-grained ethicality annotations. |
| Outcome: | Empirical results across 14 models reveal that refusal rates are poor indicators of ethical behavior, revealing a significant divergence between safety triggers and clinical appropriateness. |
Copied to clipboard
| Challenge: | Current outcome-centric verification paradigms neglect potential errors in the derivation process. |
| Approach: | They propose a process-aware RLVR training paradigm utilizing verifiers selected via **PRIME**. |
| Outcome: | The proposed approach outperforms the baseline verification paradigm on AIME24, AIME25, and Beyond-AIME models. |
Copied to clipboard
| Challenge: | Existing approaches to RLVR use multiple-choice questions as verifiable rewards . however, not all tasks provide reliable verification . |
| Approach: | They propose a framework that actively constructs high-quality distractors to block elimination shortcuts and promote deep reasoning. |
| Outcome: | The proposed method significantly improves reasoning capabilities of Large Language Models. |
Copied to clipboard
| Challenge: | Existing studies have focused on human-annotated search queries but they can not cover conversations of various domains. |
| Approach: | They propose a domain adaptation framework that uses retrieval-augmented generation to improve the model's robustness. |
| Outcome: | The proposed model is more robust and performs significantly better in a more challenging setting over strong baselines. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are used as automatic evaluators to provide accurate and reliable assessments. |
| Approach: | They propose a framework that integrates LLM-based judgment models into a multi-agent system and simulates the interactive client-server polling mechanism. |
| Outcome: | The proposed framework outperforms supervised models trained on annotated judgment data while requiring no human-labeled annotations. |
Copied to clipboard
| Challenge: | Existing models that mess up or drop the core structural information of input graphs are lacking in graph-to-text generation. |
| Approach: | They propose to leverage richer training signals to guide a graph-to-text generation model by focusing on autoencoding losses and back-propagating the losses to better calibrate the model. |
| Outcome: | Experiments on two benchmarks show the proposed model over a state-of-the-art model . two types of autoencoding losses are used to back-propagate the model based on multitask training . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit suboptimal behaviors and inconsistencies when exposed to unfamiliar external information, underscoring their limitations in effectively leveraging such knowledge. |
| Approach: | They propose a framework that enhances the external knowledge utilization of Large Language Models through a two-stage constructivist cognitive modeling process. |
| Outcome: | The proposed framework achieves a 10% improvement over baseline methods on various question-answering benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to generate error-corrected sentence pairs for improving grammatical error correction are not available. |
| Approach: | They propose a method to generate error-corrected sentence pairs for improving grammatical error correction based on machine translation models of different qualities . |
| Outcome: | The proposed method can generate multiple error-corrected sentence pairs from Chinese to English text. |
Copied to clipboard
| Challenge: | Summarization is an important application of Large Language Models. |
| Approach: | They integrate human-annotated and model-generated natural language explanations to elucidate how a summary deviates and becomes inconsistent with its source article. |
| Outcome: | The proposed model provides rationales for its judgments and improves its accuracy significantly. |
Copied to clipboard
| Challenge: | Recent Large Multimodal Models (LMMs) have shown promising potential for performing end-to-end KIE directly from document images. |
| Approach: | They propose a benchmark to evaluate the performance of Large Multimodal Models (LMMs) using a constrained-category KIE track and an open-categorical KIE Track. |
| Outcome: | Experiments on 15 state-of-the-art LMMs show performance degradation under diverse schema definitions, long-tail key fields, and complex layouts, along with pronounced performance disparities across different document types and scenarios. |
Copied to clipboard
| Challenge: | Knowledge Graphs (KGs) typically treat updates as independent facts . factual, localized updates can contradict and invalidate previously correct knowledge . |
| Approach: | They propose a model-agnostic framework for cascading KG update identification that leverages conformal prediction to provide reliable uncertainty guarantees over the cascade as a whole. |
| Outcome: | The proposed framework provides reliable uncertainty guarantees over the cascade as a whole . it integrates large language models to enrich event representations with world knowledge. |
Copied to clipboard
| Challenge: | Existing models lack multimodal understanding capabilities, resulting in closed-source model that does not support multimodal interleaved sequences. |
| Approach: | They propose a foundation model built on multimodal tokens capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. |
| Outcome: | The proposed model is able to understand speech, text, images, and videos in an end-to-end, autoregressive manner. |
Copied to clipboard
| Challenge: | a novel approach to compress neural networks by progressive module replacement is proposed . a number of techniques have been proposed to compress pretraining and fine-tuning models . |
| Approach: | They propose a model compression approach that divides BERT into modules and builds their compact substitutes. |
| Outcome: | The proposed approach outperforms existing knowledge distillation approaches on GLUE benchmark . it is based on a model that divides the original BERT into several modules and builds their substitutes . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly vulnerable to elusive and implicit intentions, causing security risks and compromising user experience. |
| Approach: | They propose a method to detect and mitigate implicit jailbreak attacks using LLMs by unearthing real intentions and a greedy gradient-based algorithm to remove the least important parts of a sentence. |
| Outcome: | The proposed method reduces attacks success rate and Harmful Score while maintaining overall model performance. |
Copied to clipboard
| Challenge: | Existing text infilling objectives for pretrained language models require self-supervision by masking out tokens or spans in text. |
| Approach: | They propose to extend text infilling to a self-supervised sequence-to-sequence (Seq2Sequen) task. |
| Outcome: | The proposed task improves the model's performance on various natural language generation tasks. |
Copied to clipboard
| Challenge: | Existing frameworks depend on rigid, pre-defined external tools to extend perceptual capabilities of VLMs. |
| Approach: | They propose a framework that leverages self-emergent linguistic toolchains to enhance visual perception and reasoning. |
| Outcome: | The proposed framework improves the visual perception capabilities of large language models by incorporating external visual documents to address a given query. |
Copied to clipboard
| Challenge: | Video Large Language Models (VLMs) have been praised for their performance in coarse-grained video understanding but still face ineffective temporal grounding and inadequate timestamp representations. |
| Approach: | They propose a novel Video-LLM that senses and reasoned over specific video moments with fine-grained temporal precision. |
| Outcome: | The proposed model surpasses existing models in fine-grained video understanding tasks and exhibits strong potential as a general video understanding assistant. |
Copied to clipboard
| Challenge: | DropHead is a structured dropout method for regularizing multi-head attention . DropHed drops entire attention heads during training to prevent overfitting . |
| Approach: | They propose a structured dropout method specifically designed for regularizing multi-head attention mechanism . DropHead drops entire attention heads during training to prevent overfitting . |
| Outcome: | The proposed method can improve transformer models by 0.9 BLEU score on translation task and around 1.0 accuracy for various text classification tasks. |
Copied to clipboard
| Challenge: | Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence. |
| Approach: | They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included . |
| Outcome: | The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model. |
Copied to clipboard
| Challenge: | Existing evaluations of hallucinations in large language models suffer from a lack of diversity and recency in the LLM and LLM families considered. |
| Approach: | They propose a summarization hallucination benchmark that challenges models to disagree on hallucines . they use models to generate answers or summaries from textual input . |
| Outcome: | The proposed model combines the best of 10 modern LLMs with ground truth annotations. |
Copied to clipboard
| Challenge: | Existing methods for generating EMRs from doctor-patient dialogues produce rigid and repetitive dialogues. |
| Approach: | They propose a framework that integrates Intent Graph Planning, Dual-Agent Simulation and Rule-Reward Quality Control to generate realistic doctor-patient dialogues. |
| Outcome: | The proposed framework significantly enhances realism, diversity and downstream EMR quality, reducing physician editing efforts. |
Copied to clipboard
| Challenge: | Existing methods for video captioning consider a sequence of frames and biases towards focused objects. |
| Approach: | They propose an Object-Oriented Non-Autoregressive approach to video captioning . it performs three steps: 1) identify the focused objects and predict their locations . 2) generate related attribute words and relation words of these focused objects to form a draft caption . |
| Outcome: | The proposed method achieves competitive results with the state-of-the-art methods but with higher diversity and faster inference speed. |
Copied to clipboard
| Challenge: | Chinese and Japanese share many characters with similar surface morphology. |
| Approach: | They propose a Chinese-Japanese pretrained masked language model with a coarse-to-fine training approach to exploit the shared knowledge across the languages. |
| Outcome: | The proposed model is effective on mono- and cross-lingual Chinese and Japanese tasks. |
Copied to clipboard
| Challenge: | Experimental results confirm the substantial superiority of GranuSum on multi-granularity summarization over strong baselines. |
| Approach: | They propose to rank events by their salience and annotate a benchmark for GranuSum that contains multiple summaries at different granularities for each document cluster. |
| Outcome: | The proposed framework is capable of producing multi-granular summaries in unsupervised manner over strong baselines. |
Copied to clipboard
| Challenge: | Existing training methods for code generation do not improve code correctness and efficiency. |
| Approach: | They propose a framework that integrates preference learning into code generation to improve code correctness and efficiency. |
| Outcome: | The proposed framework improves code correctness and efficiency by integrating preference learning into code generation. |
Copied to clipboard
| Challenge: | Existing approaches to lexical substitution tend to overlook good substitute candidates that are not the synonyms of the target words in the lexicals and fail to take into account the substitution’s influence on the global context of the sentence. |
| Approach: | They propose an end-to-end BERT-based lexical substitution approach which proposes and validates substitute candidates without using annotated data or manually curated resources. |
| Outcome: | The proposed approach performs well in proposing and ranking substitute candidates, achieving the state-of-the-art results in both LS07 and LS14 benchmarks. |
Copied to clipboard
| Challenge: | Local sequence transduction tasks involve massive overlapping between source and target sequences . experimental results show that Pseudo-Bidirectional Decoding improves performance of standard seq2seq models. |
| Approach: | They propose a simple but versatile approach for local sequence transduction tasks . they propose to copy source tokens to decoder as pseudo future context . |
| Outcome: | The proposed approach improves the performance of standard seq2seq models on LST tasks. |