Papers by Hui Wang
Copied to clipboard
| Challenge: | Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, but academic research remains non-reproducible due to the lack of publicly available training data. |
| Approach: | They propose a system for long-form song generation with fine-grained style conditioning that includes a licensed synthetic dataset and a song generation model, Muse. |
| Outcome: | The proposed system achieves competitive performance on phoneme error rate, text–music style similarity, and audio aesthetic quality while enabling controllable segment-level generation across different musical structures. |
Copied to clipboard
| Challenge: | Existing MRC models are unable to integrate general knowledge with human knowledge. |
| Approach: | They propose a data enrichment method which uses WordNet to extract inter-word semantic connections as general knowledge from each given passage-question pair. |
| Outcome: | The proposed model outperforms state-of-the-art models and is robust to noise. |
Copied to clipboard
| Challenge: | Existing safety controls fail to provide runtime intervention or cross-architecture portability for autonomous LLM agents. |
| Approach: | They propose a model-agnostic, plug-and-play module to provide arbitrary agent safety control and auditability. |
| Outcome: | The proposed module improves the secure-solution rate by 2.9–11.2 percentage points . it adds only 3.2s to end-to-end latency and a negligible average cost of 5.37 10-4 per scenario . |
Copied to clipboard
| Challenge: | Existing benchmarks for Large Language Models (LLMs) follow the data distribution of pre-training data. |
| Approach: | They propose a benchmark ConvRe focusing on converse relations which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets. |
| Outcome: | The proposed benchmark focuses on converse relations, which contains 17 relations and 1240 triples extracted from popular knowledge graph completion datasets. |
Copied to clipboard
| Challenge: | a growing number of cloud-based inference services are relying on SMPC to protect data privacy. |
| Approach: | They propose a framework for Privacy-Preserving Inference for Transformer models that eliminates exponential and maximum operations in PPI without sacrificing model performance. |
| Outcome: | The proposed framework outperforms MPCFormer in terms of performance and efficiency . it is 3.57 and 3.58 times faster than PUMA for BERTBASE and BERTLARGE . |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown impressive results, but still suffer from hallucination, i.e., the generation of false information. |
| Approach: | They propose a task of sequential model editing that aims to rectify mistakes continuously. |
| Outcome: | The proposed method significantly outperforms baselines in single-turn and sequential editing. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) generate intermediate reasoning traces before the final answer, yet they remain vulnerable to reasoning hallucinations such as subtle arithmetic errors. |
| Approach: | They propose a Routing Focus Score (RFS) that measures how strongly cross-step attention routing aligns with semantic proximity derived from hidden-state cosine similarity. |
| Outcome: | The proposed framework detects and localizes hallucinations without external tools or repeated sampling. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models are often prone to parameter interference . a common problem is that the model compromises with the language diversity to find a solution . |
| Approach: | They propose a method that allocates parameters based on consistency between the gradients of the individual language and the average gradient. |
| Outcome: | The proposed method reduces parameter interference and improves translation quality. |
Copied to clipboard
| Challenge: | Existing methods for content-based recommendation with missing or corrupted modalities are lacking in learning multimodal models. |
| Approach: | They propose a multimodal multimodal autoencoder that learns multimodal representations for complementing and imputing missing modalities. |
| Outcome: | The proposed framework achieves state-of-the-art performance on rating prediction tasks and is more robust to previous methods in alleviating data-sparsity and the cold-start problem. |
Copied to clipboard
| Challenge: | Existing DRA methods fail to accurately recover the original text of real-world privacy data. |
| Approach: | They propose to use a real-world privacy dataset to examine the performance of federated learning (FL) methods. |
| Outcome: | The proposed method improves on a real-world privacy dataset and shows that the tokens within a recovery sentence are disordered and intertwined with tokens from other sentences in the same training batch. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) enhance programming by understanding and generating code across languages. |
| Approach: | a new benchmark evaluates code understanding and generation in repositories using code large language models. |
| Outcome: | The proposed model improves code understanding and generation in repositories by evaluating 1,888 test cases across 6 programming languages. |
Copied to clipboard
| Challenge: | General pre-trained language models (PLMs) leverage relation triples from knowledge graphs (KGs) and integrate external data sources into language models via self-supervised learning. |
| Approach: | They propose to learn Knowledge-Enhanced language representations with Hierarchical Reinforcement Learning (KEHRL) to detect positions for knowledge injection and integrate external knowledge into the model to avoid injecting inaccurate or irrelevant knowledge. |
| Outcome: | The proposed model can detect essential positions in texts for knowledge injection and integrate external knowledge into the model to avoid injecting inaccurate or irrelevant knowledge. |
Copied to clipboard
| Challenge: | Existing approaches to personalize large language models (LLMs) rely on heuristic methods to compress user profiles but they ignore how LLMs process and prioritize different profile components. |
| Approach: | They propose an attention-guided context compression framework that leverages attention feedback from a marking model to mark important personalization sentences and guides a compression model to generate task-relevant compressed user contexts. |
| Outcome: | The proposed framework outperforms baselines across tasks, token limits, and settings while reducing token usage by 50 times. |
Copied to clipboard
| Challenge: | Using a pointer-generator framework for reading/sampling over large documents, we propose a framework for learning over long narratives where documents easily span over thousands of tokens. |
| Approach: | They propose a curriculum learning (CL) based pointer-generator framework for reading/sampling over large documents, enabling diverse training of the neural model based on the notion of alternating contextual difficulty. |
| Outcome: | The proposed framework improves on the NarrativeQA reading comprehension benchmark and reaches state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing literature on mechanistic interpretation (MI) treats it as an observational science, leaving practical applications underexplored. |
| Approach: | They propose a survey structured around the pipeline to identify and improve MI models. |
| Outcome: | The proposed framework enables tangible improvements in Alignment, Capability, and Efficiency. |
Copied to clipboard
| Challenge: | recurrent neural networks scale poorly due to the intrinsic difficulty in parallelizing their state computations. |
| Approach: | They propose a simple recurrent unit that provides expressive recurrence and allows highly parallel implementation. |
| Outcome: | The proposed model achieves 5—9x speed-up over cuDNN-optimized LSTM on classification and question answering datasets and delivers stronger results than LS and convolutional models. |
Copied to clipboard
| Challenge: | Existing benchmarks focus primarily on pure graph understanding, lacking a comprehensive evaluation across all graph types and detailed capability definitions. |
| Approach: | They propose a benchmark to evaluate LLMs' graph comprehension and reasoning abilities using a three-tier hierarchical taxonomy and a granular taxonomies. |
| Outcome: | The proposed model includes 11 datasets with 5,140 graphs of varying complexity. |
Copied to clipboard
| Challenge: | a new CLTL model is proposed to facilitate cross-linguistic transfer learning between distant languages . a key to CLTL is to learn a shared representation space for the given source-target language pair. |
| Approach: | They propose a new CLTL model that integrates machine translation with MT . they use an unannotated data technique to make use of the model's pre-training and fine-tuning . |
| Outcome: | The proposed model achieves better CLTL performance than the baseline model without more annotated data. |
Copied to clipboard
| Challenge: | Spatial domain queries have unique properties making them more challenging for language understanding than common conversational queries. |
| Approach: | They propose a language understanding framework for spatial domain queries that jointly learns the intent detection and entity linking tasks on a voice assistant service. |
| Outcome: | The proposed framework outperforms baseline methods with a significant margin. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning. |
| Approach: | They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations. |
| Outcome: | The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion. |
Copied to clipboard
| Challenge: | Existing models for fine-grained speaking styles are limited in terms of accuracy, coverage, and naturalness. |
| Approach: | They propose a model that pre-trains with coarse captions and annotates with a pipeline that grounds captions in audio. |
| Outcome: | The proposed model outperforms existing models with fine-grained style annotations . it integrates global and fine-granular supervision, enabling unified representations based on the proposed model . |
Copied to clipboard
| Challenge: | Existing methods for AVE are limited on rare attributes due to poor generalization ability. |
| Approach: | They propose to leverage pretraining and transfer learning to address weaknesses in existing methods. |
| Outcome: | The proposed method achieves new state-of-the-art performance without pretraining on rare attributes with limited training resources. |
Copied to clipboard
| Challenge: | Existing methods for misinformation detection are limited by domain knowledge and expert experience. |
| Approach: | They propose a Multi-Agent Framework for cross-domain misinformation detection with Automated Decision Rule Optimization (MARO) they first employ multiple expert agents to analyze target-domain news, then introduce a question-reflection mechanism that guides expert agents for higher-quality analysis. |
| Outcome: | The proposed framework improves on a common dataset and shows that iteratively improves over existing methods. |
Copied to clipboard
| Challenge: | Automatic speech recognition systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0. |
| Approach: | They propose to use Mandarin speech datasets to analyze pronunciation and tone of children aged 3 to 5 and evaluate their models on speaker verification (SV) They find that the datasets are more robust than those used by adult speech recognition systems and are open-source and available for all academic purposes. |
| Outcome: | The proposed dataset includes 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. |
Copied to clipboard
| Challenge: | Existing text embeddings with high dimensions are difficult to trace and interpret. |
| Approach: | They propose low-dimensional and interpretable text embeddings with relative representations that encode semantic meanings in a vector space where similar texts are close together in the representation space. |
| Outcome: | The proposed embeddings outperform existing models on multiple tasks with fewer dimensions and are lowdimensional and dense while maintaining interpretability. |
Copied to clipboard
| Challenge: | Existing methods for incorporating external knowledge into language models do not prioritize learning embeddings for entity-related tokens. |
| Approach: | They propose a framework for incorporating external knowledge into pre-training models that utilize entity-related tokens. |
| Outcome: | The proposed framework reduces pre-training time by 50% and outperforms other KEPLMs in knowledge probing tasks and multiple knowledge-aware language understanding tasks. |
Copied to clipboard
| Challenge: | Existing methods for aspect-based sentiment analysis (ABSA) only compare current predictions and labels on each sample, yet fail to perceive and understand its error outputs from different degrees. |
| Approach: | They propose a framework that can perceive and understand the degree of errors by learning from comparative error pairs. |
| Outcome: | The proposed framework exceeds baselines and achieves the desired performance. |
Copied to clipboard
| Challenge: | Current mathematical benchmarks focus on evaluating MLLMs’ problem-solving ability, yet there is a crucial gap in addressing more complex scenarios such as error detection. |
| Approach: | They propose to evaluate multimodal error detection by evaluating two sub-tasks error step identification and error categorization. |
| Outcome: | The proposed task evaluates MLLMs' ability to handle multimodal questions compared to text-only models. |
Copied to clipboard
| Challenge: | Learned data compression has achieved superior compression ratios, but balancing precise probability modeling with system efficiency remains challenging. |
| Approach: | They propose a Dual-Stream Multi-Scale Decoupler that disentangles local and global contexts to replace deep serial processing with shallow parallel streams. |
| Outcome: | The proposed method achieves state-of-the-art performance in both compression ratio and throughput while maintaining the lowest latency and memory usage. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated better safety performance in high-resource languages than in low-resourced languages. |
| Approach: | They propose language-agnostic semantic alignment (LASA) which anchors safety alignment directly in semantic bottlenecks. |
| Outcome: | The proposed approach significantly improves safety across all languages: average attack success rate drops from 24.7% to 2.8% on LLaMA-3.1-8B-Instruct and remains within 3–4% across Qwen2.5 and Qwend3 Instruct models (7B–32B). |
Copied to clipboard
| Challenge: | In encoder-decoder neural models, multiple encoders are used to represent contextual information in addition to the individual sentence. |
| Approach: | They propose to use multiple context encoders to encode the individual sentences in document-level neural machine translation (NMT) They propose a noisy dropout setup and a single-encoder approach to encode context sentences. |
| Outcome: | The proposed approach encodes the context and the current sentence without contexts. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) models focus on word-level information, while segment-based models focus only on word level information. |
| Approach: | They propose a Modularized Interaction Network (MIN) model which utilizes both word-level information and segment-level dependencies. |
| Outcome: | The proposed model outperforms the current state-of-the-art models on three NER benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for error identification often overlook validation of generated results . text-to-SQL is a technology that converts natural language questions into executable SQL queries . |
| Approach: | They propose to integrate a multi-grained error identification method into existing methods to detect SQL errors. |
| Outcome: | The proposed method can be integrated as a plugin into various methods, providing effective error identification and correction capabilities. |
Copied to clipboard
| Challenge: | Recent large reasoning models (LLMs) lack dynamic and diverse thinking capabilities . reusing atomic thoughts provides a practical pathway toward dynamic reasoning . |
| Approach: | They propose a framework that extracts atomic thoughts from teacher models and reuses them to guide reasoning and generate responses. |
| Outcome: | The proposed framework extracts atomic thoughts from teacher models and reuses them to guide reasoning and generate responses. |
Copied to clipboard
| Challenge: | Existing research shows LLMs struggle with complex instructions involving multiple constraints. |
| Approach: | They propose a framework to divide complex instructions into single constraints and prepare appropriate tools to verify responses. |
| Outcome: | The proposed framework doubles Llama3.1-8B’s constraint adherence and triples Mistral-7B’ s performance. |
Copied to clipboard
| Challenge: | Existing neural IR models do not have a mechanism for treating expansion terms differently from the original query terms, making it difficult to combine them with existing PRF approaches. |
| Approach: | They propose an end-to-end neural PRF framework that can be used with existing neural IR models by embedding different neural models as building blocks. |
| Outcome: | Extensive experiments on two standard test collections confirm the effectiveness of the proposed framework in improving the performance of two state-of-the-art neural IR models. |
Copied to clipboard
| Challenge: | Existing methods for training reasoning-oriented large language models assume high-resource settings with abundant data. |
| Approach: | They propose a framework that integrates high-value general-domain data to promote more diverse exploration. |
| Outcome: | The proposed framework matches or surpasses RLVR trained with 32 target-domain samples using 32 target domain samples. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) reasoning has improved the performance of large language models (LLMs) however, the detailed reasoning process in CoT often incurs long generation times and high computational costs due to the inclusion of unnecessary steps. |
| Approach: | They propose a method to identify critical reasoning steps using perplexity as a measure of their importance. |
| Outcome: | The proposed method achieves a better balance between reasoning accuracy and efficiency of CoT. |
Copied to clipboard
| Challenge: | Existing pointwise LLMs provide noisy or biased answers for documents that are partially relevant to the query. |
| Approach: | They propose to incorporate fine-grained relevance labels into the LLM prompt . they propose to better differentiate between documents with different levels of relevance . |
| Outcome: | The proposed model can differentiate between documents with different levels of relevance to the query and derive a more accurate ranking. |
Copied to clipboard
| Challenge: | Existing methods to mask and predict tokens in multilingual text limit multilingual interaction . |
| Approach: | They propose a lifelong multilingual multi-granularity semantic alignment approach which continuously extracts massive aligned linguistic units from noisy data via a maximum co-occurrence probability algorithm. |
| Outcome: | The proposed approach improves translation performance on WMT14 18 benchmarks in twelve directions. |
Copied to clipboard
| Challenge: | Efficient instruction tuning aims to enhance the ultimate performance of large language models (LLMs) current methods suffer from the curriculum rigidity, resulting in a fixed and potentially sub-optimal learning trajectory. |
| Approach: | a framework for efficient instruction tuning is proposed to address the issue of curriculum rigidity . current methods rely on static heuristic difficulty metrics and fail to adapt to evolving capabilities . |
| Outcome: | Efficient instruction tuning aims to enhance the ultimate performance of large language models . current methods suffer from the curriculum rigidity, resulting in a fixed learning trajectory . |
Copied to clipboard
| Challenge: | Existing solutions to alleviate hallucination have considered utilizing LLMs’ inherent reasoning abilities to alleviating hallucinism, such as self-correction and diverse sampling methods. |
| Approach: | They propose a counterfactual multi-agent debate framework that predetermines LLMs' stances to override their inherent biases for answer inspection. |
| Outcome: | Extensive experiments on four datasets of three tasks demonstrate the superiority of the proposed framework over existing methods. |
Copied to clipboard
| Challenge: | Existing efficient reasoning methods rely on explicit length penalties for excessive verbosity on simple queries. |
| Approach: | They propose a training-time intervention that selectively suppresses redundant tokens . they find length shift occurs when models generate unnecessary reasoning on trivial inputs - a phenomenon that is often unexplored . |
| Outcome: | The proposed method reduces inference token usage by 78% while increasing accuracy compared to the initial policy and surpasses state-of-the-art efficient reasoning methods. |
Copied to clipboard
| Challenge: | Existing methods require full-modality data during training phase or require explicit annotations to detect missing modalities. |
| Approach: | They propose a Dynamic modality Recognition and Enhancement for Adaptive Multimodal fusion framework that directs selective reconstruction of missing or underperforming modalities. |
| Outcome: | The proposed framework outperforms several baseline and state-of-the-art models on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing studies on large language models (LLMs) focus on basic plan validity, but neglect critical aspects such as route efficiency, POI appeal, and real-time adaptability. |
| Approach: | They propose a benchmark for retrieval-augmented, spatiotemporal-aware travel planning that integrates retrieved trajectories with LLMs’ intrinsic reasoning. |
| Outcome: | The proposed framework improves spatial efficiency and POI rationality while challenging universality and robustness due to conflicting references and noisy data. |
Copied to clipboard
| Challenge: | Existing methods for event argument extraction (EAE) lack cross-event information and require longer role sequences . et al. (2017): outperforms state-of-the-art methods for EE. |
| Approach: | They propose a separation-and-fusion paradigm to separate the acquisition of cross-event information and fuse it into the argument extraction of a target event. |
| Outcome: | The proposed model outperforms the state-of-the-art models on four widely used datasets. |
Copied to clipboard
| Challenge: | Existing models rank statements solely by confidence scores, and there is no information about which ones are salient from a human perspective. |
| Approach: | They propose a task where a model is required to learn whether a triple is salient . they propose supervised salience evaluation using a new Benchmark dataset . |
| Outcome: | The proposed task is based on a new Benchmark dataset of salience evaluation in e-commerce . it shows that saliency evaluation is hard, where models perform poorly on evaluation set . |
Copied to clipboard
| Challenge: | Gradient Ascent (GA) has emerged as a promising approach for concept unlearning in Multimodal Generative Models (MGMs). |
| Approach: | They propose a novel approach that selectively applies GA to targeted Conceptual Knowledge while preserving Natural Knowledge through Gradient Descent (GD). |
| Outcome: | The proposed approach removes Conceptual Knowledge and inadvertently diminishes Natural Knowledge, resulting in utility degradation. |
Copied to clipboard
| Challenge: | Speech signals convey abundant speaker-related metadata, yet current privacy research focuses on identity-centric voiceprint protection, leaving sensitive Speaker Attribute Privacy (SAP) underexplored. |
| Approach: | They propose a large-scale Chinese dataset to evaluate speaker-related privacy leakage . the dataset includes 227.3 hours of audio from 1,000 speakers . |
| Outcome: | The proposed model systematically evaluates speaker-related privacy leakage in everyday scenarios. |
Copied to clipboard
| Challenge: | a multimodal protein language model (LLM) integrates sequence, structure, and function into functional annotation. |
| Approach: | They propose a multimodal protein language model that synergistically aligns bimodal representations with the textual modality to advance protein functional annotation. |
| Outcome: | The proposed model synergizes bimodal representations with the textual modality to advance protein functional annotation. |
Copied to clipboard
| Challenge: | Existing vision-language models focus on salient attributes but ignore contextualized nuances, resulting in gender bias. |
| Approach: | They propose a task-agnostic generation framework to mitigate gender bias in vision-language models. |
| Outcome: | The proposed framework can mitigate gender bias in vision-language models . it yields all-sided but gender-obfuscated narratives, which prevents concentration on localized image features, especially gender attributes. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) is effective for aligning Large Language Models with human preferences, but its complex process limits its ability to continually learn human feedback. |
| Approach: | They propose a non-RL offline method to convert historical optimal policies into optimization constraints when continually learning new preferences. |
| Outcome: | The proposed method outperforms strong CL baselines in terms of reward-based evaluations and human assessment. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in complex reasoning through long chain-of-thought, yet they struggle with precise computations and algorithmic operations. |
| Approach: | They propose a training-free approach that activates LRMs’ latent tool-use capabilities through artificial hints and a framework that enables models to learn effective tool utilization through diverse hint patterns and rejection-based data synthesis. |
| Outcome: | Experiments show that START significantly improves state-of-the-art LRMs across challenging benchmarks, including competition-level mathematics (AMC23: 95.0%, AIME24: 75.6%) and graduate-level science questions (GPQA: 64.6%). |
Copied to clipboard
| Challenge: | Existing studies on LLM agents' social behaviors are lacking . previous studies focused on positive social behaviors, leaving research on negative social behaviors relatively scarce. |
| Approach: | They propose a framework that features a multi-agent system facilitating efficient communication and interaction with LLM agents. |
| Outcome: | The proposed framework is based on Avalon and evaluates on game success and analyzes agents’ social behaviors. |
Copied to clipboard
| Challenge: | Existing methods for solving complex problems are expensive and inefficient when handling large-scale, high-complexity problems. |
| Approach: | They propose a multi-agent framework that decomposes complex problems through agent collaboration by mapping implicitly expressed graph data into clear, structured graph representations and dynamically selecting the most suitable algorithm based on problem constraints and graph structure scale. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on multiple benchmarks with robust performance on both closed- and open-source models. |
Copied to clipboard
| Challenge: | Large language models struggle with complex reasoning tasks, such as mathematical problem-solving. |
| Approach: | They constructed a symbolic multi-step reasoning task to investigate the information propagation mechanisms in Transformer models when solving the task through direct answering and Chain-of-Thought (CoT) reasoning. |
| Outcome: | The proposed algorithm improves on 7 multi-step reasoning datasets, while introducing only 132 trainable parameters. |
Copied to clipboard
| Challenge: | Existing approaches to improve retrieval accuracy and generation quality of large language models suffer from language preference. |
| Approach: | They propose a framework that explicitly disentangles multilingual RAG into language-controllable retrieval and language-agnostic reasoning. |
| Outcome: | Experimental results show that the proposed approach outperforms baselines across multilingual benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to correct outdated or erroneous knowledge in large language models (LLMs) are slow and cumbersome, resulting in catastrophic knowledge forgetting and degradation of model performance. |
| Approach: | They propose a RetriEval-augmented ContInuous Prompt lEarning method that converts knowledge statements into short and informative continuous prompts, prefixed to the LLM’s input query embedding. |
| Outcome: | The proposed method improves the performance of large language models (LLMs) while maintaining the overall performance of the model. |
Copied to clipboard
| Challenge: | Existing approaches to remove copyrighted and privacy-sensitive data from Large Language Models (LLMs) have been proposed to remove specific data from LLMs without requiring full retraining. |
| Approach: | They propose a general framework that enhances the utility of fine-tuning-based methods by distinguishing target data and suppressing related generations. |
| Outcome: | The proposed framework improves the unlearning and utility of fine-tuning-based methods by distinguishing the target data and suppressing related generations. |
Copied to clipboard
| Challenge: | Existing methods that learn from multiple semantically-equivalent questions are limited to one-to-one mapping . |
| Approach: | They propose a constraint to explore the underlying complementary semantic information among multiple semantically-equivalent questions and learn robust feature representations with reduced spurious associations. |
| Outcome: | The proposed method outperforms strong competitors and achieves state-of-the-art results on five benchmark datasets. |
Copied to clipboard
| Challenge: | Automatic speech recognition (ASR) for children remains challenging due to developmental variability and the scarcity of high-quality corpora. |
| Approach: | They propose a large-scale Chinese child speech corpus that contains 112.5 hours of speech from 498 children and 500 caregivers. |
| Outcome: | The proposed model improves in-domain and cross-domain performance on children's speech. |
Copied to clipboard
| Challenge: | Existing systems for conversational recommender systems (CRS) have strong results in movies, but games present distinct challenges . MATCHA framework provides specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking, and stronger safety. |
| Approach: | They propose a framework for conversational recommender systems that assigns specialized agents for intent parsing, tool-augmented retrieval, multi-LLM ranking and risk control. |
| Outcome: | MATCHA outperforms baselines on real user request dataset, improves Hit@5 by 20%, reduces popularity bias by 24%, and achieves 97.9% adversarial defense. |
Copied to clipboard
| Challenge: | Existing routers generalize poorly in cold-start scenarios where in-domain training data is unavailable. |
| Approach: | They propose a task-type–aware router approach that models query-conditioned cost and performance via latent task-like variables with prior regularization derived from the synthesized task taxonomy. |
| Outcome: | The proposed framework improves performance and cost under cold-start and in-domain settings and enables efficient routing. |
Copied to clipboard
| Challenge: | Existing web agents struggle with complex tasks due to rigid planning strategies and hallucination-prone reasoning. |
| Approach: | They propose a task-uncertainty-driven Adaptive Planning Mechanism that adaptively selects planning modes to navigate unknown environments. |
| Outcome: | The proposed framework performs better on the WebArena and WebVoyager benchmarks than existing frameworks. |
Copied to clipboard
| Challenge: | In the industry, numerous natural language processing tasks are deployed online . traditional approaches tackle each task separately by its own network and pipeline . |
| Approach: | They propose a three-stage multi-task learning framework for large language models . it involves task filtering, fine-tuning on high-resource tasks, and finally fine- tuning on all tasks . |
| Outcome: | The proposed framework reduces up to 90% of overhead while reducing latency and resource usage. |
Copied to clipboard
| Challenge: | Existing approaches to identifying capabilities rely on external signals with limited structural grounding . emergence of specific capabilities remains poorly understood . |
| Approach: | They propose a lightweight approach that links LLM capabilities to internal components by identifying correspondences at the level of attention heads. |
| Outcome: | The proposed approach improves accuracy on MMLU and BBH by 1 to 1.5 points over gradient-based method and 5 to 6 points over other intermediate-state baselines. |
Copied to clipboard
| Challenge: | Lossless compression has made significant advancements in Genomics Data storage, sharing and management. |
| Approach: | They propose a novel agent-based GD Compressor with 3 layers with a multi-agent named Leader and Worker. |
| Outcome: | The proposed method improves on existing methods with low-level modeling and limited adaptability and user-unfriendly interface. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models struggle with dynamic interactions due to the scarcity of high-quality interleaved data. |
| Approach: | They propose a large-scale interleaved live interaction Chinese dataset with human-annotated video responses. |
| Outcome: | The proposed model can be used to evaluate live interactions in Chinese over 1,100 hours and 80,037 dialogue turns. |
Copied to clipboard
| Challenge: | Existing methods that optimize for scalar scores or ranking reward ignore multi-dimensional nature of human preferences. |
| Approach: | They propose to extend the preference of Direct Preference Optimization to two dimensions: segments and aspects. |
| Outcome: | The proposed framework decomposes the overall objective into multi-segment and multi-aspect objectives. |
Copied to clipboard
| Challenge: | Existing RAG methods focus on improving the task performance, without fine-grained process of knowledge. |
| Approach: | They propose a method that detects long-tail knowledge in large language models by analyzing retrieved documents and enhancing queries indiscriminately with retrieved information. |
| Outcome: | The proposed method achieves over 4x speedup in average inference time and consistent performance improvement in downstream tasks compared to existing pipelines. |
Copied to clipboard
| Challenge: | Recent advances in graph neural networks have made it difficult to capture user preferences. |
| Approach: | They propose a graph contrastive learning recommendation model based on noise augmentation that integrates truncated singular value decomposition in the feature engineering stage. |
| Outcome: | The proposed model reduces dimensionality and denoises the original data. |
Copied to clipboard
| Challenge: | Existing methods for detecting fake news are limited due to non-transparent reasoning processes and inherent risks of integration with large language models. |
| Approach: | They propose a framework for trustworthy fake news detection that prioritizes explainability, generalizability and controllability of models. |
| Outcome: | The proposed framework prioritizes explainability, generalizability and controllability of models. |
Copied to clipboard
| Challenge: | Mainstream speaker diarization systems rely only on acoustic information, making it challenging in complex aural environments. |
| Approach: | They propose a multimodal approach that integrates audio, visual, and semantic cues to enhance speaker diarization. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on multi-party conversations . it integrates audio-visual-semantic cues into the clustering process for acoustic speaker embeddings . |
Copied to clipboard
| Challenge: | federated learning (FL) is a promising technique for preserving data privacy . however, there is no work on applying FL to legal NLP . |
| Approach: | They propose to use federated learning to train models in a collaborative way without sharing data . they propose to test the FL benchmark on real-world legal data from Chinese courts . |
| Outcome: | The proposed benchmark combines five legal NLP tasks and one privacy task on Chinese courts. |
Copied to clipboard
| Challenge: | Unsupervised neural machine translation methods have been observed to make particular errors in comparison to supervised machine translation, such as confusing nouns that pertain to the same semantic category. |
| Approach: | They propose a method that incorporates images at the word level to augment lexical mappings. |
| Outcome: | Experiments on a multi-lingual dataset show that the proposed method generates more accurate translations with only monolingual data. |
Copied to clipboard
| Challenge: | Existing studies on how SAEs derive most fine-grained latent features for safety remain unexplored. |
| Approach: | They propose a framework for interpreting SAE features in safety-critical domains . they train a suite of SAEs with human-readable explanations and systematic evaluations based on pornography, politics, violence, and terror . |
| Outcome: | The proposed framework reduces interpretation cost by 55% and improves safety-critical features. |
Copied to clipboard
| Challenge: | E-commerce search relevance is a critical component of retrieval systems. |
| Approach: | They propose a large-generative model for search relevance that trains reasoning knowledge, multi-modal understanding and rule awareness into three core competencies. |
| Outcome: | The proposed model outperforms GPT-5 in Macro-F1 and achieves 27% online gain. |
Copied to clipboard
| Challenge: | End-to-end sign language translation (SLT) aims to convert sign language videos into spoken language texts without intermediate representations. |
| Approach: | They propose a cross-modality data-augmented framework to transfer gloss-to-text translation capabilities to end-to end sign language translation. |
| Outcome: | The proposed framework outperforms baseline models on two widely used SLT datasets. |
Copied to clipboard
| Challenge: | Existing methods rely on one-hot discriminative supervision, leading to overfitting on seen classes and poor generalization to unseen ones. |
| Approach: | They propose a Generative–Discriminative Dual-View Co-Training framework that unifies discriminative classification and semantic label generation within an LLM. |
| Outcome: | The proposed framework outperforms existing methods on five benchmarks on the generalized category discovery (GCD) task. |
Copied to clipboard
| Challenge: | Existing word embedding methods for Mongolian PB prediction are expensive and time-consuming. |
| Approach: | They propose to use Mongolian word embedding to build a robust Mongolian PB prediction model . they encode sub-word units and feed it to LSTM to decode the best corresponding PB label . |
| Outcome: | The proposed model outperforms traditional model using manual features and achieves 7.49% gain. |
Copied to clipboard
| Challenge: | Rapid advances in Large Language Models have spurred demand for processing extended context sequences . however, performance degradation due to sequence lengths out-of-distribution and excessively long inference times are limiting LLMs in long-context scenarios. |
| Approach: | They propose a training-free method for efficient and accurate long-context inference . they selectively involves a few critical KV cache tokens in attention calculation . |
| Outcome: | The proposed method speeds up attention computation and accelerates inference time while reducing selection overhead. |
Copied to clipboard
| Challenge: | Existing approaches to verify agent behaviors in complex environments rely on rule-based verifiers or LLM-as-a-Judge models. |
| Approach: | They propose a benchmark to evaluate Agent-as-a-Judge across three domains . the benchmark covers search, data systems, and graphical user interfaces - with 155 tasks and 516 trajectories . |
| Outcome: | The proposed benchmark outperforms existing benchmarks in search, data systems, and GUI domains while revealing open challenges in agent-based verification. |
Copied to clipboard
| Challenge: | RLVR is a paradigm for improving reasoning ability of large language models . but voting results often induce confirmation bias and suffer from sparse rewards . |
| Approach: | They propose a framework integrating model confidence and dynamic subgroup partitioning to address these issues. |
| Outcome: | The proposed framework outperforms recent baselines on multiple models and benchmarks. |
Copied to clipboard
| Challenge: | Recent studies have focused on classifying cardiac conditions using ECG data but have overlooked ECG report generation, which is time-consuming and requires clinical expertise. |
| Approach: | They propose a Multimodal ECG Instruction Tuning framework that extends the capability of large language models (LLMs) for the task. |
| Outcome: | The proposed framework outperforms open-source LLMs and LLM backbones across two large-scale ECG datasets. |
Copied to clipboard
| Challenge: | Existing work on 3D radiograph report generation focuses on 2D images, but 3D medical images provide more comprehensive diagnostic information. |
| Approach: | They propose a comprehensive training recipe for building high-performing VLMs for 3DRRG using a publicly available 3D CT-report dataset. |
| Outcome: | The proposed model achieves superior performance across different model sizes and input 3D medical image resolutions. |
Copied to clipboard
| Challenge: | State-of-the-art (SOTA) methods use the cross-encoder architecture to concatenate a mention (and its context) with each type and feed it into a pretrained language model (PLM) to score their relevance. |
| Approach: | They propose to perform entity typing in a recall-expand-filter manner and use a novel model to encode and score all these K candidates in one forward pass. |
| Outcome: | The proposed method is thousands of times faster than the CE-based architecture and is very efficient in fine-grained (130 types) and coarse-grain (9 types) entity typing. |
Copied to clipboard
| Challenge: | Sarcasm is a linguistic phenomenon indicating a discrepancy between literal meanings and implied intentions. |
| Approach: | They propose a hierarchical framework for sarcasm detection by exploring atomic-level congruity and composition-level convergence. |
| Outcome: | The proposed model outperforms existing methods on a public sarcasm detection dataset based on Twitter . |
Copied to clipboard
| Challenge: | Long-context Multimodal Large Language Models (MLLMs) require substantial computational resources as their multimodal Key-Value (KV) cache grows with increasing input lengths, challenging memory and time efficiency. |
| Approach: | They propose a dynamic multimodal KV cache allocation strategy that dynamically allocating KV size based on attention entropy to better adapt to multimodal interactions. |
| Outcome: | The proposed model achieves up to 72% KV cache memory reduction and 2.82 faster decoding speeds while maintaining or enhancing performance on various multimodal tasks in a long context. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated superior language understanding abilities in many real-world NLP applications. |
| Approach: | They propose a learning-based sample selection method that incorporates signals from both teacher and student to enhance model performance. |
| Outcome: | The proposed method improves model performance across datasets with higher data efficiency. |
Copied to clipboard
| Challenge: | Large language models (LLMs) with one or more fine-tuning phases can unlock various capabilities, but can be catastrophic forgetting during sequential training. |
| Approach: | They propose a method to regularly reset partial parameters to mitigate forgetting issues by using half fine-tuning instead of full fine-uning. |
| Outcome: | The proposed approach reduces the risk of catastrophic forgetting during training and the parametric knowledge lost during training may be overwhelmed by incoming training data. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing privacy and copyright concerns. |
| Approach: | They propose a framework that learns a universal noise pattern to recover unlearned information from MLLMs. |
| Outcome: | The proposed framework learns a universal noise pattern and can reveal unlearned content when applied to images. |
Copied to clipboard
| Challenge: | Existing approaches for explainable machine learning systems focus on interpreting outputs or connections between inputs and outputs. |
| Approach: | They propose a generative explanation framework that learns to make classification decisions and generates fine-grained explanations at the same time. |
| Outcome: | The proposed framework surpasses all baselines on two datasets and generates concise explanations at the same time. |
Copied to clipboard
| Challenge: | Existing memory systems rely on summarization to preserve contextual nuances and obscuring key retrieval features. |
| Approach: | They propose a method that decouples the retrieval unit from the generation context. |
| Outcome: | The proposed method outperforms baseline models on the LoCoMo benchmark. |
Copied to clipboard
| Challenge: | Existing models for natural language processing are heavily parameterized and memory inefficient. |
| Approach: | They propose a series of lightweight and memory efficient neural architectures for NLP tasks . they propose quaternion algebra and hypercomplex spaces for computation . |
| Outcome: | The proposed models enable expressive inter-component interactions and significantly reduce parameter size without loss of performance. |
Copied to clipboard
| Challenge: | Existing datasets face issues such as low quality, limited scale, and incomplete modalities, hindering model performance. |
| Approach: | They propose to use Chinese multimodal datasets to capture authentic emotional interplay from 19 professional actors. |
| Outcome: | The EmotionTalk dataset spans 23.6 hours of dyadic conversations across diverse scenarios. |
Copied to clipboard
| Challenge: | Existing privacy-preserving Transformer Inference frameworks suffer from high computational overhead and performance losses. |
| Approach: | They propose a framework that integrates random permutations and SMPC to address the "impossible trinity" CENTAUR resists diverse data reconstruction attacks and boosts inference speed by 5.030.4 times . |
| Outcome: | CENTAUR achieves an unprecedented balance between privacy, efficiency, and performance. |
Copied to clipboard
| Challenge: | Goal-oriented script planning is used by humans to plan for typical activities . however, this capability remains underexplored due to several challenges . |
| Approach: | They propose a framework that enables product-enriched scripts by associating products with each step based on the semantic similarity between the actions and their purchase intentions. |
| Outcome: | The proposed framework can generate product-enriched scripts from 2.4 million scripts . human annotations are conducted to provide gold labels for a sampled subset . |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) combines the language understanding and reasoning capabilities of large language models (LLMs) with external retrieval to produce domain-grounded responses. |
| Approach: | They propose a scalable and modular data-centric framework for generating domain-grounded question–answer–context triples tailored to diverse RAG adaptation strategies. |
| Outcome: | The proposed framework generates domain-grounded question–answer–context triples for multiple RAG adaptation strategies. |
Copied to clipboard
| Challenge: | Existing majority voting methods generate only a single answer in each trial, ignoring the possibility of other possible answers. |
| Approach: | They propose to generate ranked answers in each reasoning process and conduct ranked voting among multiple ranked responses from different responses. |
| Outcome: | Extensive experiments show that the proposed method outperforms baselines on multiple-choice and open-ended questions. |
Copied to clipboard
| Challenge: | Existing methods to rank documents using large language models do not understand these challenging ranking formulations. |
| Approach: | They propose to use Pairwise Ranking Prompting to improve ranking performance . they propose to outperform fine-tuned baseline rankers on benchmark datasets . |
| Outcome: | The proposed technique outperforms supervised baselines on benchmark datasets and outperformed other LLM-based solutions by over 10% on average. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data. |
| Approach: | They propose a method to disentangle risks through step-by-step reasoning within multimodal inputs. |
| Outcome: | The proposed approach improves safety alignment in MLLMs by fine-tuning and iterative Reinforcement Learning from AI feedback. |
Copied to clipboard
| Challenge: | Existing agent benchmarks neglect hierarchical rule application in real-world domains . a critical gap persists in numerous real-life professional domains where decision-making is governed by expert-written rules. |
| Approach: | They propose a benchmark requiring agents to assign a unique 10-digit Harmonized System (HS) Code to products by aligning their fuzzy attributes with strict tariff classification rules. |
| Outcome: | The proposed benchmarks lack hierarchical rule application capability in real-world domains . the proposed benchmark is based on e-commerce and is open-source . |
Copied to clipboard
| Challenge: | Existing studies on the sensitivity of Large Language Models (LLMs) to irrelevant contexts neglect the importance of key information. |
| Approach: | They investigate the sensitivity of large language models to key medical information by introducing different perturbation strategies to investigate their sensitivity. |
| Outcome: | The proposed models are based on three LLMs, namely GPT-3.5, GPT-4, Gemini, Claude3 and LLaMA2-7b, and demonstrate their reliability and sensitivity to medical information. |
Copied to clipboard
| Challenge: | Recent learning-based demonstration selection methods have proven beneficial to in-context learning (ICL) by choosing more useful exemplars. |
| Approach: | They propose two methods to capture task-agnostic similarities between input and output of LLMs. |
| Outcome: | The proposed methods integrate task-agnostic similarities of different levels between input and output of exemplars and test cases to eliminate costly data collection. |
Copied to clipboard
| Challenge: | Large language models extract useful information from conversation history to enhance the response in long-term conversations. |
| Approach: | They propose a Fragment-then-Compose framework to optimize memory utilization for long-term open-domain conversation. |
| Outcome: | The proposed framework can be used to extract useful information from conversation history . it can be adapted to different situations and improve response generation . |
Copied to clipboard
| Challenge: | Experimental results show that deep training is 1:4 faster than training from scratch. |
| Approach: | They propose a shallow-to-deep training method that learns deep models by stacking shallow models. |
| Outcome: | The proposed method is 1:4 faster than training from scratch and achieves BLEU scores of 30:33 and 43:29 on two translation tasks. |
Copied to clipboard
| Challenge: | Extensive experiments on six real-world datasets show our approach outperforms the best baselines by 7.33% in NDCG@10, 4.65% in Recall@10 and 8.42% in MRR. |
| Approach: | They propose a framework for mapping sequential item texts to sequential item IDs that incorporates multi-query input and item linear projection to model conditional probability distribution of items. |
| Outcome: | The proposed framework outperforms baseline models on six real-world datasets by 7.33% and 4.65% respectively. |
Copied to clipboard
| Challenge: | Existing sentence embedding methods rely on fixed prompt templates or involve modifications to the model architecture, compromising its generative capabilities. |
| Approach: | They propose a sentence-level direct preference optimization approach that boosts the sentence representations while preserving the generative ability of LLMs. |
| Outcome: | The proposed method improves representations of semantically meaningful vectors without sacrificing generation capability. |
Copied to clipboard
| Challenge: | Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. |
| Approach: | They propose a Monte Carlo Tree Search (MCTS) based Prompt Auto-generation (MPA) method to generate adversarial suffixes for valid jailbreak attacks. |
| Outcome: | The proposed method outperforms existing methods on open-source and closed-source models and shows that it can generate harmful responses. |
Copied to clipboard
| Challenge: | Current named entity recognition methods struggle with text-image mismatch problem due to a lack of visual context. |
| Approach: | They propose an adaptive mixup image augmentation method that generates augmented images based on matching score between text and image . |
| Outcome: | The proposed method can be integrated into existing models and demonstrate consistent performance improvements. |
Copied to clipboard
| Challenge: | Recent research shows that dialogue systems trained on human conversation data are biased and can produce responses that reflect people’s gender prejudice. |
| Approach: | They propose a novel adversarial learning framework Debiased-Chat to train dialogue models free from gender bias while keeping their performance. |
| Outcome: | The proposed framework significantly reduces gender bias in dialogue models while maintaining the response quality. |
Copied to clipboard
| Challenge: | Existing graph-based encoders for text-to-SQL do not model the syntax of natural language questions. |
| Approach: | They propose to inject Syntax to question-Schema graph encoder for text-to-SQL parsers and employ the decoupling constraint to induce diverse relational edge embedding. |
| Outcome: | The proposed approach outperforms all existing methods when pre-training models are used, resulting in a performance ranks first on the Spider leaderboard. |
Copied to clipboard
| Challenge: | Pre-trained models have not been used to outperform other deep learning models such as CNN in Automated Essay Scoring (AES). |
| Approach: | They propose a novel multi-scale essay representation for BERT that can be jointly learned . they employ multiple losses and transfer learning from out-of-domain essays to further improve performance . |
| Outcome: | The proposed model outperforms existing models in the area of automated essay scoring . the proposed model generalizes well to the CommonLit Readability Prize data set . |
Copied to clipboard
| Challenge: | despite its linguistic significance, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. |
| Approach: | They propose to use WenetSpeech-Wu as a large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect of Chinese. |
| Outcome: | The proposed dataset includes 8,000 hours of speech data and strong open-source models . the proposed dataset is competitive and empirically validated . |
Copied to clipboard
| Challenge: | Multi-agent LLMs generate multiple candidate responses that are aggregated by an LLM judge. |
| Approach: | They propose to advocate KV cache reuse across partially shared contexts and report substantial speedups for generation agents. |
| Outcome: | The proposed reuse strategies weaken cross-candidate attention, especially for later candidate blocks, and highlight judge-centric inference as a distinct regime that requires dedicated, risk-aware system design. |
Copied to clipboard
| Challenge: | Current frontier models sometimes generate false outputs or answers that are not substantiated by evidence. |
| Approach: | They propose Chinese SimpleQA, a Chinese benchmark to evaluate LLMs' factuality . they focus on Chinese language over 6 major topics with 99 diverse subtopics . |
| Outcome: | The Chinese SimpleQA benchmark evaluates the factuality ability of LLMs . the questions and answers are short and easy-to-evaluate . |
Copied to clipboard
| Challenge: | Existing approaches to cross-document relation extraction (RE) focus on identifying relations between head and tail entities from single sentence or document. |
| Approach: | They propose a hierarchical relation tree-based LLM-based hierarchic classification model for cross-document relation extraction (HCRE) based on predefined relations, the model can perform hierarchically classification level by level. |
| Outcome: | The proposed model outperforms existing baselines and validates its effectiveness. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made strong progress in reasoning. |
| Approach: | They propose a dual-phase test-time scaling framework that separates planning and execution and performs search over each phase independently. |
| Outcome: | Experiments on math reasoning and code generation benchmarks show that the proposed approach improves accuracy while reducing redundant computation. |
Copied to clipboard
| Challenge: | Existing studies show that large language models can be instructed to perform zero-shot passage re-ranking . Existing work like UPR demonstrate promising results for zero- shot ranking using LLMs . |
| Approach: | They propose a demonstration selection strategy based on difficulty rather than semantic similarity . they propose to include only one demonstration in the prompt to improve re-ranking . |
| Outcome: | The proposed method improves LLM-based re-ranking by adding one demonstration to the prompt. |
Copied to clipboard
| Challenge: | Recent multimodal information extraction approaches overestimate the significance of images. |
| Approach: | They propose a general data splitting strategy to divide social media posts into two sets to achieve better performance under information extraction models of the corresponding modalities. |
| Outcome: | The proposed method outperforms existing models on two different multimodal information extraction tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks for Continual Language Learning (CLL) are limited due to the complexity of the task and the lack of unified benchmarks. |
| Approach: | They propose a Continual Language Learning Evaluation benchmark CLLE in multilingual translation. |
| Outcome: | The proposed method is effective when compared with other strong benchmarks. |
Copied to clipboard
| Challenge: | Fill-in-the-Middle (FIM) models suffer from performance degradation and prohibitive latency. |
| Approach: | They propose a search-and-replace infilling framework that integrates agentic verification and editing into a single-pass inference process. |
| Outcome: | The proposed framework harmonizes completion tasks with the instruction-following priors of Chat LLMs, extending the paradigm from static infilling to dynamic context-aware editing. |
Copied to clipboard
| Challenge: | Language models have shown impressive abilities in a range of natural language processing tasks. |
| Approach: | This tutorial will provide an overview of the latest advances in natural language processing . it will provide preliminaries of training foundation models on code and their common practices . |
| Outcome: | This tutorial aims to provide an overview of recent advances in code modeling . it provides preliminaries of training foundation models on code and their common practices . |
Copied to clipboard
| Challenge: | Experimental results show that the MOS-aware GRM significantly improves fine-grained speech quality discrimination. |
| Approach: | They propose a MOS-aware reward model that incorporates MOS gap into reward function during reinforcement learning. |
| Outcome: | The proposed model significantly improves fine-grained speech quality discrimination. |
Copied to clipboard
| Challenge: | Existing methods for compressing Large Language Models suffer from significant truncation losses. |
| Approach: | They propose a novel method that optimizes singular value truncation in SVD compression . they use dynamic compression ratio allocation to balance the large tuncation loss . |
| Outcome: | The proposed method outperforms current state-of-the-art methods on ten datasets and five models on various scales. |
Copied to clipboard
| Challenge: | Existing systems rely on a monolithic policy to execute subgoals across varying contexts, causing inconsistent outcomes and scaling only partially mitigates. |
| Approach: | They propose a memory-routed mixtureof-experts controller for Adaptive Minecraft Control that routes via a subgoal-indexed expert memory and regulates capacity through failure-triggered expert growth and redundancy-aware consolidation. |
| Outcome: | The proposed controller shows significant gains in adaptability, robustness, and execution consistency over strong baselines. |
Copied to clipboard
| Challenge: | Current speaker diarization systems consider only acoustic information, resulting in performance degradation when encountering adverse acustic environment. |
| Approach: | They propose methods to extract speaker-related information from conversational semantics in multi-party meetings. |
| Outcome: | The proposed method improves on AISHELL-4 and AliMeeting datasets on speakers diarization and speaker-turn detection. |
Copied to clipboard
| Challenge: | Large language models exhibit intrinsic limitations such as knowledge cutoff, single-threaded reasoning that hinders finer-grained branch and aggregation, and rigid collaboration mechanisms that struggle to coordinate specialized capabilities. |
| Approach: | They propose a taxonomy spanning *Graph-Assisted Knowledge Augmentation*, *Graph Assisted Reasoning and Planning*, and *Graphed LLM Collaboration*. |
| Outcome: | The proposed models show that graphs can augment and correct LLMs and support dynamic coordination among experts and agents in collaborative settings. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) have demonstrated impressive performance on general video comprehension benchmarks, but their robustness needs to be thoroughly investigated for broader applications. |
| Approach: | They propose a temporal robustness benchmark which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. |
| Outcome: | The proposed method improves the model’s robustness and reliability in temporal analysis. |
Copied to clipboard
| Challenge: | Existing methods for learning relational patterns from data are prone to catastrophic forgetting issues due to limited number of samples and continual training mode. |
| Approach: | They propose a unified causal framework for CFRL to restore causal effects from old data . they establish two additional causal paths from old to predictions by colliding with old data separately in the old feature space. |
| Outcome: | The proposed method is superior to existing state-of-the-art methods in CFRL task settings. |
Copied to clipboard
| Challenge: | Existing methods to solve label dependency and noisy labeling problems are limited . experimental results show the proposed method is competitive to state-of-the-art methods . |
| Approach: | They propose a deep learning XML method with word-vector-based self-attention followed by ranking-based AutoEncoder architecture to solve these problems. |
| Outcome: | The proposed method is competitive to state-of-the-art methods on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing question reformulation models are based on supervised question labels without considering feedback information from answers. |
| Approach: | They propose a question reformulation model that integrates conversational history information with reinforcement learning. |
| Outcome: | The proposed model is more effective in conversational machine comprehension with reinforcement learning. |
Copied to clipboard
| Challenge: | Existing multimodal task-oriented dialog data fails to demonstrate the diverse expressions of user subjective preferences and recommendation acts in the real-life shopping scenario. |
| Approach: | They propose a multimodal task-oriented dialog dataset with subjective preferences and recommendation acts that is well-annotated with sales experts. |
| Outcome: | The proposed model is powered by a state-of-the-art multimodal model for these tasks. |
Copied to clipboard
| Challenge: | Existing methods for misinformation detection lack interpretability due to the black-box nature of the neural network. |
| Approach: | They propose a logic-based neural model which integrates interpretable logic clauses to express the reasoning process of the target task. |
| Outcome: | The proposed model can be generalizable across multiple misinformation sources and is based on three public datasets. |
Copied to clipboard
| Challenge: | Existing methods for tuning large language models from dense to MoE face significant data requirements and require large-scale post-training. |
| Approach: | They propose an upcycling instruction tuning approach for tuning a dense pre-trained model into a MoE instruction model using genetic algorithm and parameter merging. |
| Outcome: | The proposed approach improves the performance of large language models with a small amount of seed data and improves their scaling. |
Copied to clipboard
| Challenge: | Existing security evaluation benchmarks lack relevance to real-world AI programming tasks . current LLMs struggle with secure coding, research shows . |
| Approach: | They propose a repository-level evaluation benchmark to assess security of AI-generated code. |
| Outcome: | The proposed framework mirrors real-world AI programming tasks and offers valuable insights into the state of AI code generation. |
Copied to clipboard
| Challenge: | Methods for controlling large language models (LLMs) are often studied in isolation, obscuring connections and making comparison difficult. |
| Approach: | They propose a preference-utility analysis that separates control effects into preference and utility, and measures both on a shared log-odds scale using polarity-paired contrastive examples. |
| Outcome: | The proposed approach improves preference while preserving utility. |
Copied to clipboard
| Challenge: | Using a simplified version of GRU, we replace the GRUs at the middle layers of hierarchical recurrent models with Fixed-size Ordinally-Forgetting Encoding (FOFE). |
| Approach: | They propose to make the lower layers simpler than the upper ones to simplify two typical hierarchical recurrent models, namely Hierarchical Recurrent Encoder-Decoder (HRED) and R-NET, whose basic building block is GRU. |
| Outcome: | The proposed models contain less trainable parameters, consume less training time, and achieve slightly better performance than baseline models. |
Copied to clipboard
| Challenge: | Existing studies rely on shallow unsupervised data generated by token surface matching regardless of global context-aware semantics of the surrounding text tokens. |
| Approach: | They propose an Unsupervised Pseudo Semantic Data Augmentation mechanism to enrich training data without human intervention. |
| Outcome: | The proposed model improves on general zero-shot cross-lingual understanding tasks on different languages without human intervention. |
Copied to clipboard
| Challenge: | a preference evaluation metric is often biased towards longer responses, revealing a reliability problem . a decomposition of the preference evaluation into two components is needed to understand this bias. |
| Approach: | They propose to decompose the preference evaluation metric into two key components . the first component is length-dependent and related to trustworthiness . |
| Outcome: | The proposed evaluation metric is based on two components: desirability and information mass. |
Copied to clipboard
| Challenge: | Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks. |
| Approach: | They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset. |
| Outcome: | The proposed model performs well across tasks and languages. |
Copied to clipboard
| Challenge: | Large language models generate biased stances due to spurious correlations and preference towards certain individuals and topics. |
| Approach: | They propose a counterfactual Augmented Calibration Network to calibrate potential bias in stance detection of large language models. |
| Outcome: | The proposed calibration network can mitigate biases of large language models, achieving state-of-the-art results. |