Papers by Cheng Wang
Copied to clipboard
| Challenge: | Existing unlearning paradigms are mired in vague forgetting boundaries, erasing knowledge indiscriminately. |
| Approach: | They propose a benchmark to evaluate if unlearning erases essential knowledge . they propose 'knowUnDo' which uses copyrighted content and privacy domains . |
| Outcome: | The proposed method is superior to existing methods in both precise knowledge unlearning and general knowledge retaining of LLMs. |
Copied to clipboard
| Challenge: | Continual pre-training is the paradigm where pre-trained language models acquire fresh knowledge and gradually get upgraded. |
| Approach: | They propose to use adapted weights to recycle old PLMs for continual pre-training . they propose to combine initialization and distillation methods to achieve better performance . |
| Outcome: | The proposed method improves the convergence and performance of the upgraded PLM. |
Copied to clipboard
| Challenge: | Emotion-Cause Pair Extraction in Conversations (ECPEC) aims to identify the set of causal relations between emotion utterances and their triggering causes within a dialogue. |
| Approach: | They propose a framework for Emotion-Cause Pair Extraction in Conversations that decouples emotion-oriented semantics from cause-oriented ones and employs optimal transport to enable many-to-many and globally consistent emotion-cause matching. |
| Outcome: | The proposed framework achieves state-of-the-art on several benchmark datasets. |
Copied to clipboard
| Challenge: | Recent advances in retrieval-augmented generation (RAG) have substantially improved question-answering systems, particularly for factoid ‘5Ws’ questions. |
| Approach: | They propose a data organization paradigm where large language models transform documents into more structured and loosely interconnected LUs. |
| Outcome: | Experiments in open-domain and industrial settings show that the proposed paradigm outperforms existing paradigms and shows high adaptability across diverse document formats. |
Copied to clipboard
| Challenge: | Existing research to improve CoT efficiency falls into three categories, each with distinct limitations. |
| Approach: | They propose a training-free framework that addresses both dimensions of CoT reasoning by applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. |
| Outcome: | Empirical results show that the proposed framework achieves 11.3 efficiency gain without compromising accuracy. |
Copied to clipboard
| Challenge: | Existing methods for style transfer are based on an inductive learning approach, which represents the style as embeddings, decoder parameters, or discriminator parameters and directly applies these general rules to the test cases. |
| Approach: | They propose a retrieval-based context-aware style representation that involves top-K relevant sentences in the target style in the transfer process. |
| Outcome: | The proposed method outperforms several strong baselines and is general and effective to the task of unsupervised style transfer. |
Copied to clipboard
| Challenge: | Existing methods for compressing context by removing redundant tokens are inconsistent with the objective of retaining the most important tokens when conditioning on a given query. |
| Approach: | They propose a method that uses information bottleneck theory to compress context . they propose to remove redundant tokens using metrics such as self-information or perplexity . |
| Outcome: | The proposed method achieves a 25% increase in compression rate compared to the state-of-the-art . |
Copied to clipboard
| Challenge: | Cross-modal retrieval tasks are used to retrieve data from one modality or another based on a query from another modality. |
| Approach: | They propose a generative cross-modal retrieval framework based on coarse-to-fine semantic modeling . they propose combining K-Means and RQ-VAE to discretize multimodal data into token sequences that support autoregressive generation. |
| Outcome: | The proposed framework achieves excellent performance and efficiency in multimodal retrieval tasks. |
Copied to clipboard
| Challenge: | Recent research has focused on pushing weight-only quantization to extremely low-bit due to numerical representation limitations. |
| Approach: | They propose a vector-based quantization approach that pushes LLMs to extremely low-bit . they propose scalar-based weight quantization that reduces memory requirements and optimizes storage costs . |
| Outcome: | The proposed method reduces model quantization perplexity by 0.01-0.34 on LLaMA-2, 0.38-0.68 on mistral-7B, 4.41-7.34, on llaMA-3 on QA tasks on average. |
Copied to clipboard
| Challenge: | Speech Recognition often gets stuck in the lack of new domain utterances when training a model of new-domain speech. |
| Approach: | They propose a training system Open-modality Speech Recognition that enables zero-shot modality transfer . they use multi-modal alignment in phoneme space to maintain multi-modality alignment . |
| Outcome: | The proposed system achieves zero-shot modality transfer compared to existing methods . it achieves state-of-the-art performance on audio-visual speech recognition and lip-reading with 2.7% and 25.0%, respectively. |
Copied to clipboard
| Challenge: | Existing studies typically provide traditional, but less practical evaluation testbeds for multimodal systems. |
| Approach: | They propose a novel evaluation dataset, ProMQA, to measure the advancement of systems in application-oriented scenarios. |
| Outcome: | The proposed evaluation dataset reveals a significant gap between human and competitive multimodal models. |
Copied to clipboard
| Challenge: | Existing studies on LLM adoption and their social implications lack empirical grounding, weakening their validity. |
| Approach: | They propose to integrate a dataset of over 94,000 LLM use cases collected from Reddit and news articles to provide insights into LLM adoption across different domains. |
| Outcome: | The proposed dataset includes over 94,000 LLM use cases collected from Reddit and news articles. |
Copied to clipboard
| Challenge: | Existing mRAG systems suffer from a language bias during reranking, systematically favoring English and the query’s native language. |
| Approach: | They propose a language-agnostic utility-driven reranker alignment technique to mitigate language bias during re-ranking. |
| Outcome: | The proposed approach mitigates language bias and consistently improves mRAG performance across languages. |
Copied to clipboard
| Challenge: | Existing top-k attention methods struggle to strike a balance between efficiency and accuracy. |
| Approach: | They propose a top-k attention approach that integrates low-overhead techniques into the Top-k Attention process to achieve 7.2 speedup compared to vanilla full attention. |
| Outcome: | The proposed approach achieves 7.2 speedup compared to current top-k attention methods while maintaining model accuracy. |
Copied to clipboard
| Challenge: | Existing methods for sentiment classification use binary treatment of words . Existing approaches limit generalizability to novel words and low-frequency words if there is a word in a sentence that is not treated . |
| Approach: | They propose a meta-causal approach that uses a single training task to identify causal words for arbitrary words. |
| Outcome: | The proposed method reduces the spurious correlation between word treatment and sentiment classification by removing words with low treatment effects from a pre-trained language model. |
Copied to clipboard
| Challenge: | a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities . |
| Approach: | They present a comparative analysis to identify and distinguish LLM activities from human activities. |
| Outcome: | The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) has demonstrated promising potential in scaling LLMs . however, it is hindered by two critical challenges: substantial GPU memory consumption and low activated parameters. |
| Approach: | They propose an Expert-Selection Aware Compressor for Mixture-of-Experts (MoE) that aligns with the characteristics of MoE from the perspectives of quantization and pruning. |
| Outcome: | The proposed approach significantly reduces memory usage and improves inference speed with minimal performance degradation. |
Copied to clipboard
| Challenge: | Temporal Knowledge Graphs (TKGs) are used in many different areas of research. |
| Approach: | They propose to use a beam search policy to induce multiple clues from historical facts . they propose to adopt a graph convolution network based sequence method to deduce answers from clues . |
| Outcome: | The proposed model can predict future facts in two stages, Clue Searching and Temporal Reasoning. |
Copied to clipboard
| Challenge: | Existing approaches to visual segmentation from language queries require expensive labeling and degradation when deployed to an unseen domain. |
| Approach: | They propose a task to adapt a visual segmentation model from a labeled domain to an unseen domain. |
| Outcome: | The proposed framework achieves precise feature- and relation-invariant across domains via universal semantic structure. |
Copied to clipboard
| Challenge: | Existing LLMs model overly capable learners who over-apply feedback, resulting in pedagogically implausible behavior. |
| Approach: | They propose a framework that decouples cognitive ability from writing proficiency and models their interaction during writing and revision. |
| Outcome: | The proposed model produces distinguishable proficiency levels and is consistent with instructional theories. |
Copied to clipboard
| Challenge: | Pre-trained language models (LMs) have shown effectiveness in literature understanding tasks, especially when tuned via contrastive learning. |
| Approach: | They propose a multi-task contrastive learning framework that enables common knowledge sharing across different scientific literature understanding tasks while preventing task-specific skills from interfering with each other. |
| Outcome: | The proposed framework outperforms state-of-the-art pre-trained language models on a comprehensive dataset. |
Copied to clipboard
| Challenge: | Existing LLMs operate as an opaque process without explicit separation between knowledge retrieval and reasoning steps, making the decision-making process unclear and disorganized. |
| Approach: | They propose a language model inference paradigm that decomposes the complex inference process into two distinct and clear actions: (1) memory recall: which retrieves relevant knowledge, and (2) reasoning: which performs reasoning steps based on the recalled knowledge. |
| Outcome: | The proposed paradigm decomposes the inference process into two distinct and clear actions, memory and reason, guiding the model to distinguish between steps that require knowledge retrieval and those that involve reasoning. |
Copied to clipboard
| Challenge: | Existing query languages for question answering over knowledge bases are not capable of processing queries presented in human language directly. |
| Approach: | They advocate a new model architecture that includes a verification mechanism for checking the correctness of predicted relations. |
| Outcome: | The proposed approach dramatically improves the question answering performance. |
Copied to clipboard
| Challenge: | Existing methods for analyzing textual attributes in product catalogs are not effective on structured tabular data since they are trained on free-form natural language texts. |
| Approach: | They propose a model to handle error detection over tabular data following a pre-training paradigm. |
| Outcome: | The proposed model improves on a real-world Amazon Product Catalog table by 16% over state-of-the-art methods and by 11% on PR AUC over attribute value validation task. |
Copied to clipboard
| Challenge: | Large-scale pre-trained language models require enormous computational resources and long training time. |
| Approach: | They propose an algorithm to reduce inference time and train large NLP models by slimming the self-attention and fully-connected sub-layers inside a transformer. |
| Outcome: | The proposed algorithm achieves comparable performance to standard BERT with 35 45% less training time. |
Copied to clipboard
| Challenge: | Real-world machine learning systems are achieving excellent performance in terms of coarse-grained metrics like overall accuracy and F-1 score. |
| Approach: | They extend slice-based learning (SBL) with a mixture of attentions to learn slice-aware dual attentive representations. |
| Outcome: | The proposed approach outperforms the baseline method and the original SBL approach on monitored slices with two natural language understanding tasks. |
Copied to clipboard
| Challenge: | Chinese word segmentation datasets have ambiguous annotation criteria resulting in multi-grained compounds. |
| Approach: | They propose a domain adaptive segmenter to exploit diverse annotation criteria of datasets . they use bidirectional encoder representations from transformers to introduce open-domain knowledge . |
| Outcome: | The proposed model outperforms the state-of-the-art models on 10 Chinese word datasets with superior efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair. |
| Approach: | They propose a repository-level code debugging dataset with 22 subtypes of errors that supports 8 commonly used programming languages and 3 debug tasks. |
| Outcome: | The proposed dataset supports 8 commonly used programming languages and 3 debugging tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) produce hallucinated text, compromising their practical utility in professional contexts. |
| Approach: | They have developed an unconstrained hallucination generation evaluation benchmark that contains hallucines generated by large language models with minimal restrictions. |
| Outcome: | The proposed benchmarks are based on a Chinese-language dataset that is lacking in the field. |
Copied to clipboard
| Challenge: | Current methods for multimodal representation learning for electrocardiograms often result in suboptimal alignment of ECG signals with their corresponding text reports. |
| Approach: | They propose a framework to learn ECG representations by aligning ECG signals with paired free-text reports. |
| Outcome: | The proposed framework outperforms existing methods in zero-shot classification and linear probing tasks using 12 leads. |
Copied to clipboard
| Challenge: | Named Entity Recognition models are feature-engineering and machine learning based. |
| Approach: | They propose a new NER learning framework that uses entity mentions to improve model performance. |
| Outcome: | The proposed model achieves better performance on OOV entities on various settings and datasets. |
Copied to clipboard
| Challenge: | In-context knowledge editing (ICE) is currently the most effective method for knowledge editing, but it is constrained by the black-box modeling of LLMs and lacks interpretability. |
| Approach: | They propose a method to decode new knowledge by comparing logits with unedited knowledge to improve the accuracy of LLMs. |
| Outcome: | The proposed method improves the performance of LLaMA3-8B-instruct on MQuAKE by up to 219%. |
Copied to clipboard
| Challenge: | Existing studies focus on specialized agents designed for particular tasks. |
| Approach: | They propose to scale annotated interaction trajectories and fine-tune LLMs on AgentBank to get a series of agent models, Samoyed. |
| Outcome: | The proposed model can scale to get generalized agent capabilities. |
Copied to clipboard
| Challenge: | Natural Language Inference (NLI) is a foundational understanding task in language understanding. |
| Approach: | They propose a framework to construct counterfactual reasoning data and fine-tune LLMs to reduce attestation bias. |
| Outcome: | The proposed framework reduces hallucinations from attestation bias on original and bias-neutralized datasets while keeping hypotheses unchanged. |
Copied to clipboard
| Challenge: | Recent studies have focused on prompt engineering to extract sentence embeddings from large language models (LLMs) but these models are mostly decoder-only and the earlier tokens in the sentence cannot attend to the latter, resulting in biased encoding of sentence information and cascading effects on the final decoded token. |
| Approach: | They propose a plug-and-play and training-free technique that prepends each layer’s decoded sentence embedding to the beginning of the sentence in the next layer’ s input. |
| Outcome: | The proposed technique can significantly improve the performance of existing prompt-based sentence embedding methods across different LLMs while incurring negligible additional inference cost. |
Copied to clipboard
| Challenge: | Recent studies highlight a fundamental dichotomy in deep learning optimization: parameter updates along the top eigendirections of the loss Hessian (Dom-space) capture most of the update magnitude, while updates in the orthogonal component (Bulk-space), have smaller magnitudes but drive most learning progress. |
| Approach: | They propose a plug-and-play framework that scales update components projected onto distinct subspaces and a block-wise strategy that applies this estimation on a per-parameter-block basis. |
| Outcome: | The proposed framework accelerates training by differentially scaling update components projected onto distinct subspaces, while enhancing stability by moderating updates in dominant subspace and boosting convergence speed by amplifying updates in bulk-space. |
Copied to clipboard
| Challenge: | Existing methods for content-based recommendation with missing or corrupted modalities are lacking in learning multimodal models. |
| Approach: | They propose a multimodal multimodal autoencoder that learns multimodal representations for complementing and imputing missing modalities. |
| Outcome: | The proposed framework achieves state-of-the-art performance on rating prediction tasks and is more robust to previous methods in alleviating data-sparsity and the cold-start problem. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown impressive capabilities in various downstream tasks but typically face Catastrophic Forgetting (CF) during fine-tuning. |
| Approach: | They propose a pruning-based approach to balance CF and downstream task performance by integrating the ratio of the task vector to pre-trained model parameters into the pruning criteria. |
| Outcome: | The proposed pruning-based approach limits CF to just 0.25% while maintaining 99.67% accuracy on downstream tasks. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) lack visual input, leading to errors in basic numerical comparisons. |
| Approach: | They propose a spatial OODA framework that integrates the OODAC cognitive loop into multiple control tasks and integrates it into LLMs. |
| Outcome: | The proposed model significantly improves the spatial reasoning capabilities of large language models across multiple scenarios including SPOD-Bench, SPACE and applications. |
Copied to clipboard
| Challenge: | Relation extraction (RE) is an essential topic in natural language processing and has attracted extensive attention. |
| Approach: | They propose a case-oriented construction framework to build a hard case relation extraction dataset with 65,225 relational facts annotated from 9,231 documents. |
| Outcome: | The proposed model achieves a high 96% F1 score on data quality and is far lower than humans. |
Copied to clipboard
| Challenge: | SDiaReward is an end-to-end spoken dialogue system that integrates paralinguistic nuances and spontaneous nature of human conversation. |
| Approach: | They propose an end-to-end multi-turn reward model trained on SDiaReward-Dataset . it is a collection of episode-level preference pairs targeting modality and colloquiality gaps . |
| Outcome: | The proposed model outperforms general-purpose audio LLMs in episode-level evaluation. |
Copied to clipboard
| Challenge: | Long contexts break transformers, attention scores dilute, model cannot adapt to novel patterns at inference time. |
| Approach: | They propose a framework that gates the memory consolidation process by estimating Contextual Utility . they propose GDWM to maintain a form of working memory with constant contexts . |
| Outcome: | The proposed framework achieves comparable or superior performance on sparse-information tasks with 4 fewer gradient steps compared to uniform baselines. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been used to evaluate the safety of their users . however, evaluation questions in current benchmarks are too straightforward and difficult to update with practical relevance due to their lack of correlation with real-world events. |
| Approach: | They propose a question-generation framework to evaluate the safety of LLMs in the Chinese context. |
| Outcome: | The proposed framework reduces decline rate while maintaining similar attack success rate. |
Copied to clipboard
| Challenge: | Existing models for encoding long sequences in deep learning suffer from high latency and memory demands. |
| Approach: | They propose a clustering-based sparse Transformer framework to perform attention across chunked sequences. |
| Outcome: | The proposed framework achieves state-of-the-art on several major QA benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for retrieval-based in-context learning ignore model biases and fail to retrieve the most appropriate demonstrations for different LLMs. |
| Approach: | They propose a model-specific demonstration retrieval method that considers the biases of different LLMs at inference time. |
| Outcome: | The proposed method improves performance on seen and unseen tasks with multi-scale inference LLMs by up to 41.2%. |
Copied to clipboard
| Challenge: | Analogy-making between narratives is crucial for human reasoning . despite its importance, there has been limited research on story analogies . |
| Approach: | They construct a large-scale story-level analogy corpus with 24K story pairs . they find that the tasks are incredibly difficult for large language models such as ChatGPT . |
| Outcome: | The proposed corpus contains 24K story pairs from diverse domains with human annotations on two similarities from the extended Structure-Mapping Theory. |
Copied to clipboard
| Challenge: | Existing transfer learning techniques focus on uni-modal analysis and lack consideration of multi-modal content and cross-modal relation. |
| Approach: | They propose a transferable audio-visual text generation framework that incorporates two components: Audio-Visual Meta-Mapper and Dual Counterfactual Contrastive Learning. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods across multiple domains and modal settings. |
Copied to clipboard
| Challenge: | Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100. |
| Approach: | They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth. |
| Outcome: | The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, but it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. |
| Approach: | They propose a benchmark that combines a constructionist Out-of-Sample dataset with reverse understanding probes to evaluate large-scale large-format models. |
| Outcome: | The proposed model performs well on classical Chinese poetry benchmarks, but a performance gap persists . the model can complete famous couplets and can be used to understand a variety of texts. |
Copied to clipboard
| Challenge: | Existing approaches to personalize large language models (LLMs) rely on heuristic methods to compress user profiles but they ignore how LLMs process and prioritize different profile components. |
| Approach: | They propose an attention-guided context compression framework that leverages attention feedback from a marking model to mark important personalization sentences and guides a compression model to generate task-relevant compressed user contexts. |
| Outcome: | The proposed framework outperforms baselines across tasks, token limits, and settings while reducing token usage by 50 times. |
Copied to clipboard
| Challenge: | Experimental results show that Flow-matching generative models can scale training by increasing data, computational resources, and model size. |
| Approach: | They propose a flow-matching transformer with masked generative modeling for scaling text-to-audio inference-time prediction. |
| Outcome: | The proposed model scales inference-time computations by masking generation and re-predicting them through iterative decoding. |
Copied to clipboard
| Challenge: | Retrieval-augmented Large Language Models struggle with complex inputs and noisy knowledge retrieval hindering model effectiveness. |
| Approach: | They propose a query generation method that integrates query generation blending with knowledge filtering to enhance retrieval-augmented LLMs. |
| Outcome: | The proposed approach surpasses state-of-the-art benchmarks on open-domain question answering benchmarks. |
Copied to clipboard
| Challenge: | Pretrained language models have been successfully applied to a wide range of tasks . however, the pretraining tasks were based on the context of documents . |
| Approach: | They propose a self-supervised joint training framework with a method called Masked Query Prediction to establish semantic relations between given queries and positive documents. |
| Outcome: | The proposed framework outperforms existing models on document reranking tasks without further pre-training . it uses a self-supervised method to establish semantic relations between given queries and positive documents. |
Copied to clipboard
| Challenge: | Current methods for multi-modal entity alignment ignore relative interactions between modalities and the accuracy of weights. |
| Approach: | They propose a relative interaction and calibration framework for multi-modal entity alignment that uses attention mechanisms to perceive the uncertainty of the weight for each modality. |
| Outcome: | The proposed framework outperforms baselines across 5 datasets and 23 settings. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in various domains but face challenges when applied to data science workflows due to their complex, multi-stage nature. |
| Approach: | They propose a hierarchical graph-based agent that represents complexity and a progressive strategy for step-by-step verification, refinement, and consistent context management. |
| Outcome: | The proposed agent surpasses state-of-the-art baselines on the MATH dataset and performs better on InfiAgent-DABench. |
Copied to clipboard
| Challenge: | Currently, research on complex chart understanding tasks is limited . a pipeline for visual reasoning datasets addresses these limitations . |
| Approach: | They propose a code-driven pipeline for generating visual reasoning datasets . pipeline integrates retrieval-augmented generation to retrieve professional chart templates . |
| Outcome: | The proposed pipeline enhances chart diversity and data quality through model-based evaluation. |
Copied to clipboard
| Challenge: | Multimodal machine translation (MMT) aims to improve translation quality by incorporating information from other modalities, such as vision. |
| Approach: | They propose a framework for multimodal machine translation that utilizes large-scale non-triple data and a multimodal translation dataset. |
| Outcome: | The proposed method can significantly improve translation performance with more non-triple data. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show promising results in language generation but often “hallucinate”, making their outputs less reliable. |
| Approach: | They propose to shift attention to more relevant components at token- and sentence-levels for better UQ. |
| Outcome: | The proposed approach improves the performance of a range of popular “off-the-shelf” LLMs with model sizes extending up to 33B parameters. |
Copied to clipboard
| Challenge: | Current approaches to temporal knowledge representation face limited generalization to unseen facts and insufficient interpretability of reasoning processes. |
| Approach: | They propose a framework that uses a denoising diffusion process to complete reasoning tasks . they propose introducing a noise source and historical conditionguiding mechanism to improve interpretability . |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three benchmark datasets. |
Copied to clipboard
| Challenge: | MT-DNN is an open-source natural language understanding toolkit . it allows researchers and developers to train customized deep learning models . |
| Approach: | They present MT-DNN, an open-source natural language understanding toolkit . it is designed to facilitate rapid customization for a broad spectrum of NLU tasks . MT supports multi-task knowledge distillation, which can substantially compress a deep neural model without significant performance drop. |
| Outcome: | The proposed model can significantly compress a large model without significant performance drop. |
Copied to clipboard
| Challenge: | Existing LLMs often rely on complex prompting or extensive fine-tuning to introduce new capabilities while preserving strong generalizability. |
| Approach: | They propose a large-scale pre-training corpus to enhance LLM agents' capabilities . they use 103B agent-specific data encompassing 76,537 APIs . |
| Outcome: | The proposed training corpus outperforms open-source LLMs and commercial LLM agents on three agent benchmarks. |
Copied to clipboard
| Challenge: | Recent work on spoken video grounding challenges extracting semantic information from speech . previous studies focused on textual queries, but recent work focuses on spoken queries . |
| Approach: | They propose a framework for weakly-supervised spoken video grounding to represent cross-modal semantics without expensive temporal annotations. |
| Outcome: | The proposed framework is more efficient than existing methods. |
Copied to clipboard
| Challenge: | Multi-modal entity alignment (MMEA) is a critical task that aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs). |
| Approach: | They propose a novel MMEA transformer that hierarchically introduces neighbor features, multi-modal attributes, and entity types to enhance alignment task. |
| Outcome: | The proposed transformer hierarchically introduces neighbor features, multi-modal attributes, and entity types to enhance the alignment task. |
Copied to clipboard
| Challenge: | Existing studies on graph learning with large language models have focused on the link prediction task on large graphs. |
| Approach: | They propose a framework for scalable link prediction on large-scale heterogeneous graphs based on large language models. |
| Outcome: | The proposed framework outperforms baselines in link prediction tasks on large graphs. |
Copied to clipboard
| Challenge: | Existing methods for question generation over knowledge bases have low diversity and poor fluency due to the limited information contained in the subgraphs and semantic drift due to decoder’s oblivion of the semantics of the answer entity. |
| Approach: | They propose a knowledge-enriched, type-constrained and grammar-guided KBQG model that generates natural-language questions over a set of triples in the KB. |
| Outcome: | The proposed model outperforms existing methods on two widely-used benchmark datasets. |
Copied to clipboard
| Challenge: | Generative audio modeling has been fragmented into specialized tasks such as text-to-speech (TTS), text- to-music (TTM), and text-ta (TTA) specialized models require reference audio for timbre cloning and strict phoneme alignment, whereas TTA models generate unstructured textures from open-ended captions. |
| Approach: | They propose a unified flow-matching framework capable of synthesizing speech, music, sound effects . they propose 'token injection mechanism' that projects unstructured environmental sounds into structured temporal latent space . |
| Outcome: | The proposed framework achieves state-of-the-art performance in instruction-based TTS and TTM while maintaining competitive fidelity in TTA. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) have shown strong potential in various downstream tasks. |
| Approach: | They propose to model adversarial attack task as a sequential decision-making problem where the whole attack process is sequential with two decision- making problems, i.e., word finder and word substitution. |
| Outcome: | The proposed approach achieves the highest attack success rate with a comparable modification rate and semantic similarity to attack fine-tuned BERT. |
Copied to clipboard
| Challenge: | Recent large language models like GPT-4 have demonstrated astonishing zero-shot capabilities in general domain tasks, but they often generate content with hallucinations in specific domains such as Chinese law. |
| Approach: | They propose a framework for adapting large language models (LLMs) to Chinese legal domains by reformulating generation as an adapt-retrieve-revise process. |
| Outcome: | The proposed framework outperforms existing models in the Chinese legal domain by +33.6 points in the zero-shot setting. |
Copied to clipboard
| Challenge: | Agentic learning increasingly hinges on interaction, yet real-world experience is expensive, limited, and often irreversible at inference time. |
| Approach: | They propose a framework that reframes language modeling as next-state prediction under interaction. |
| Outcome: | The proposed framework evaluates world models in text-based environments . it shows that sufficiently trained models capture coherent environment dynamics . |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations in LLMs rely on external knowledge for reference retrieval or require sampling multiple responses for consistency verification. |
| Approach: | They propose a reference-free, uncertainty-based method for detecting hallucinations in Large Language Models that imitates human focus in factuality checking from three aspects: focus on the most informative keywords; focus on unreliable tokens in historical context; focus of token properties such as token type and token frequency. |
| Outcome: | The proposed method achieves state-of-the-art performance across all evaluation metrics and eliminates the need for additional information. |
Copied to clipboard
| Challenge: | Existing approaches focus on functional tool selection following user instructions while overlooking the critical role of context-aware personalization in tool selection. |
| Approach: | They propose a benchmark to evaluate LLMs’ capabilities in personalized tool utilization. |
| Outcome: | The proposed benchmark evaluates LLMs' capabilities in personalized tool utilization. |
Copied to clipboard
| Challenge: | Existing approaches to machine translation support autoregressive, semi-autoregressive and refinement-based non-auto-regressives. |
| Approach: | They propose a unified approach for supporting different generation manners of machine translation including autoregressive, semi-autoregressive and refinement-based non-auto-regressives. |
| Outcome: | The proposed approach achieves better or competitive translation performance compared with strong baseline models in all the settings. |
Copied to clipboard
| Challenge: | Existing data selection methods for RLVR are heuristic-based, lacking theoretical guarantees and generalizability. |
| Approach: | They propose an off-policy influence estimation method that approximates data influence using offline trajectories. |
| Outcome: | The proposed method reduces the computational cost of policy rollouts and improves storage and computation efficiency. |
Copied to clipboard
| Challenge: | Existing representation methods fail to fully capture olfactory pathway . current approaches focus on isolated segments of the olefactory pathways . |
| Approach: | They propose a representation learning framework that aligns molecular structure, receptor sequence, and natural language description. |
| Outcome: | The proposed framework achieves state-of-the-art and excellent zero-shot generalization . it decouples contributions of molecular structure, receptor sequence, and natural language description . |
Copied to clipboard
| Challenge: | Entity Alignment (EA) is a critical task in Knowledge Graph (KG) integration. |
| Approach: | They propose a novel approach that leverages the data characteristics of synthetic benchmarks to improve performance in real-world datasets. |
| Outcome: | The proposed approach outperforms state-of-the-art models on real-world datasets and achieves a 29.94% improvement in Hits@1 on DOREMUS and 5.64% improvement on AGROLD. |
Copied to clipboard
| Challenge: | Existing datasets focus on captions describing images or videos, which are not large and diverse enough. |
| Approach: | They propose a large-scale video subtitle translation dataset to facilitate multi-modality machine translation. |
| Outcome: | The proposed dataset is 10 times larger than the widely used *How2* and *VaTeX* datasets. |
Copied to clipboard
| Challenge: | Multi-agent systems (MAS) powered by Large Language Models (LLMs) have been demonstrated to push the boundaries of LLM capabilities, yet they often face significant costs and challenges in dynamic LLM selection. |
| Approach: | They propose a multi-agent system routing solution that integrates all components of MAS into a unified routing framework. |
| Outcome: | The proposed solution is high-performing, cost-effective, and efficient . it reduces overhead by up to 52.07 compared to current methods on HumanEval . |
Copied to clipboard
| Challenge: | Existing reward models evaluate empathy from a single perspective, overlooking bidirectional interaction nature of empathy. |
| Approach: | They propose a reward model that evaluates empathy from a single perspective . they propose PERM to integrate a bystander perspective to monitor overall interaction quality . |
| Outcome: | a new reward model outperforms state-of-the-art models on an emotional intelligence benchmark and an industrial daily conversation dataset. |
Copied to clipboard
| Challenge: | Existing approaches to recognize relational relationships with a few support samples are limited for unlimited queries. |
| Approach: | They propose a simple but effective framework that uses relation descriptions as external knowledge to enhance the model’s comprehension of the relation semantics. |
| Outcome: | The proposed framework outperforms strong baselines while being robust against various NOTA rates. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to adequately evaluate the proficiency of Large Language Models (LLMs) Existing standards do not cover the skills needed to evaluate LLMs in scientific literature analysis. |
| Approach: | They propose a benchmark to evaluate the proficiency of large language models in scientific literature analysis. |
| Outcome: | SciAssess evaluates 11 LLMs on multiple tasks across scientific fields. |
Copied to clipboard
| Challenge: | Existing methods, such as a n-terminal coding, do not provide accurate data for large language models. |
| Approach: | They propose a lightweight framework that leverages attention distributions and uncertainty signals in a single-pass decoding. |
| Outcome: | Experiments on open-book QA datasets show that DAGCD improves faithfulness and robustness while preserving computational efficiency. |
Copied to clipboard
| Challenge: | Existing approaches generate workflows either at task level or query level, but their relative costs and benefits remain unclear. |
| Approach: | They propose a query-level workflow generation framework that generates tasks at task level and query level. |
| Outcome: | The proposed framework reduces token usage by up to 83% compared to existing approaches . it maintains competitive performance with an average degradation of just 0.61% compared with existing approaches across multiple datasets . |
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) aims to evaluate the quality score of input essays without human intervention. |
| Approach: | They propose an unsupervised approach to evaluate the quality of input essays . they use multiple heuristic quality signals as pseudo-groundtruths to train a neural AES model . |
| Outcome: | The proposed approach achieves state-of-the-art performance on eight prompts of ASPA dataset compared with previous unsupervised methods . |
Copied to clipboard
| Challenge: | Autoregressive LLMs perform well on relational tasks that require linking entities via relational words, but it is unclear whether they learn the logical semantics of such relations or whether left-to-right order bias is involved. |
| Approach: | They propose a framework that generates text from symmetric/inverse triples and trains autoregressive models from scratch. |
| Outcome: | The proposed framework generates text from symmetric/inverse triples, trains autoregressive models from scratch, and evaluates memorization, logical inference, and in-context generalization to unseen entities. |
Copied to clipboard
| Challenge: | Automatic speech recognition systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0. |
| Approach: | They propose to use Mandarin speech datasets to analyze pronunciation and tone of children aged 3 to 5 and evaluate their models on speaker verification (SV) They find that the datasets are more robust than those used by adult speech recognition systems and are open-source and available for all academic purposes. |
| Outcome: | The proposed dataset includes 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. |
Copied to clipboard
| Challenge: | Existing MKGC methods train with all modalities available, implicitly assuming consistent complementarity . however, this often induces modality dependence and modality competition under heterogeneous noise, which can hinder robust multi-modal fusion and limit overall performance. |
| Approach: | They propose a framework to infer missing links in multimodal knowledge graphs by leveraging structured triples together with auxiliary modalities such as text and images. |
| Outcome: | The proposed framework outperforms baselines and achieves new state-of-the-art results. |
Copied to clipboard
| Challenge: | Continual learning approaches fail to achieve autonomy lifelong improvement in dynamic environments . current task-oriented dialog systems are static, unable to learn from ongoing interactions . |
| Approach: | They propose a lifelong self-evolving dialog framework that integrates evolutionary computation and LLM driven self-improvement into a single framework. |
| Outcome: | The proposed framework surpasses state-of-the-art methods and exhibits continuous performance gains throughout evolution. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit significant but subtle weaknesses, such as mistakes in instruction-following or coding tasks. |
| Approach: | They propose a framework to automatically expose weaknesses in Large Language Models (LLMs) they use three LLM-powered agents to perform comprehensive weakness identification . |
| Outcome: | The proposed framework shows that it is more effective than untargeted data augmentation methods like Self-Instruct to identify weaknesses in LLMs. |
Copied to clipboard
| Challenge: | Existing methods for quantizing large embeddings rely on Euclidean quantization, which is poorly aligned with the angular geometry induced by contrastive embeddment training. |
| Approach: | They propose a geometry-aware distillation method that compresses large embeddings into short discrete representations via iterative Householder transformations on the unit hypersphere. |
| Outcome: | The proposed method reduces decoding cost and maintains strong semantic retrieval accuracy. |
Copied to clipboard
| Challenge: | Existing approaches to empathetic response generation only consider causalities between the user’s emotion and the user's experiences and neglect interdependence among causalities and reason them independently. |
| Approach: | They propose to use a conditional variable Graph Auto-Encoder to reason all plausible causalities interdependently and simultaneously given the user’s emotion, dialogue history, and future dialogue content. |
| Outcome: | The proposed method achieves state-of-the-art in a real-world situation. |
Copied to clipboard
| Challenge: | Existing rule-based chunking methods lead to suboptimal splits, where overly large chunks introduce irrelevant information and small chunks lack semantic coherence. |
| Approach: | They propose a method that leverages document summaries as pseudo-instructions to guide chunking by computing semantic similarity between sentences and the summary. |
| Outcome: | Experiments on multiple open-domain question-answering benchmarks show that PIC significantly improves retrieval accuracy (Hits@k) and end-to-end QA performance (Exact Match) without any additional training. |
Copied to clipboard
| Challenge: | Existing methods to compress language models use a simple L_2 loss to distill knowledge in the intermediate representations of a large BERT model to a smaller one. |
| Approach: | They propose a method that uses knowledge distillation to distill knowledge through intermediate layers of the teacher via a contrastive objective. |
| Outcome: | The proposed method outperforms state-of-the-art methods on the GLUE benchmark. |
Copied to clipboard
| Challenge: | Existing studies on knowledge inference on binary facts have focused on finding out connotative valid facts. |
| Approach: | They propose a neural network model, NeuInfer, for knowledge inference on n-ary facts. |
| Outcome: | The proposed model can cope with the task to infer an unknown element in a whole fact, while ignoring the binary facts. |
Copied to clipboard
| Challenge: | Existing studies on multilingual machine translation have ignored the importance of LTs. |
| Approach: | They propose to use language tag (LT) strategies to indicate translation directions in MNMT to enhance consistency and alleviate off-target issues in zero-shot directions. |
| Outcome: | The proposed model could translate between unsupervised languages and achieve a +8 BLEU score difference over other LT strategies in translation tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have significantly impacted various domains, especially through organized LLM-driven autonomous agents. |
| Approach: | They propose a framework that enables orchestrated teams to jointly propose various task-oriented solutions and interact with their insights in a self-independence while cross-team collaboration environment for superior solutions generation. |
| Outcome: | Experiments show that the framework can generate better software quality compared to state-of-the-art frameworks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated better safety performance in high-resource languages than in low-resourced languages. |
| Approach: | They propose language-agnostic semantic alignment (LASA) which anchors safety alignment directly in semantic bottlenecks. |
| Outcome: | The proposed approach significantly improves safety across all languages: average attack success rate drops from 24.7% to 2.8% on LLaMA-3.1-8B-Instruct and remains within 3–4% across Qwen2.5 and Qwend3 Instruct models (7B–32B). |
Copied to clipboard
| Challenge: | Existing methods for extracting conditional text embeddings from large language models (LLMs) relying on prompts often fails to produce high-quality conditional embeddables, resulting in degradation of quality. |
| Approach: | They propose a plug-and-play method that constructs unconditional general text embeddings and uses them to refine conditional text embeds. |
| Outcome: | The proposed method improves performance of prompt-based methods on clustering, Semantic Textual Similarity, and triplet alignment datasets. |
Copied to clipboard
| Challenge: | Existing methods to extract information from evidence are unable to grasp relational and logical information among the evidence. |
| Approach: | They propose a graph-based evidence aggregating and reasoning framework to integrate evidence from multiple pieces of evidence. |
| Outcome: | The proposed framework achieves significant performance improvements on a large-scale benchmark dataset. |
Copied to clipboard
| Challenge: | Medical dialogue systems (MDS) aim to provide patients with medical services, such as diagnosis and prescription. |
| Approach: | They propose a Dual Flow enhanced Medical (DFMed) dialogue generation framework that extracts the medical entities and doctor's dialogue acts used in the dialogue history and models their transitions with an entity-centric graph flow and a sequential act flow. |
| Outcome: | The proposed framework exceeds baselines in both automatic and manual evaluations on two datasets. |
Copied to clipboard
| Challenge: | Existing RL-based approaches to function calling are misaligned between reasoning processes and tool-call decisions. |
| Approach: | They propose a reasoning-aware RL framework for interpretable function calling . they integrate a composite reward integrating format/correctness constraints, CER, and SMV . |
| Outcome: | Experiments on BFCL/ACEBench show R2IF outperforms baselines by 34.62% with positive Average CoT Effectiveness. |
Copied to clipboard
| Challenge: | Effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. |
| Approach: | They propose a multi-dimensional benchmark for evaluating LLMs’ alignment in Chinese with 8 main categories, 683 real-scenario rooted queries and corresponding human verified references. |
| Outcome: | The benchmark uses a human-in-the-loop data curation pipeline, 683 real-scenario rooted queries and human verified references. |
Copied to clipboard
| Challenge: | Existing methods for large language models focus on post-training stage or inference time, while few tackle the source of toxicity, namely, the dataset itself. |
| Approach: | They propose to localize and rewrite toxic spans in raw corpora with SoCD, which guides an LLM to localized and preserving semantics while preserving toxicity. |
| Outcome: | The proposed method reduces TP from 0.42 to 0.18 and Expected Maximum Toxicity (EMT) from 0.43 to 0.20 on three LLMs. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on online one-on-one chatting or human-AI interactions, neglecting real-world scenarios. |
| Approach: | They propose a framework to curate a lifelog benchmark that combines two subsets of audio data to address temporal leakage in offline settings. |
| Outcome: | The proposed framework outperforms existing benchmarks on live chats and AI interactions. |
Copied to clipboard
| Challenge: | Speculative decoding (SD) is a training-free SD framework that orchestrates dynamic alternation combining serial dynamic drafting with parallel draft verification. |
| Approach: | They propose a serial and parallel intertwined speculative DEcoding framework that orchestrates dynamic alternation combining serial dynamic drafting and parallel draft verification. |
| Outcome: | The proposed framework accelerates inference while reducing the LLM usage costs. |
Copied to clipboard
| Challenge: | Existing benchmarks for knowledge editing do not accurately evaluate how well models apply knowledge in real-life situations. |
| Approach: | They propose a benchmark to evaluate how well updated models apply new knowledge in real-life situations. |
| Outcome: | The proposed method achieves 39.47 accuracy on GPT-4o-mini but drops significantly to 3.83 on Qwen2.5-3B. |
Copied to clipboard
| Challenge: | Continual fine-tuning of large language models suffers from catastrophic forgetting . some approaches use routers to assign tasks to experts, but continual learning often requires retraining . |
| Approach: | They propose a framework that integrates routing and response mechanisms within each expert . it eliminates the need for an additional router and allows each expert to decide whether a query should be handled . |
| Outcome: | The proposed framework outperforms previous approaches in continual fine-tuning . it can handle learning tasks and out-of-distribution instances, paving the way for distributed model ensembling. |
Copied to clipboard
| Challenge: | Existing methods to detect toxic generation of pretrained language models rely on templates, data extraction, crowdsourcing workers or automatic generation. |
| Approach: | They propose a method to construct adversarial contexts conditioned on a given response . they augment existing dataset BAD+ and construct a new dataset B AD+ . |
| Outcome: | The proposed method can detect toxic or biased content in large pretrained language models. |
Copied to clipboard
| Challenge: | Existing approaches to multi-hop question answering struggle to identify and organize dynamic knowledge . et al., 2023; Liu e.t. al. 2023) suggest a dual-process framework for multi-step reasoning . |
| Approach: | They propose a synergistic dual-process framework that integrates reasoning and retrieval. |
| Outcome: | The proposed framework improves answer accuracy and coherence even in smaller-scale models. |
Copied to clipboard
| Challenge: | Current methods embed factual knowledge into continuous vector space and apply geometric operations to learn potential patterns in temporal knowledge graphs. |
| Approach: | They propose a temporal knowledge graph completion method that uses two geometric operations to learn missing facts in temporal graphs. |
| Outcome: | The proposed method significantly outperforms existing temporal knowledge graph embedding models. |
Copied to clipboard
| Challenge: | Existing approaches to training pre-trained language models (PLMs) focus on static image modality; inevitably encounter modality gaps and noise; and treat all modalities. |
| Approach: | They propose a multimodal-augmented framework that can infuse multimodal semantics into PLMs and facilitate a self-balancing calibration of information allocation. |
| Outcome: | The proposed framework outperforms baselines on multiple NLP tasks and outperformed existing frameworks. |
Copied to clipboard
| Challenge: | Existing approaches to train multilingual models to learn the inductive bias of a shared vocabulary and set of parameters across languages. |
| Approach: | They propose to use a multilingual crossover encoder-decoder to fuse language pairs at an instance level to encourage sharing of input and output spaces. |
| Outcome: | The proposed approach improves quality on English-to-Many, Many-to English and zero-shot translation tasks from +0.5 BLEU up to +5.5 BLUE points. |
Copied to clipboard
| Challenge: | Long-context extension attempts to extend contextual window in pre-trained LLMs . primary method involves expanding initial positional encodings, disrupting positional learning . |
| Approach: | They propose a new extension strategy based on Rotary Position Embedding to extend contextual window in pre-trained large language models. |
| Outcome: | The proposed method can extend the contextual window in pre-trained large language models . expansion disrupts positional encodings learned during pre-training, authors show . |
Copied to clipboard
| Challenge: | Existing work on stance detection tasks using large language models shows promising results, but it may not be able to provide detailed background knowledge. |
| Approach: | They propose a method which leverages the generated experienced experts and lets LLMs reason in a semi-parametric way. |
| Outcome: | The proposed method outperforms methods with self-consistency reasoning and reduces bias. |
Copied to clipboard
| Challenge: | Existing methods for visual token pruning rely on predefined configurations without determining whether they achieve optimal performance. |
| Approach: | They propose a framework that formulates visual token pruning as a Pareto configuration optimization problem to automatically identify optimal configurations. |
| Outcome: | The proposed framework approximates the empirical Pareto frontier obtained through grid search and generalizes well across pruning methods and VLM architectures. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can now access a wide range of external tools thanks to the Model Context Protocol (MCP). |
| Approach: | They expose a vulnerability in prevalent tool/function-calling protocols by editing tool descriptions to find out which tools are used by LLMs. |
| Outcome: | The proposed changes in the tool descriptions can increase the usage of tools from LLMs when competing with alternatives. |
Copied to clipboard
| Challenge: | Existing studies focus on replicating macro-level stylized facts while neglecting verification of micro-level decision-making. |
| Approach: | They propose a framework that replicates macro-level stylized facts while ignoring micro-level decision-making. |
| Outcome: | The proposed framework improves alignment with human trends and captures behavioral heterogeneity. |
Copied to clipboard
| Challenge: | Existing graph-based detection models are vulnerable to deceptive message propagation, where bots deliberately interact with legitimate users. |
| Approach: | They propose a framework to mitigate deceptive message propagation by node-level uncertainty estimation and graph structure purification. |
| Outcome: | The proposed framework improves on three benchmark datasets and six GNN backbones on real-world social bots. |
Copied to clipboard
| Challenge: | Emergent Large Language Models (LLMs) use extraordinary performance and powerful deduction capacity to discern from traditional language models. |
| Approach: | They propose a method that uses weights to compensate quantization error and learnable singular value incremental (LSI) LSI is a technique that helps weights compensate each other conditioned on activation. |
| Outcome: | The proposed method achieves state-of-the-art performance in diverse quantization settings, no matter in weight-only, weight-activation or extremely low bit scenarios. |
Copied to clipboard
| Challenge: | Existing methods to improve LLMs' reasoning abilities suffer from diminishing returns or high computing overhead. |
| Approach: | They propose a genetic evolutionary framework that casts CoT generation as a population-based search over reasoning trajectories. |
| Outcome: | The proposed framework improves correct-CoT synthesis success by over 30% and enhances structural diversity with markedly improved efficiency. |
Copied to clipboard
| Challenge: | Existing studies on conversational recommender systems lack a unified and standardized implementation or comparison. |
| Approach: | They propose to use a unified framework and highly-decoupled modules to develop CRSs. |
| Outcome: | The proposed framework collects 6 commonly used human-annotated CRS datasets and implements 19 models that include advanced techniques such as graph neural networks and pre-training models. |
Copied to clipboard
| Challenge: | Product attribute extraction in e-commerce is bottlenecked by ontologies that are inconsistent, incomplete, and costly to maintain. |
| Approach: | They propose a multi-agent Large Language Model framework that constructs a Product-attribute Knowledge Graph from multimodal product content. |
| Outcome: | The proposed framework achieves 0.953 WKE for product types, 0.724 WKEs for attribute keys, and 0.531 edge-level accuracy for value assertions after canonicalization on a large real-world marketplace catalog dataset from Lazada (Alibaba). |
Copied to clipboard
| Challenge: | Extract-Refine-Retrieve-Read is a query optimization framework for large language models . it is designed to bridge the pre-retrieval information gap in Retriev-Augmented Generation systems . |
| Approach: | They propose a framework to extract parametric knowledge from Large Language Models and refine them using a specialized query optimizer. |
| Outcome: | The extract-refine-retrieve-read framework outperforms baselines on QA datasets . it is designed to meet the knowledge requirements of large language models (LLMs) |
Copied to clipboard
| Challenge: | Existing methods for multi-modal relation extraction lack useful visual information. |
| Approach: | They propose a novel multi-modal relation extraction framework to capture deeper correlations of text, entity pair, and image/objects. |
| Outcome: | The proposed framework captures the deeper correlations of text, entity pair, and image/objects, and extracts useful information. |
Copied to clipboard
| Challenge: | Existing alignment strategies that rely on discrete reranking struggle to address this granularity mismatch or effectively balance external evidence with internal knowledge. |
| Approach: | They propose a framework that synergizes discrete retrieval with continuous reranking to discern the information density differences between unstructured narrative passages and structured knowledge triplets. |
| Outcome: | Extensive experiments on four open-domain QA benchmarks show that AED-RAG significantly outperforms competitive baselines. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) are becoming tool-interactive agents . quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data . et al.: a new approach to scale trajectory diversity improves tool-use generalization . |
| Approach: | They propose a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. |
| Outcome: | Experiments on general tool-use benchmarks and code agent tasks show that TDScaling improves tool-user generalization and inherent coding proficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made remarkable advances in reasoning, yet continue to struggle with hallucinations, logical errors, and inability to self-correct during complex multi-step tasks. |
| Approach: | They propose a framework that enhances LLM reasoning through real-time environmental feedback validating each reasoning step, dynamic branch exploration for investigating alternative solution paths when faced with errors, and experience-based learning from successful reasoning trajectories. |
| Outcome: | The proposed model outperforms comparable models by 24.4 percentage points across benchmarks while outperforming comparable models. |
Copied to clipboard
| Challenge: | Recent Large Reasoning Models (LRMs) have demonstrated the ability to generate long chains of thought (LongCoT) LongCoT still faces challenges such as redundancy and logical incoherence. |
| Approach: | They propose a neural-symbolic reasoning framework that generates chains of thought . they propose Logic-Thinker, which transforms symbolic solvers into chains of thoughts . |
| Outcome: | The proposed framework outperforms models fine-tuned with ThinkerCoT on logic reasoning tasks. |
Copied to clipboard
| Challenge: | Recent studies reveal the risk of the model stealing attack, posing a financial threat to EaaS providers. |
| Approach: | They propose a dynamic embedding watermarking method that detects watermarks in embedded text . this method is a cross-platform approach that trains a verifier to detect watermark . |
| Outcome: | The proposed method enables an attacker to replicate the proposed method for profit without compromising embedding functionality. |
Copied to clipboard
| Challenge: | Existing approaches to incentivize LLMs’ deep thinking abilities require large-scale data or significant training efforts. |
| Approach: | They introduce an efficient framework that enhances LLM reasoning by teaching models to self-verify and self-correct during inference. |
| Outcome: | The proposed framework outperforms models trained on long-CoT distilled data with 3.1k initialization samples and achieves an accuracy improvement of 51.0% to 81.6%. |
Copied to clipboard
| Challenge: | Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings. |
| Approach: | They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data. |
| Outcome: | Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. |
| Approach: | They propose a framework to train critic models using refinement signals to generate feedback loops where critiques guide the model in refining its responses. |
| Outcome: | The proposed framework outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes. |
Copied to clipboard
| Challenge: | Tabular data analysis is an important application task of large language models, but advanced models are not yet on par with expert level performance. |
| Approach: | They propose to employ Large Language Models to facilitate an automated guide and execute high-precision data analyzes on tabular datasets. |
| Outcome: | The proposed framework is based on large language models and an automated machine learning pipeline for predictive modeling. |
Copied to clipboard
| Challenge: | Large language models (LLMs) traditionally represent text as sequences of discrete tokens . a long-context scaling problem requires processing more tokens more efficiently . |
| Approach: | They propose a framework that renders long texts into compact visual pages and processes them with a vision-language model. |
| Outcome: | The proposed framework renders long texts into compact visual pages and processes them with a vision-language model. |
Copied to clipboard
| Challenge: | Existing text embedding approaches often leverage the embeddment of the final token, typically a reserved special token such as ‘[EOS]‘. |
| Approach: | They propose to add a new training stage before contrastive learning to enrich the semantics of the final token embedding. |
| Outcome: | The proposed training stage improves performance on the Massive Text Embedding Benchmark (MTEB), achieving new state-of-the-art results across different LLM base models and scales. |
Copied to clipboard
| Challenge: | Existing research on building ES conversation systems only considered single-turn interactions with users, which is over-simplified and has limited support for multi-turn systems. |
| Approach: | They propose a multi-turn ES conversation system that uses lookahead heuristics to estimate future user feedback after using particular strategies. |
| Outcome: | The proposed system significantly outperforms baselines in both dialogue generation and strategy planning. |
Copied to clipboard
| Challenge: | Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy. |
| Approach: | They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions. |
| Outcome: | The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks. |
Copied to clipboard
| Challenge: | Theory of mind evaluations currently focus on testing models using machine-generated data or game settings prone to shortcuts and spurious correlations. |
| Approach: | They propose a benchmark to stress-test machine ToM in real-world negotiation surrounding covered multi-dimensional mental states. |
| Outcome: | The proposed benchmark builds upon the Belief-Desire-Intention theory and conducts the necessary empirical experiments to evaluate large language models. |
Copied to clipboard
| Challenge: | Existing supervised neural methods are underexplored for coreference resolution, especially in incremental clustering. |
| Approach: | They propose a dual-threshold incremental clustering approach based on a lightweight Transformer. |
| Outcome: | Experiments on common benchmarks show that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints. |
Copied to clipboard
| Challenge: | Existing methods to jailbreak Large Vision Language Models do not consider interaction between images and text. |
| Approach: | They propose a prior-guided bimodal interactive black-box jailbreak attack for toxicity maximization that exploits the interaction of images and text. |
| Outcome: | The proposed method outperforms state-of-the-art jailbreak methods in black box scenarios and in closed-source LVLMs. |
Copied to clipboard
| Challenge: | Existing answer selection approaches for community question answering lack additional answer summaries due to redundancy and lengthiness issues of crowdsourced answers. |
| Approach: | They constructed a dataset which contains a corresponding reference summary for each original lengthy answer. |
| Outcome: | The proposed model improves the performance of a question and candidate answer on a WikiHowQA dataset. |
Copied to clipboard
| Challenge: | Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, but studies on risks associated with cross biases are limited to immediate context preferences. |
| Approach: | They investigate multilingual bias in state-of-the-art Large Language Models by analyzing their responses to decision-making tasks across multiple languages. |
| Outcome: | The proposed model can provide personalized advice across university applications, travel, and relocation scenarios. |
Copied to clipboard
| Challenge: | Current Large Language Models (LLMs) lack self-awareness to balance reasoning and tool use, increasing computational overhead. |
| Approach: | They propose a paradigm that enhances an agent’s self-awareness to optimize task handling and reduce tool overuse. |
| Outcome: | The proposed model reduces tool use by 24% while improving performance by over 37%. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated notable capabilities across various tasks, showcasing complex problem-solving abilities. |
| Approach: | They propose a benchmark to evaluate the rule-based logical reasoning capabilities of Large Language Models (LLMs) they create simulated scenarios in which models execute or plan operations to achieve specific outcomes. |
| Outcome: | The proposed benchmark evaluates the performance of large language models on a variety of scenarios with varying difficulty levels. |
Copied to clipboard
| Challenge: | Empirical evidence indicates that Large Language Models exhibit spontaneous cross-lingual alignment in Information Extraction (IE) however, a significant imbalance across languages persists, highlighting an underlying deficiency. |
| Approach: | They propose a code LLM with advanced cross-lingual and multilingual capabilities for universal IE that standardizes the representation of multilingual schemas using Python classes and conducts IE alignment instruction tuning on translated instance prediction task. |
| Outcome: | The proposed model surpasses ChatGPT and SoTA by 30.17% without training in 29 unseen languages and significantly improves cross-lingual IE transferability. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable proficiency in handling a wide range of tasks within the software engineering domain, but their ability to perform code migration—adapting code to different environments—remains underexplored. |
| Approach: | They propose a benchmark to evaluate large language models’ performance in handling code migration tasks. |
| Outcome: | The proposed benchmark comprises 922 data points across 19 Python and Java packages and offers three tasks to systematically evaluate code migration: identifying version-incompatible functions, determining function changes, and adapting code to target environments. |
Copied to clipboard
| Challenge: | Questionnaires are a professional research methodology used for qualitative and quantitative analysis of human opinions, preferences, and behaviors. |
| Approach: | They propose a questionnaire-based dataset that consists of 13,168 human-written questionnaires. |
| Outcome: | The proposed dataset contains 13,168 human-written questionnaires gathered from online platforms. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can be fine tuned with human feedback, but human preferences can be diversified due to annotators’ different tastes, which hinders the effectiveness of LLM alignment methods. |
| Approach: | They propose a calibration error metric to evaluate large language models (LLMs) and a multi-objective reward learning method to enhance the calibration performance of RMs on shared preferences. |
| Outcome: | The proposed model can be adopted as a key calibration error and MORE can achieve superior alignment performance. |
Copied to clipboard
| Challenge: | Recent studies focus on retrieval to solve knowledge-intensive tasks, but the potential of retrieval for non-knowledge-intensive (NKI) tasks remains under-explored. |
| Approach: | They propose a task-agnostic retrieval framework for NKI tasks that uses a static index and a prompt-guided reranker to re-rank the nearest evidence according to task-specific relevance. |
| Outcome: | The proposed framework outperforms state-of-the-art retrieval-augmented methods on NKI tasks and will be released for further research. |
Copied to clipboard
| Challenge: | Empirical evaluations show that Mixture of Expert Prompt Tuning outperforms state-of-the-art parameter efficient baselines on SuperGLUE. |
| Approach: | They propose a pretrain-then-fine-tune paradigm for manifold mapping using multiple prompt experts. |
| Outcome: | Empirical results show that the proposed approach outperforms state-of-the-art methods on SuperGLUE while reducing activated prompts by 79.25%. |
Copied to clipboard
| Challenge: | Existing DA methods for named entity recognition (NER) are costly and labor-intensive to acquire, necessitating innovative approaches to data scarcity. |
| Approach: | They propose an order-agnostic data augmentation solution that exploits the order-based property in the training phase of sequence-to-sequence NER methods for data augmented. |
| Outcome: | The proposed method significantly enhances the few-shot capabilities of pre-trained language models in low-resource settings. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized the landscape of artificial intelligence. |
| Approach: | They propose a self-guided method to identify and select cherry samples from open-source datasets, minimizing manual curation and potential cost for instruction tuning an LLM. |
| Outcome: | The proposed method enables LLMs to identify discrepancies between expected responses and intrinsic generation capability, and a marked uptick in model training efficiency. |
Copied to clipboard
| Challenge: | Existing methods to edit multimodal models have been used to incrementally infuse a language model with a new set of facts. |
| Approach: | They construct a benchmark for editing multimodal Large Language Models and establish metrics for evaluation. |
| Outcome: | The proposed benchmarks show that editing multimodal models is not as difficult as editing single-modal models. |
Copied to clipboard
| Challenge: | Experimental results show that the model can be used to generate dialogues in new domains quickly. |
| Approach: | They propose to use LLMs to generate dialogue data to reduce dialogue collection and annotation costs. |
| Outcome: | The proposed model performs better than the baseline model trained on real data. |
Copied to clipboard
| Challenge: | a task-oriented dialogue system needs high conversational capability and can be easily adaptable to changing situations. |
| Approach: | They propose a retrieval-based conversational model that can rank a large set of responses . they propose supervised training and fine-tuning on limited data collected through a human-in-the-loop platform . |
| Outcome: | The proposed model can scale to rank a large set of responses in real-world situations. |
Copied to clipboard
| Challenge: | Recent studies focus on single-modality threats, but this approach fails to address cross-modal safety alignment. |
| Approach: | They propose a safety alignment challenge to evaluate cross-modality safety alignment . they propose 'Safe Inputs but Unsafe Output' to consider safety of single modalities . |
| Outcome: | The proposed safety alignment challenge examines cases where modalities are safe independently but could lead to unsafe outputs when combined. |
Copied to clipboard
| Challenge: | Recent advances in model editing for LLMs have created challenges and opportunities for the community. |
| Approach: | They propose to alter the behavior of LLMs efficiently within a specific domain without negatively impacting performance across other inputs. |
| Outcome: | The proposed method alters behavior of LLMs efficiently within a specific domain without negatively impacting performance across other inputs. |
Copied to clipboard
| Challenge: | Existing approaches focus on action selection or use pre-trained models as world models to enhance planning capabilities. |
| Approach: | They propose a new learning framework that optimizes state prediction and action selection through preference learning. |
| Outcome: | The proposed method outperforms existing methods and GPT-4o on VoTa-Bench and Qwen2-VL (7B), LLaVA-1.6 (7B) and LLama-3.2 (11B). |
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated RL's substantial capacity to enhance multi-step reasoning beyond what supervised instruction tuning achieves. |
| Approach: | They propose a framework that converts multimodal questions into descriptive text . they propose RL-enhanced geoscience reasoning that can be fine-tuned to a text-only level . |
| Outcome: | The proposed framework improves accuracy and accuracy on multimodal questions while preserving answerability and difficulty. |
Copied to clipboard
| Challenge: | Text-to-audio (T2A) models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio. |
| Approach: | They propose to use AI feedback learning to enhance basic capabilities of text-to-audio models . they use a large audio preference dataset to evaluate the model's capabilities . |
| Outcome: | The proposed model improves in simple and complex scenarios with AI feedback learning. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Dynamic nature of language limits the adaptability of Large Language Models (LLMs) Traditionally, LLMs are trained on static data, which limits their adaptability . |
| Approach: | They propose a benchmark to integrate novel data and assess LLMs’ ability to comprehend emerging concepts, alongside a causal inference-based approach to enhance LLM comprehension of new phrases and their colloquial context. |
| Outcome: | The proposed model outperforms baseline models in terms of precision and relevance in the comprehension of Internet slang and memes. |
Copied to clipboard
| Challenge: | Existing memory systems can support long-horizon human-LLM interactions by persisting historical interactions beyond limited context windows. |
| Approach: | They propose a framework that augments memory systems with a self-evolving meta-memory . meta-meso is iteratively distilling transferable knowledge utilization experiences . results show MetaMem outperforms strong baselines by over 3.6% . |
| Outcome: | The proposed framework outperforms baselines by over 3.6% in the long-horizon human-LLM interaction. |
Copied to clipboard
| Challenge: | a small model can be used to select effective prompt templates for a larger model. |
| Approach: | They propose a method to use a smaller model to select effective prompt templates for a larger model. |
| Outcome: | The proposed method significantly reduces the cost of prompt engineering while matching performance with optimal prompts among candidates. |
Copied to clipboard
| Challenge: | Autoregressive (AR) language models are a dominant paradigm in the field of parallelism and non-causal modeling. |
| Approach: | They propose a blockwise discrete diffusion model that preserves AR-compatible serving while enabling parallel intra-block generation. |
| Outcome: | The proposed model achieves theoretical speedups over 5 and wall-clock speedup of 2.3 on H200 GPUs in latency-critical regimes. |
Copied to clipboard
| Challenge: | Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes. |
| Approach: | They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning. |
| Outcome: | The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench. |
Copied to clipboard
| Challenge: | Large language models are often not well aligned with human intents, which requires additional training. |
| Approach: | They propose to use Black-Box Prompt Optimization (BPO) to perform alignments on large language models that are not well aligned with human intents. |
| Outcome: | The proposed model outperforms existing models and is model-agnostic. |
Copied to clipboard
| Challenge: | Existing methods to generate RAG documents require knowledge of the target RAG system’s internal composition and implementation details, whereas black-box methods are unable to utilize interactive information. |
| Approach: | They propose a RIPRAG attack framework that treats the target RAG system as a black box and leverages a Reinforcement Learning from Black-box Feedback (RLBF) method to optimize the generation model for poisoned documents. |
| Outcome: | The proposed method achieves an attack success rate (ASR) improvement of up to 0.72 compared to baseline methods. |
Copied to clipboard
| Challenge: | Modern large language models (LLMs) employ diverse logical inference mechanisms for reasoning. |
| Approach: | They analyze the comparative dynamics of inductive (System 1) versus abductive/deductive (system 2) inference in large language models by using a controlled analogical reasoning environment and a MCQ/free-text task format. |
| Outcome: | The proposed methods can significantly scale LLM reasoning. |
Copied to clipboard
| Challenge: | Current methods for harmful meme detection lack the knowledge required to identify such hate . current methods lack the ability to identify cultural stereotypes and visual metaphors . |
| Approach: | They propose a framework that decomposes meme analysis into a human-inspired reasoning process . they propose DR-HM to transfer knowledge from closed-source models while mitigating biases . |
| Outcome: | The proposed framework outperforms existing methods on three benchmark datasets. |
Copied to clipboard
| Challenge: | Full-duplex spoken dialogue systems allow simultaneous bidirectional communication . low latency and natural interactions in full-duplice systems remains a challenge . |
| Approach: | They propose a multi-stage post-training scheme that adapts a text large language model into a speech-text dialogue LLM. |
| Outcome: | The proposed model can model human conversation behaviors with low latency and natural interactions with low delay. |
Copied to clipboard
| Challenge: | a production-grade pricing system for tourism is challenging due to unstructured nature of travel orders and ever-evolving pricing policies. |
| Approach: | They propose a production-grade pricing system with a strict decision boundary . they propose to combine structured extraction and bounded policy/path selection with interpretable condition trees . |
| Outcome: | The proposed system processed 3,960 orders in six months and reduced the order management team from 15-20 to 3 . the system reduced the per-order handling time from 10 minutes to 2 minutes. |
Copied to clipboard
| Challenge: | ECom-Bench is a benchmark framework for evaluating LLM agent with multimodal capabilities in e-commerce customer support domain. |
| Approach: | They introduce a benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. |
| Outcome: | The proposed benchmark features dynamic user simulation based on persona information from real e-commerce customer interactions and a realistic task dataset derived from authentic ecommerce dialogues. |
Copied to clipboard
| Challenge: | Earlier studies on controlling styles in neural machine translation (NMT) have focused on regulating the level of formality, but they still encounter two major challenges. |
| Approach: | They propose a method to control the style of neural machine translation by retrieving prompts from stylized monolingual corpus. |
| Outcome: | The proposed method can control the style of translation and achieve remarkable performance. |
Copied to clipboard
| Challenge: | Lossless compression has made significant advancements in Genomics Data storage, sharing and management. |
| Approach: | They propose a novel agent-based GD Compressor with 3 layers with a multi-agent named Leader and Worker. |
| Outcome: | The proposed method improves on existing methods with low-level modeling and limited adaptability and user-unfriendly interface. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved significant success in reasoning tasks, including mathematical reasoning and logical deduction. |
| Approach: | They conduct the first comprehensive analysis of how the order of graph descriptions impacts LLM performance. |
| Outcome: | The results show that graph descriptions significantly improve LLMs’ comprehension of graph structures, and the robustness of LLM models to graph description order varies across different tasks. |
Copied to clipboard
| Challenge: | Large language models have shown tremendous success in following user instructions and generating helpful responses, but their robustness is still far from optimal. |
| Approach: | They propose a two-stage training framework that helps a model generalize on following instructions via similar instruction augmentations. |
| Outcome: | The proposed training framework improves diversity and aligns the model with human expectations by differentiating subtle differences in similar responses. |
Copied to clipboard
| Challenge: | Existing methods train one encoder-decoder-based model to fit all questions . however, such a one-size-fits-all strategy may not perform well for complex questions involving multiple KB relations or functional constraints. |
| Approach: | They propose a meta-learning framework for complex question generation over knowledge bases . they propose he meta-trained generator can acquire universal meta-knowledge . |
| Outcome: | The proposed framework can acquire universal and transferable meta-knowledge and quickly adapt to long-tailed samples under different dimensions. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are too large to be fine-tuned with budget constraints and some are only accessible via APIs. |
| Approach: | They propose a pluggable Reward-Driven Contextual Adapter that integrates large language models as generators and trains them to refine the retrieved information. |
| Outcome: | The proposed method improves ReQA performance on three datasets by up to 20% compared to existing methods. |
Copied to clipboard
| Challenge: | Document editing requires full-context awareness of dependencies, but processing entire documents for each edit incurs prohibitive token costs and latency. |
| Approach: | a framework that constructs lightweight dependency graphs captures semantic relationships and structural hierarchies across document elements is proposed for agentic document editing . a scaLing agentic agentic framework is based on a dependency graph framework that captures dependencies and refactors function dependencies. |
| Outcome: | a new framework achieves 76 consistency versus 56 baseline while reducing token usage by 85 . the framework is based on a framework that captures semantic relationships and structural hierarchies across document elements . it can be used to improve document consistency, but it also reduces token costs and latency . |
Copied to clipboard
| Challenge: | Existing methods for decomposing fine-tuned LLMs are sensitive to the magnitude of delta values. |
| Approach: | They propose a hierarchical quantization framework that shares low-bit integer weights across similar models. |
| Outcome: | The proposed framework achieves an average accuracy degradation of approximately 3% on fine-tuned models across mathematics, coding, chatbot, and Chinese LLMs. |
Copied to clipboard
| Challenge: | Existing approaches to exploit LLMs' inherent safety mechanism, including GCG and AutoDAN, are ineffective for certain malicious requests. |
| Approach: | They propose a method that generates jailbreak prompts to suppress a refusal stance and induce affirmative responses by modifying adversarial prompts. |
| Outcome: | The proposed method outperforms the best baseline approach in Llama-2-7b-chat and achieves a 92.2% success rate across all models. |
Copied to clipboard
| Challenge: | Extensive experiments on real-world datasets demonstrate that DPL significantly enhances LLM personalization. |
| Approach: | They propose a novel approach that emphasizes extracting inter-user differences to enhance LLM personalization. |
| Outcome: | The proposed approach extracts inter-user differences to enhance LLM personalization. |
Copied to clipboard
| Challenge: | Existing knowledge graphs lack the ability to integrate structural information into LLMs and output predictions deterministically. |
| Approach: | They propose a method which encodes structural information of KGs and merges it with LLMs to enhance KGC performance. |
| Outcome: | The proposed method improves the performance of KG Completion datasets on KGs by integrating structural information with LLMs. |
Copied to clipboard
| Challenge: | Existing neural dialogue models only capture syntactic and semantic information, but fail to model the logical consistency between the dialogue history and the generated response. |
| Approach: | They propose a fine-grained comparison model to capture syntactic and semantic information and then compare each candidate's representation with the whole history to obtain a history consistency representation. |
| Outcome: | The proposed model obtains higher ranking scores than baseline models on two public dialogue datasets. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) lack visual-aware tutorial retrieval and historical visual context curation and pruning. |
| Approach: | They propose a framework that integrates an orchestrator and a Reflection-Memory Agent for robust automation. |
| Outcome: | Experimental results show that OS-Symphony delivers substantial performance gains across model scales. |
Copied to clipboard
| Challenge: | Existing models to pretrain sentence encoders with large unlabeled corpus are lacking in linguistic information retrieval. |
| Approach: | They propose a novel approach to pre-training sequence encoder using transformers . they propose to train a Transformer-based sequence encoded over a large set of short sequences based on a set of masked words . |
| Outcome: | The proposed approach outperforms state-of-the-art encoders on hotpotQA by improving intermediate information retrieval performance. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have a high level of advanced reasoning capabilities, but they are vulnerable and vulnerable. |
| Approach: | This paper presents the first comprehensive survey of Large Reasoning Models . it explores the new safety risks, attacks, and defense strategies specific to LRMs based on reasoning . |
| Outcome: | The proposed study examines the safety and security risks of large reasoning models. |
Copied to clipboard
| Challenge: | Pre-trained language models have strong generalizability, but fine-tuning involves updating all parameters, rendering full fine-uning resource-intensive. |
| Approach: | They propose a parameter-efficient fine-tuning method that updates all pre-trained parameters during inference. |
| Outcome: | The proposed method outperforms baseline methods on five benchmarks across 20 datasets. |
Copied to clipboard
| Challenge: | Existing methods for pre-training KEPLMs with relational triples are difficult to adapt to close domains due to the lack of sufficient domain graph semantics. |
| Approach: | They propose a Knowledge-enhanced language representation learning framework for various closed domains that captures the implicit graph structure among the entities. |
| Outcome: | The proposed framework outperforms existing methods for pre-training KEPLMs in closed domains significantly. |
Copied to clipboard
| Challenge: | Recent studies have incorporated reward models to guide response selection or decoding, aiming to obtain higher-quality data. |
| Approach: | They propose a Hierarchical Sampling framework for self-taught reasoners that allocates a fixed sampling budget to problem boundary-level problems and then reallocates the remaining budget toward high-utility problems during a re-sampling phase. |
| Outcome: | The proposed framework outperforms baseline models without additional sampling budgets across multiple reasoning benchmarks and backbone LLMs. |
Copied to clipboard
| Challenge: | Mainstream speaker diarization systems rely only on acoustic information, making it challenging in complex aural environments. |
| Approach: | They propose a multimodal approach that integrates audio, visual, and semantic cues to enhance speaker diarization. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on multi-party conversations . it integrates audio-visual-semantic cues into the clustering process for acoustic speaker embeddings . |
Copied to clipboard
| Challenge: | Existing approaches to large vision-language models fail to capture interleaved nature of human visual-verbal reasoning processes. |
| Approach: | They propose a framework that integrates visuospatial and linguistic domains to facilitate multimodal slow thinking by enabling progressive visual-textual reasoning. |
| Outcome: | Experiments show that VisuoThink significantly improves reasoning capabilities even without fine-tuning. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained models are limited due to suboptimal activation subspaces. |
| Approach: | They propose a method that leverages tail eigenvectors of model output activations to construct low-rank adapters. |
| Outcome: | The proposed method outperforms existing methods across 16 benchmarks and surpasses full fine-tuning in certain scenarios. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) has emerged as a significant advancement in the field of large language models (LLMs). |
| Approach: | They propose a method that uses hallucination detection labels to correct hallucines by integrating up-to-date information into their initial training. |
| Outcome: | The proposed method is based on the Retrieval Augmented Generation (RAG) method, which has shown to be effective in mitigating hallucinations and improving answer quality. |
Copied to clipboard
| Challenge: | Existing role-playing models rely on superficial textual descriptions or simplistic metrics, inadequately modeling both intrinsic and extrinsic character dimensions. |
| Approach: | They propose a framework that integrates fine-grained psychological attributes and explicit memory control for role-playing. |
| Outcome: | The proposed framework outperforms baseline models in human-likeness and character fidelity. |
Copied to clipboard
| Challenge: | Large pre-trained language models have demonstrated impressive capabilities, but there is still much to learn about how they operate. |
| Approach: | They investigate the ability of the autoregressive transformer to perform basic addition operations by using causal analysis to find that a few different attention heads in the middle layers control the addition carry . they found that due to the lack of global focus on the sequence within these attention heads, the model struggles to handle long-sequence addition tasks. |
| Outcome: | The model performs basic addition tasks, but it still faces challenges with length generalization. |
Copied to clipboard
| Challenge: | Existing methods focus on knowledge and linguistic patterns of characters. |
| Approach: | They propose to evaluate character fidelity of role-playing agents with psychological scales . they propose to use psychological scale to measure personality traits of RPAs based on personality traits. |
| Outcome: | The proposed model reproduces character fidelity with psychological scales and shows that it is effective in measuring personality traits. |
Copied to clipboard
| Challenge: | Existing multi-modal dialogue models are limited to incapacity of reading visual information and multi-dimensional interactions. |
| Approach: | They propose a novel event-oriented video-dialogue dataset called SportsVD to overcome these challenges by generating human-like response according to event contents in the video and related external knowledge. |
| Outcome: | The proposed method outperforms existing methods on SportsVD and other baselines under several automatic metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. |
| Approach: | They propose a benchmark to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. |
| Outcome: | The proposed framework measures task completion and quality of collaboration and competition using novel, milestone-based key performance indicators. |
Copied to clipboard
| Challenge: | “Jailbreak” is a major safety concern of Large Language Models (LLMs). |
| Approach: | They propose a benchmarking framework to evaluate "jailbreak" outputs . they propose specialized validation framework to ensure outputs are useful malicious instructions . |
| Outcome: | The proposed framework enhances existing benchmarks to ensure outputs are useful . it also aims to evaluate the true potential of jailbroken outputs to cause harm to human society. |
Copied to clipboard
| Challenge: | Symbols are used in abstract reasoning, chemical property prediction, and tabular question-answering. |
| Approach: | They propose a method that converts symbols to language-based representations to improve their accuracy. |
| Outcome: | The proposed method improves the accuracy of symbols in language-based models. |
Copied to clipboard
| Challenge: | a novel generalization framework for visual temporal-aligned translation is proposed to transfer recognition skills to unseen performers . ambiguity in the visual sequence can hinder current methods for visual language translation . |
| Approach: | They propose a generalizable framework to transfer recognition skills to unseen performers . they use visual temporal-aligned translation to generate multiple words autoregressively . |
| Outcome: | The proposed framework is generalized to transfer recognition skills to unseen performers . it is compared with existing methods on lipreading and fingerspelling datasets . |
Copied to clipboard
| Challenge: | Speculative Safety-Aware Decoding (SSD) equips large language models with desired safety property while accelerating inference. |
| Approach: | They propose a lightweight decoding-time approach that equips large models with the desired safety property while accelerating inference. |
| Outcome: | Experimental results show that a small language model has the desired safety property while accelerating inference. |
Copied to clipboard
| Challenge: | LLMEdgeRefine is an iterative clustering method enhanced by large language models . existing clustering methods struggle with domain-specific fine-tuning and outliers . |
| Approach: | They propose an iterative clustering method enhanced by large language models focusing on edge points refinement . authors propose to use LLMs to iterate clusters and iterating to improve semantic coherence . |
| Outcome: | The proposed method outperforms state-of-the-art methods and offers robustness, adaptability, and cost-efficiency for diverse text clustering applications. |
Copied to clipboard
| Challenge: | Existing detection methods fail to account for **self-consistent error** . study identifies self-consistency errors and evaluates them . |
| Approach: | They propose a method that fuses hidden state evidence from an external verifier LLM to detect self-consistent errors. |
| Outcome: | The proposed method significantly enhances performance on self-consistent errors across three LLM families. |
Copied to clipboard
| Challenge: | Existing supervised neural methods for coreference resolution are underexplored . current methods rely on small language models, but their potential is underexploited . |
| Approach: | They propose a framework that integrates an enhanced supervised model with LLM-based reasoning. |
| Outcome: | The proposed method surpasses existing state-of-the-art methods in coreference resolution. |
Copied to clipboard
| Challenge: | Existing methods for speculative decoding incur substantial training overhead to mitigate information misalignment between autoregressive draft model training and decoding. |
| Approach: | They propose an Entropy-Driven Speculative Decoding framework that uses entropy as a unified, interpretable signal for both draft model training and architectural design. |
| Outcome: | Experiments on seven large language models show that EDSD improves training efficiency by 24.8% and increases acceptance length by 4.0% compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing methods for designing and optimizing multi-agent systems are static and do not learn from experience. |
| Approach: | They propose a framework that enables a multi-agent system to learn to evolve . they use "textual gradients" to pinpoint failures and suggest granular modifications . |
| Outcome: | a new framework enables a multi-agent system to learn to evolve . it learns from historical optimization experiences to improve its performance . |
Copied to clipboard
| Challenge: | Existing block-wise discrete diffusion models lack robust autoregressive (AR) decoders. |
| Approach: | They propose a block-wise discrete diffusion framework for large-scale vision-language understanding with a progressive beta noise curriculum. |
| Outcome: | The proposed framework improves training efficiency, convergence stability, and task performance over conventional block diffusion. |
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. |
| Approach: | They propose a framework which calibrates bias and errors and achieves adaptive-balanced decoupling training by a prototype-based loss model. |
| Outcome: | The proposed framework outperforms existing approaches and achieves state-of-the-art performance on three datasets. |
Copied to clipboard
| Challenge: | Recent advances in augmenting Large Language Models (LLMs) with auxiliary information have significantly revolutionized their efficacy in knowledge-intensive tasks. |
| Approach: | They propose a systematic framework to identify whether LLMs’ responses are attributed to either generated or retrieved contexts. |
| Outcome: | The proposed framework identifies whether LLMs’ responses are attributed to either generated or retrieved contexts. |
Copied to clipboard
| Challenge: | Existing jailbreaking methods view a malicious prompt as a whole but they are not effective at reducing LLMs’ attention on combinations of words with malice. |
| Approach: | They propose an automatic prompt Decomposition and Reconstruction framework for jailbreaking Attack that decomposes a malicious prompt into separate sub-prompts and reassembles them implicitly by In-Context Learning. |
| Outcome: | The proposed framework reduces LLMs' attention on malice words by presenting them to LLM in a fragmented form, addressing these limitations and improving attack effectiveness. |
Copied to clipboard
| Challenge: | Out of Scope (OOS) detection is a problem with chatbots that cannot make sense of a query . a real-world solution to this problem is to identify out-of-domain queries . |
| Approach: | They propose a simple yet effective OOS detection method that outperforms standard methods . they propose analyzing data from an enterprise virtual assistant platform to test the method . |
| Outcome: | The proposed method outperforms standard OOS detection methods in a real-world deployment of virtual assistants. |
Copied to clipboard
| Challenge: | Existing studies focus on prompt engineering to encode the full semantics of a sentence into the embedding of the last token. |
| Approach: | They propose a technique that introduces an extra auxiliary prompt to elicit better sentence embedding . they propose to use the hidden state of the token as the sentence embedded in LLMs . |
| Outcome: | The proposed technique can improve performance of existing prompt-based methods on STS tasks and downstream classification tasks. |
Copied to clipboard
| Challenge: | Existing methods for multimodal entity linking rely on mention words as retrieval cues, which limits their ability to effectively utilize information from both images and text. |
| Approach: | They propose a visual prompt-guided multimodal entity linking task for a text-image pair . they propose VPWiki to facilitate this task and a framework to capture latent information. |
| Outcome: | The proposed framework outperforms baseline methods on a VPWiki dataset. |
Copied to clipboard
| Challenge: | Graphical User Interface (GUI) agents powered by Vision-Language Models (VLMs) have demonstrated human-like computer control capability. |
| Approach: | They propose a GUI data synthesis pipeline that reverse engineers GUI trajectory construction process by executing pre-defined tasks. |
| Outcome: | The proposed GUI data synthesis pipeline overcomes the bottlenecks of previous methods that rely on pre-defined tasks and limited data diversity. |
Copied to clipboard
| Challenge: | Existing methods for talking head translation rely on cascading, resulting in delays and cascadic errors. |
| Approach: | They propose a model for talking head translation, TransFace, which can translate audio-visual speech into audio-visual speech in other languages. |
| Outcome: | The proposed model can translate audio-visual speech into audio-visual speech in other languages. |
Copied to clipboard
| Challenge: | Existing studies on table reasoning focus on flat tables and hierarchical tables . a new dataset, HiTab, aims to examine numerical reasoning over hierarchic tables based on hierarchically structured tables - a strong challenge for existing baselines and a valuable benchmark for future research. |
| Approach: | They propose a hierarchical question answering and natural language generation dataset to study hierarchic tables. |
| Outcome: | The proposed model shows that it is effective in QA and natural language generation over hierarchical tables. |
Copied to clipboard
| Challenge: | Open benchmarks are essential for evaluating large language models, but their accessibility makes them likely targets of test set contamination. |
| Approach: | They propose a framework that leverages backdoor attacks to flag models that used benchmark test sets during training. |
| Outcome: | The proposed framework detects models that trained on benchmark test sets without loss of logits or internal details . it can prevent false accusations while providing strong evidence for every detected case of contamination. |
Copied to clipboard
| Challenge: | Existing studies show that Mamba architectures have room for further optimization in linear projections and state caches. |
| Approach: | They propose a decoupled scale quantization scheme to mitigate outliers in states and channels by applying separate quantization scales. |
| Outcome: | The proposed method reduces memory consumption by 50% across various quantization settings, model sizes, and generation and zero-shot tasks. |
Copied to clipboard
| Challenge: | Existing frameworks that focus on static tools and static assets are ineffective for self-evolving agents. |
| Approach: | They propose a paradigm of co-evolutionary Capability Expansion and Experience Distillation that leverages accumulated experience to guide dynamic creation of assets. |
| Outcome: | The proposed framework improves performance in single-task and cross-task settings by 18.53% over standard LLMs, 11.80% over agents evolving solely through experience, and 6.46% over those evolving solelly through asset creation. |
Copied to clipboard
| Challenge: | Existing studies on large language models have limited evaluation of their geospatial cognition . a unified framework for evaluating geospcial cognition in LLMs remains absent . |
| Approach: | They propose a benchmark to evaluate the geospatial route cognition of Large Language Models . they propose 'pathbuilder' tool for converting natural language instructions into navigation routes . |
| Outcome: | The proposed framework and metrics evaluate 9 state-of-the-art LLMs on route reversal task. |
Copied to clipboard
| Challenge: | Large Audio-Language Models (LALMs) are augmented with the ability to perceive audio, but their reliability when faced with conflicting inputs remains largely unexplored. |
| Approach: | They examine how LALMs prioritize information when presented with inconsistent audio-text pairs. |
| Outcome: | The proposed models display a significant bias toward textual input when presented with inconsistent audio-text pairs. |
Copied to clipboard
| Challenge: | Recent approaches to enhance agent performance focus on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals. |
| Approach: | They propose a step-level framework that provides detailed step-by-step guidance to enhance agent training by using Monte Carlo methods. |
| Outcome: | The proposed framework outperforms strong baselines on three tasks and shows that it is effective in augmenting efficiency and its applicability to diverse models. |
Copied to clipboard
| Challenge: | Global scientific publications are growing annually by about 4%-5% (Pinedo et al., 2024). |
| Approach: | They introduce an AI-assisted platform that answers diverse questions from researchers using Retrieval-Augmented Generation (RAG) they develop various tools to understand queries, search from the scientific literature, filter retrieved information, provide accurate and comprehensive answers, and self-refine answers. |
| Outcome: | OpenResearcher is built on Retrieval-Augmented Generation (RAG) to integrate Large Language Models (LLMs) with up-to-date, domain-specific knowledge. |
Copied to clipboard
| Challenge: | Large Language Models have exceptional capabilities in open generation, yet they encounter difficulties with tasks that require intensive knowledge. |
| Approach: | They propose a framework that integrates unknown knowledge into LLMs without overlap . they propose integrating domain-specific knowledge graphs into Llms to reduce knowledge forgetting . |
| Outcome: | The proposed framework outperforms state-of-the-art baselines in integrating new knowledge into LLMs. |
Copied to clipboard
| Challenge: | Existing decoding strategies and hyperparameters may not be optimal for each sample. |
| Approach: | They propose a model that auto-regulates decoding strategies and hyperparameters . this approach eliminates the need for extensive manual tuning, they argue . |
| Outcome: | The proposed model eliminates the need for extensive manual tuning, offering a more autonomous, self-regulate model behavior. |
Copied to clipboard
| Challenge: | Numerical reasoning requires both natural language understanding and arithmetic computation. |
| Approach: | They propose a graph representation for the context of the passage and question needed for numerical reasoning. |
| Outcome: | The proposed model achieves remarkable results in benchmark datasets such as DROP. |
Copied to clipboard
| Challenge: | Existing methods for event detection use sentence-level contextual information. |
| Approach: | They propose a document embedding enhanced bi-RNN model to detect events in sentences . they use hierarchical and supervised attention based RNN to learn document embeds . |
| Outcome: | The proposed model compares with state-of-the-art models on a ACE-2005 dataset. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) suffer from knowledge cutoff or fallacy issues, which means they are unaware of unseen events or generate text with incorrect facts owing to outdated/noisy data. |
| Approach: | They propose an easy-to-use knowledge editing framework for Large Language Models that allows users to easily edit updated knowledge and adjust undesired behavior while minimizing the impact on unrelated inputs. |
| Outcome: | The proposed framework surpasses traditional fine-tuning in terms of reliability and generalization. |
Copied to clipboard
| Challenge: | Traditional generation methods focus primarily on textual quality, but they fail to meet complex, multifaceted educational requirements. |
| Approach: | They propose a method for automatic generating high-quality mathematical problems that align with educational objectives using a dataset of 16k mathematical questions with multi-dimensional educational objectives. |
| Outcome: | The proposed method improves generating high-quality mathematical questions that meet multi-dimensional educational objectives. |
Copied to clipboard
| Challenge: | Earlier studies of instruction tuning on Large Language Models focus on creating large, varied, and high-quality datasets with responses curated by human experts. |
| Approach: | They propose to use a smaller and weaker model to fine tune a larger and stronger model . they find it can largely speed up the data filtering and improve performance . |
| Outcome: | The proposed model can filter instruction data faster and better on benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning. |
| Approach: | They propose a virtual API server and stable evaluation system to assess the stability of large-scale real-time APIs. |
| Outcome: | The proposed benchmarks demonstrate the stability of the proposed system and its caching system. |
Copied to clipboard
| Challenge: | Existing methods for crafting adversarial passages are slow and computationally expensive, requiring either access to retriever’s gradients or large computational resources. |
| Approach: | They propose a method that leverages two key characteristics of retrievers: insensitivity to token order and bias towards influential tokens to generate effective adversarial passages. |
| Outcome: | The proposed method achieves superior efficiency and scalability compared to existing methods while maintaining comparable or better attack success rates across multiple datasets. |
Copied to clipboard
| Challenge: | Using curated datasets, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 105 times more computationally efficient than retraining. |
| Approach: | They propose a framework for machine unlearning in pre-trained LLMs and integrate gradient ascent with gradient descent on in-distribution data to achieve robustness. |
| Outcome: | The proposed framework is over 105 times more efficient than retraining on in-distribution data and provides detailed guidelines for efficient hyperparameter tuning in the unlearning process. |
Copied to clipboard
| Challenge: | Existing methods for multi-class sentiment analysis (MCSA) are difficult due to subtle semantic differences between adjacent sentiment levels and the scarcity of high-quality annotated data. |
| Approach: | They propose a framework to integrate classification rationales with adaptively selected demonstrations to enhance MCSA performance under limited supervision. |
| Outcome: | The proposed framework outperforms baseline and standard ICL methods on five benchmark datasets. |
Copied to clipboard
| Challenge: | Positional biases in large language models hinder their ability to process long inputs. |
| Approach: | They propose a benchmark to assess positional bias in large language models involving multiple pieces of relevant information. |
| Outcome: | The proposed benchmark assesses the performance of long-context language models by examining their models with different input lengths and tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are capable of generating inaccurate discharge summary content or fabricating information without valid sources. |
| Approach: | They propose a tool for empowering LLMs with Logic-Controlled Discharge Summary generation. |
| Outcome: | The proposed tool identifies the writing logic of discharge summaries and integrates it with EMRs to generate silver discharge summararies. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have shown promising results in zero-shot settings, which motivates us to explore prompt-based methods. |
| Approach: | They propose a two-stage framework which transforms the SLU task into a question-answering problem by directly prompting LLMs. |
| Outcome: | The proposed framework can be built by directly prompting LLMs to understand user needs without training data. |
Copied to clipboard
| Challenge: | Autoregressive language models excel in text-to-audio generation, but lag behind diffusion models by a non-trivial margin. |
| Approach: | They propose a framework that integrates multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. |
| Outcome: | The proposed framework outperforms existing LM-based and diffusion-based systems in audio synthesis. |
Copied to clipboard
| Challenge: | Existing studies have focused on the alignment of multimodal sequential learning using transformers. |
| Approach: | They propose a constrained scheme to align the multiple attentional results from both local and global perspectives. |
| Outcome: | The proposed scheme could align the multiple attentional results from both local and global perspectives, making the information capture more efficient. |
Copied to clipboard
| Challenge: | Recent literature reveals that Large Language Models (LLMs) hallucinate intermittently, which impedes their reliability for further utilization. |
| Approach: | They propose a self-detection method to detect which questions an LLM does not know by combining the two components to identify whether the model generates a non-factual response to the question. |
| Outcome: | The proposed method can detect which questions an LLM does not know across factoid question-answering, arithmetic reasoning, and commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | RAG systems that integrate external knowledge with Large Language Models often become bottlenecks due to their limited parameters compared to LLMs and their inability to perform step-by-step reasoning. |
| Approach: | They propose a model that integrates external knowledge with Large Language Models to enhance factual correctness and mitigate hallucination. |
| Outcome: | The proposed model outperforms baselines and can transfer well to different retrievers. |
Copied to clipboard
| Challenge: | Current language models lack the structured deliberation needed for high-stakes tasks such as healthcare and finance. |
| Approach: | They propose a decision-making framework that guides models to reason over structured representations of actions, attributes, and constraints. |
| Outcome: | The proposed framework achieves up to 30% accuracy gains over strong prompting baselines and enhances alignment in outcomes. |
Copied to clipboard
| Challenge: | Experimental results show VocalNet outperforms existing open-source speech LLMs despite limited training data. |
| Approach: | They propose a scalable and model-agnostic training framework and a novel multi-token prediction paradigm for speech generation. |
| Outcome: | The proposed model outperforms open-source speech LLMs while outperforming existing open-sourced models. |
Copied to clipboard
| Challenge: | Existing efforts to estimate document utility rely on downstream generation performance, which conflates the influence of external documents with the intrinsic knowledge of the LLM. |
| Approach: | They propose an uplift-based definition of document utility that quantifies each document’s marginal benefit over the LLM’s internal knowledge. |
| Outcome: | The proposed framework improves the performance of the LLM by incorporating external retrieved documents into the model. |
Copied to clipboard
| Challenge: | Natural language is the de facto communication medium for LLM-based agents, but it presents a fundamental constraint . natural language downsampling limits the depth and nuance of information that can be transmitted . et al.: inter-agent latent space communication is a promising paradigm for solving complex tasks . |
| Approach: | They propose a paradigm that leverages the last hidden states of an LLM as a representation of its thought for direct communication. |
| Outcome: | The proposed paradigm outperforms fine-tuned chain-of-thought prompting and single-agent baselines even across heterogeneous models. |
Copied to clipboard
| Challenge: | Conventional fine-tuning works through updating all of the parameters in the pre-trained model, but as the size of pre-train models grows, it can be time-consuming and computationally expensive. |
| Approach: | They propose a framework for resource- and parameter-efficient fine-tuning by leveraging the sparsity prior in both weight updates and the final model weights. |
| Outcome: | The proposed framework saves 25% inference FLOPs while maintaining competitive downstream performance. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating long-context language models employ irrelevant noise texts to artificially extend the length of test cases, diverging from the real-world scenarios of long-constituency applications. |
| Approach: | They propose a long-context benchmark, Loong, aligning with realistic scenarios through extended multi-document question answering (QA) . |
| Outcome: | The proposed model can scale up the context window of large language models to perform in-depth analysis of multiple long documents. |
Copied to clipboard
| Challenge: | Existing methods to generate toxic content in pretrained language models are resource-intensive and require additional components. |
| Approach: | They propose a method that enables the PLM itself to achieve "self-detoxification" they identify the toxification direction from the normal generation process to the one prompted with the negative prefix and then steer the generation to the reverse direction by manipulating the information movement within the attention layers. |
| Outcome: | The proposed method can achieve comparable performance with state-of-the-art methods without any fine-tuning or extra components. |
Copied to clipboard
| Challenge: | Existing studies show that training LLMs on data containing unfamiliar knowledge during instruction tuning can encourage hallucinations. |
| Approach: | They propose a framework that measures how familiar the LLM is with instruction data and introduce an expert-aligned reward model to ensure the quality of selected samples. |
| Outcome: | The proposed framework reduces hallucinations while maintaining a competitive ability to follow instructions. |
Copied to clipboard
| Challenge: | Recent advances in parameter-efficient fine-tuning techniques allow for adjustments to only a minor fraction of the parameters of language models. |
| Approach: | They propose a low-rank direct attention adapted method for efficient LLM fine-tuning . they propose LMAM, which can bring negative attention to self-attention modules . |
| Outcome: | The proposed method outperforms the full fine-tuning method by 2.1% on GLUE benchmark. |
Copied to clipboard
| Challenge: | Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. |
| Approach: | They propose a framework for massively multilingual table question answering that includes tables expanded to 97 languages from Chinese and English sources. |
| Outcome: | Experiments on state-of-the-art LLMs show that synthetically generated training data significantly boosts performance, especially for low-resource languages. |
Copied to clipboard
| Challenge: | Existing approaches to assess the risk of bias in RCTs focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability. |
| Approach: | They propose a framework for enhancing Large Language Models to assess the risk of bias in RCTs by reformulation, document parsing and multi-expert collaboration. |
| Outcome: | The proposed framework outperforms existing methods on the RoB-Item and RoB domains. |
Copied to clipboard
| Challenge: | Existing methods analyze training data with member and non-member contexts, overlooking potential insights from both member and not-member. |
| Approach: | They propose a method that leverages asymmetric distributional shifts induced by member and non-member contexts through contrastive decoding to enhance membership inference. |
| Outcome: | The proposed approach outperforms the current state-of-the-art on the WikiMIA benchmark and is robust against various text manipulation techniques. |
Copied to clipboard
| Challenge: | Language models contain ranking-based knowledge and are powerful solvers of in-context ranking tasks. |
| Approach: | They propose to use a model to elicit language models' ranking knowledge without supervision by using a pairwise, pointwise and listwise prompting method. |
| Outcome: | The proposed method is inspired by an unsupervised probing method called Contrast-Consistent Search (CCS). |
Copied to clipboard
| Challenge: | Large language model agents assume a stationary context, failing to account for real-world time elapsed between messages. |
| Approach: | They construct a dataset of multi-turn user–agent message trajectories across 76 scenarios . they collect human preferences between "calling a tool" and "directly answering" they also examine whether existing models lack human temporal perception . |
| Outcome: | The results show that existing models display poor alignment with human temporal perception . the findings provide insights to foster the development of more time-aware and human-aligned agents. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models fail to reflect real-world complexity . existing benchmarks often fail to capture real-life problems . |
| Approach: | They propose a benchmark that features real-world-inspired, open-ended problems from competitions . they propose 'ModelingBench' that supports multiple valid solutions . |
| Outcome: | The proposed framework outperforms baselines and produces well-grounded, creative solutions. |
Copied to clipboard
| Challenge: | Existing memory systems rely on summarization to preserve contextual nuances and obscuring key retrieval features. |
| Approach: | They propose a method that decouples the retrieval unit from the generation context. |
| Outcome: | The proposed method outperforms baseline models on the LoCoMo benchmark. |
Copied to clipboard
| Challenge: | Prior zero-shot TTS models only mimic the speaker’s voice without further control and adjustment capabilities while prior controllable TTS systems cannot perform speaker-specific voice generation. |
| Approach: | They propose a style control module that captures codec representations corresponding to timbre, content, and style in a discrete decoupling codec space. |
| Outcome: | The proposed system can fully clone the speaker's voice and perform speech-specific adjustment and control functions. |
Copied to clipboard
| Challenge: | Chain-of-Thought reasoning has driven recent gains of large language models (LLMs) on reasoning-intensive tasks by externalizing intermediate steps. |
| Approach: | They propose a training-free framework that adaptively determines when to stop reasoning to mitigate overthinking. |
| Outcome: | The proposed framework reduces token usage by 20-55% while maintaining or improving accuracy compared to standard CoT prompting. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have brought significant changes to various domains, especially through autonomous agents. |
| Approach: | They propose a framework that lets agents learn shortcuts from their past tasks and use them for future task execution. |
| Outcome: | The proposed framework enables agents to tackle unseen software-developing tasks more effectively. |
Copied to clipboard
| Challenge: | Existing methods for evaluation of large language models are inefficient and inefficient due to inaccuracy of standard metrics in human perception of text quality and inefficiency in sampling informative test examples. |
| Approach: | They propose a sample-efficient human evaluation method for large language models based on the principle of MAximum Discrepancy (MAD) competition. |
| Outcome: | The proposed method achieves the “golden” ranking of LLMs with a minimum set of input instructions, which in turn reveal their relative strengths and weaknesses. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have demonstrated proficiency in diverse tasks across different domains. |
| Approach: | They propose a method that integrates multimodal instruction tuning with Conditional Mixture-of-LoRA. |
| Outcome: | Experimental results show that MixLoRA outperforms LoRA with the same or higher ranks . MLLMs have demonstrated remarkable proficiency in diverse tasks across domains . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating long sequences. |
| Approach: | They propose a benchmark to evaluate LLM safety in open-ended long-context tasks . they find that relevant context and extended input sequences can exacerbate safety risks . |
| Outcome: | The proposed benchmark identifies significant safety vulnerabilities in 16 LLMs . strong safety performance in short-context scenarios does not correlate with safety in long-contact tasks . |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
Copied to clipboard
| Challenge: | Existing methods for visual storytelling construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content. |
| Approach: | They propose a topic description task to detect the global semantic context of an image stream and a story is then constructed with the guidance of the topic description. |
| Outcome: | The proposed framework can generate stories with higher quality compared to state-of-the-art methods on a VIST dataset. |
Copied to clipboard
| Challenge: | Low-resource languages, like Tibetan, remain underrepresented in large language models' evaluations. |
| Approach: | They propose a Tibetan Language Understanding Evaluation Benchmark to assess LLMs' proficiency in Tibetan . they use a multi-task understanding benchmark and a safety benchmark to evaluate models . |
| Outcome: | The proposed benchmark shows that most large language models perform below the random baseline, especially in Tibetan language processing. |
Copied to clipboard
| Challenge: | Existing Lifelong Knowledge Editing methods struggle to overwrite outdated knowledge with the latest one. |
| Approach: | They propose a new Mixture-of-Knowledge-Experts scheme with an ARM . ARM ensures that each update completely overwrites old information with the latest one . Experimental results show that ARM performs favorably against SOTA knowledge editing methods . |
| Outcome: | The proposed scheme overwrites old knowledge with the latest data on a benchmark . it performs favorably against existing knowledge editing methods on the same concept . |
Copied to clipboard
| Challenge: | MAS-BENCH isolates coordination under explicit communication constraints . CAMOC significantly improves coordination success and efficiency across backends . |
| Approach: | They propose a distributed-sorting benchmark that isolates coordination under explicit communication constraints. |
| Outcome: | MAS-BENCH improves coordination success and efficiency across backends . CAMOC significantly improves efficiency under shared-state interaction . |
Copied to clipboard
| Challenge: | Existing methods to locate code snippets from databases represent the semantics of code and query by averaging the features of each token and word. |
| Approach: | They propose a fine-grained code search model that consists of a cross-modal encoder, mapping layer and classification layer to capture fine-granular interactions between code and query. |
| Outcome: | The proposed model significantly outperforms existing methods across multiple programming language datasets. |
Copied to clipboard
| Challenge: | Existing methods for detecting fake news videos fall short due to lack of knowledge to verify the news is real or not. |
| Approach: | They propose a VLM-based framework for detecting fake news on short video platforms . they design four experts tailored to handle each scenario and integrate them into VLM . |
| Outcome: | The proposed framework outperforms current state-of-the-art models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly deployed in high-stakes domains such as finance where unsafe behavior can lead to serious regulatory risks. |
| Approach: | They propose a black-box multi-turn risk-concealed redteaming framework that progressively conceals surface-level risk while exploiting regulatory-violating behaviors. |
| Outcome: | Experiments on nine widely used LLMs show that the proposed framework achieves 93.19% average attack success rate (ASR) and improves the average ASR to 95.00%. |
Copied to clipboard
| Challenge: | Existing benchmarks rely on static datasets that degrade over time as models evolve, leading to data contamination and diminished evaluative power. |
| Approach: | They construct a framework that generates model-specific challenging datasets and aggregates them across diverse LLM families. |
| Outcome: | The framework captures emerging defensive patterns and aligns with each model’s unique behavior. |
Copied to clipboard
| Challenge: | Existing review-based recommenders favor large and complex language encoders that can only learn latent and uninterpretable text representations. |
| Approach: | They propose a tightly coupled two-stage approach to extract latent user sentiments and item properties from reviews and an Attention-Property-aware Rating Estimator (APRE). |
| Outcome: | Extensive experiments on seven real-world Amazon review datasets show that the proposed approach extracts the latent user sentiments, item properties, and the complicated interactions between the two components. |
Copied to clipboard
| Challenge: | Existing methods for semantic parsing rely on extensive manually annotated datasets and limited generalization capability to unseen examples. |
| Approach: | They propose a framework that generates high-relevance synthetic data without manual annotation . they generate queries for the queries and use them as demonstrations for in-context learning . |
| Outcome: | The proposed framework outperforms non-fine-tuned methods on KBQA datasets and shows superior sample efficiency, robustness, and generalization capabilities under non-I.I.D. settings. |
Copied to clipboard
| Challenge: | Pre-training large language models can be expensive and wasteful. |
| Approach: | They propose a method which can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and a two-stage learning method to further accelerate the pre-training. |
| Outcome: | The proposed method can transfer the knowledge of an existing smaller pre-trained model to a large model through parameter initialization and significantly improve the pre-training efficiency of the large model. |
Copied to clipboard
| Challenge: | Large Language Model Unlearning (LLMU) is a promising way to remove private or sensitive information from large language models. |
| Approach: | They propose a Fully Probabilistic Evaluation framework that incorporates input and output distributions in LLMU evaluation. |
| Outcome: | The proposed framework improves unlearning effectiveness by 50.1% and robustness by 37.2% on Llama-2-7B. |
Copied to clipboard
| Challenge: | CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview. |
| Approach: | They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
| Outcome: | The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
Copied to clipboard
| Challenge: | Existing benchmarks assess basic Theory of Mind abilities but neglect temporal evolution of mental states in real-world social contexts. |
| Approach: | They propose a benchmark specifically designed to evaluate Large Language Models' ability to understand and track the temporal progression of mental states across interconnected scenarios. |
| Outcome: | The proposed benchmarks underperform humans by 44.7% and show that they can model the dynamic nature of human mental states better than existing models. |
Copied to clipboard
| Challenge: | Existing vision-language-action models are unsuitable for simulated or physical-world deployments . current methods fail when confronted with inherent real-world dynamic variability. |
| Approach: | They propose a test-time reinforcement learning framework that enables on-the-fly policy adaptation during inference. |
| Outcome: | Empirical results show that the proposed framework improves adaptability, stability and task success in dynamic, previously unseen scenarios. |
Copied to clipboard
| Challenge: | Existing studies on dialogue quality assessment are uncapable of providing an end-to-end and human-epistemic assessment dataset . open-domain dialogue assessment is complicated and costly, but it can be done by recruiting human evaluators. |
| Approach: | They propose a large-scale dialogue quality assessment dataset for automatically assessing open-domain dialogue quality. |
| Outcome: | The proposed dataset is openly accessible at https://github.com/yukunZhao/Dialogue_quality_evaluation. |
Copied to clipboard
| Challenge: | Low-rank adaptation methods for large language models have limitations in preserving world knowledge and limiting updates to preserve world knowledge. |
| Approach: | They propose a Fisher-optimized adaptive low Rank and Singular-VectorSelection framework for knowledge-preserving fine-tuning that allows efficient and task-sensitive updates. |
| Outcome: | The proposed framework outperforms existing methods for knowledge-preserving fine-tuning. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are often easily deceived by tricky questions such as “How many eyes does the sun have?” . |
| Approach: | They annotate a FalseQA dataset containing 2365 human-written FPQs and find that PLMs are capable of discriminating FPqs by fine-tuning on moderate numbers. |
| Outcome: | The proposed model can discriminate on FPQs by fine-tuning on moderate numbers of examples and generate reasonable explanations for false premise questions. |
Copied to clipboard
| Challenge: | Temporal reasoning is crucial for large language models to understand event concurrency and complex temporal interactions in natural language. |
| Approach: | They propose an ontology-guided and description logic–constrained temporal reasoning paradigm that shifts focus from internal inference to the explicit modeling of temporal structure. |
| Outcome: | The proposed method outperforms state-of-the-art methods by 2.07–31.83 F1 points and 1.00–30.73 EM points, exhibiting strong generalization and highlighting the potential of explicit temporal reasoning. |
Copied to clipboard
| Challenge: | Large language models have demonstrated exceptional performance across a wide range of tasks . however, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance and cost. |
| Approach: | They propose a multi-LLM routing framework that efficiently routes user queries to the most suitable LLM. |
| Outcome: | The proposed framework outperforms baseline methods in terms of effectiveness and interpretability. |
Copied to clipboard
| Challenge: | Existing large language models favor high-resource languages, such as English, at the expense of low-resourced and regional languages. |
| Approach: | They propose a series of language models that specifically focuses on Southeast Asian languages. |
| Outcome: | SeaLLM models outperform ChatGPT-3.5 in non-Latin languages by large margins . linguistic disparity impedes access to state-of-the-art AI technologies for non-English-speaking populations . |
Copied to clipboard
| Challenge: | Recent studies have demonstrated that many layers are functionally redundant in large language models (LLMs), enabling model compression by removing these layers to reduce inference cost. |
| Approach: | They propose a framework that removes redundant layers to reduce inference cost by preserving sensitivity-aware singular values. |
| Outcome: | The proposed framework outperforms existing methods in 90% of the original model under a 20% compression ratio. |
Copied to clipboard
| Challenge: | Existing methods for quantifying uncertainty in large language models with black-box API access are limited due to the complex data distributions and inner model mechanism. |
| Approach: | They propose a conformal prediction method that minimizes the size of prediction sets and ensures a statistical guarantee of the user-defined coverage. |
| Outcome: | The proposed method outperforms existing methods on close-ended and open-ended questions. |
Copied to clipboard
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can solve complex tasks through iterative information retrieval. |
| Approach: | They propose a turn-level stage-aware policy optimization approach to solve this problem . they introduce a first-occurrence latent reward mechanism to allocate partial rewards . |
| Outcome: | Experiments show that TSPO outperforms state-of-the-art models on Qwen2.5-3B and 7B models. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on basic abilities using nonverbal methods, such as yes-no and multiple-choice questions. |
| Approach: | They propose a benchmark that provides more nuanced evaluations of alignment capabilities for large Vision-Language Models (VLMs) they use a rule-calibrated evaluator that exceeds GPT-4's evaluation ability and a “alignment score” to assess the robustness and stability of models across diverse prompts. |
| Outcome: | The proposed benchmark covers 13 tasks across three categories and includes both single-turn and multi-turn dialogue scenarios. |
Copied to clipboard
| Challenge: | Existing methods for table entity linking ignore row and column contexts . existing methods for TEL focus on understanding sequential text contexts, making it difficult to adapt to the row and columns structure of tables. |
| Approach: | They propose to leverage row and column contexts to enhance the semantics of mentions in entity disambiguation. |
| Outcome: | The proposed method outperforms the state-of-the-art (SOTA) baseline by 1.5% on the in-domain dataset and 3.7% on average across three out-of domain datasets. |
Copied to clipboard
| Challenge: | Existing methods to update parametric knowledge of large language models (LLMs) are outdated and incontext editing (KE) is not effective due to the substantial cost associated with retraining. |
| Approach: | They propose a new decoding technique that enhances in-context editing (ICE) they propose to use parametric knowledge to update the models' knowledge . |
| Outcome: | The proposed technique improves ICE performance while incurring only half the latency. |
Copied to clipboard
| Challenge: | Existing approaches to attack pre-trained language models suffer from low success rates or fail to search efficiently in the exponentially large perturbation space. |
| Approach: | They propose an efficient framework to generate natural adversarial text by constructing different semantic perturbation functions. |
| Outcome: | The proposed framework generates natural adversarial texts for different languages with high success rates. |
Copied to clipboard
| Challenge: | Recent vision-language models (VLMs) have shown impressive capabilities as general visual assistants, but there are two challenges to their performance: (1) lacking task diversity in pretraining and visual instruction tuning; (2) annotation error and bias in GPT-4 synthesized instruction tuning data. |
| Approach: | They propose a two-stage instruction tuning framework that fine tunes VLMs firstly and further tuned on GPT-4 synthesized data. |
| Outcome: | The proposed framework outperforms the traditional single-stage visual instruction tuning framework and achieves state-of-the-art performance across a wide range of multi-modal evaluation benchmarks. |
Copied to clipboard
| Challenge: | Typical verifiers require resource-intensive supervised dataset construction, which is costly and faces limitations in data diversity. |
| Approach: | They propose an unsupervised verifier regularized by logical rules that uses internal activations and logical constraints on multiple reasoning paths. |
| Outcome: | Experiments on 10 datasets show that the proposed verifier outperforms baselines and is comparable to the supervised verifier. |
Copied to clipboard
| Challenge: | Existing methods to detect MGT from human-written texts are inadequate . existing methods are fine-tuned and zero-shot metric-based, but they can be more accurate. |
| Approach: | They propose a novel fine-tuned detector that can detect MGT from human-written texts by contrastive learning on selective perturbation. |
| Outcome: | The proposed method outperforms the state-of-the-art by 1.20% on four public datasets. |
Copied to clipboard
| Challenge: | MU has gained significant attention as a means to remove the influence of specific data from a trained model without requiring full retraining. |
| Approach: | They propose a novel approach that disentangles and selectively forgets both visual and textual associations, ensuring that unlearning does not compromise model performance. |
| Outcome: | Experiments on CIFAR-100, Flickr30K, and Conceptual 12M show that CLIPErase effectively removes designated associations from multimodal samples in downstream tasks while preserving model performance on retain set. |
Copied to clipboard
| Challenge: | In this paper, we introduce SCALE, a collaborative framework that connects a compact Specialized Translation Model (STM) and a general-purpose Large Language Model (LLM) as one unified translation engine. |
| Approach: | They propose a collaborative framework that connects a Specialized Translation Model (STM) and a general-purpose Large Language Model (LLM) as one unified translation engine. |
| Outcome: | The proposed framework outperforms both LLMs and supervised models in high-resource or challenging low-resourced settings. |
Copied to clipboard
| Challenge: | Explicit reasoning trajectories increase performance but often trigger overthinking . despite its importance, this study examines how each step of reasoning affects the final outcome . |
| Approach: | They propose a Reasoning Completion Point Detector that detects the RCP by monitoring rank dynamics of termination tokens. |
| Outcome: | The proposed method reduces token usage by up to 44% while preserving accuracy. |
Copied to clipboard
| Challenge: | Dynamical systems theory provides a framework for understanding iterative processes and evolution over time. |
| Approach: | They propose to apply this perspective to large language models which iteratively map input text to output text and re-express meaning with linguistic variation. |
| Outcome: | The proposed model reveals that paraphrases re-express meaning with linguistic variation limiting linguistic diversity . |
Copied to clipboard
| Challenge: | Video Large Language Models (VLMs) have been praised for their performance in coarse-grained video understanding but still face ineffective temporal grounding and inadequate timestamp representations. |
| Approach: | They propose a novel Video-LLM that senses and reasoned over specific video moments with fine-grained temporal precision. |
| Outcome: | The proposed model surpasses existing models in fine-grained video understanding tasks and exhibits strong potential as a general video understanding assistant. |
Copied to clipboard
| Challenge: | Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence. |
| Approach: | They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included . |
| Outcome: | The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model. |
Copied to clipboard
| Challenge: | RuleArena assesses the ability of large language models (LLMs) to follow complex, real-world rules in reasoning. |
| Approach: | They propose a benchmark to evaluate the ability of large language models (LLMs) to follow complex, real-world rules in reasoning. |
| Outcome: | The proposed benchmark covers airline baggage fees, NBA transactions, and tax regulations. |
Copied to clipboard
| Challenge: | Existing studies rely on idealized "gold" evidence for predictions, which is unrealistic due to its limited availability in real-world scenarios. |
| Approach: | They propose a fact-checking framework based on planning and customized action reasoning using LLMs. |
| Outcome: | The proposed framework outperforms baseline methods across three datasets and with varying complexity levels. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable linguistic comprehension and generation capability, but when applied to specialized industries, they face challenges such as hallucination, insufficient domain knowledge, and failing to incorporate the latest domain knowledge. |
| Approach: | They propose a paradigm for fine-tuning LLMs that effectively injects domain-specific knowledge while safeguarding inference-time data privacy. |
| Outcome: | The proposed model protects private data while enhancing the model's knowledge. |
Copied to clipboard
| Challenge: | Significant concerns emerge when addressing cultural sensitivity and local values. |
| Approach: | They propose a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. |
| Outcome: | The proposed model sets the state-of-the-art standard for open Arabic LLMs across various benchmarks. |
Copied to clipboard
| Challenge: | Existing RAG models are sensitive to the order in which evidence is presented, resulting in unstable performance and biased reasoning. |
| Approach: | They propose to quantify position bias in multimodal RAG systems by using position sensitivity index . they also develop a visualization framework to trace attention allocation patterns across decoder layers . |
| Outcome: | The proposed framework shows that multimodal interactions intensify position bias compared to unimodal settings and that this bias increases logarithmically with retrieval range. |
Copied to clipboard
| Challenge: | Recent years have witnessed a growing interest in the development of explainable recommendation models. |
| Approach: | They propose a model that combines prediction and generation tasks to produce more persuasive explanations by obtaining additional information from the training sets. |
| Outcome: | The proposed model outperforms state-of-the-art models on three datasets and shows that it is more persuasive than previous models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit potential artificial generic intelligence, however, their usage is costly with high response latency. |
| Approach: | They develop a dynamic contextual-bandit-based routing system for query-LLM assignment that leverages query tags to enhance query embeddings. |
| Outcome: | The proposed model maximizes response quality and minimizes cost and latency. |
Copied to clipboard
| Challenge: | Existing studies rely on deep graph neural networks (GNNs) to capture rich structural information, but they lack the structural information needed for QA. |
| Approach: | They propose a framework which captures structural information from KBs and models long-distance node relations from two perspectives. |
| Outcome: | The proposed framework models long-distance node relations from two perspectives . it is based on two widely used multi-hop KBQA datasets . |
Copied to clipboard
| Challenge: | a novel evaluation framework assesses the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. |
| Approach: | They propose to use a benchmark dataset to assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. |
| Outcome: | The proposed evaluation framework is based on a dataset of 1,267 test samples in 24 news domains. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on well-defined or abstract reasoning and fail to capture real-world engineering problems. |
| Approach: | They propose a hierarchical benchmark to evaluate large language models on engineering problems. |
| Outcome: | The proposed model performs well under well-defined conditions and is based on three levels of difficulty and covers diverse engineering subfields. |
Copied to clipboard
| Challenge: | Existing models for NLP evaluations lack the ability to generate informative critiques in pointwise grading and pairwise comparison especially without references. |
| Approach: | They propose a method which can acquire pointwise grading critiques with pseudo references and revise these critiques via multi-path prompting to obtain informative evaluation data in different tasks and settings. |
| Outcome: | The proposed method outperforms all open-source models and even GPT-4 in system-level correlations of pointwise grading. |
Copied to clipboard
| Challenge: | Existing methods to detect and correct spelling errors in Chinese take external input or just heuristic rules. |
| Approach: | They propose to incorporate phonological and visual similarity knowledge into Chinese language models by using a specialized graph convolutional network. |
| Outcome: | The proposed method outperforms existing models on three human-annotated datasets. |
Copied to clipboard
| Challenge: | Large language models lack reliability in scientific domains that require strict adherence to physical constraints. |
| Approach: | They propose a large-scale dataset constructed via a task-adaptive strategy and a hybrid verification protocol that combines deterministic solvers with semantic auditing to guarantee scientific rigor. |
| Outcome: | The proposed model outperforms baselines and general-purpose preference models and is competitive with proprietary models. |
Copied to clipboard
| Challenge: | Existing paradigms for pre-training and fine-tuning have limitations . knowledge rekindle aims to break through performance upper bounds of experts without introducing additional annotated data. |
| Approach: | They propose a new paradigm for pre-training and fine-tuning that aims to re-incorporate the fine- tuned expert model into the training cycle and break through performance upper bounds of experts. |
| Outcome: | The proposed model breaks through performance upper bounds of experts without additional annotated data. |
Copied to clipboard
| Challenge: | Imbalanced data distributions can cause models to overfit to majority classes and output unreliable (mostly overconfident) predictions. |
| Approach: | They propose to streamline the model development and deployment using focal loss to address imbalanced data distributions. |
| Outcome: | The proposed model training with focal loss improves calibration and accuracy compared to standard cross-entropy loss. |
Copied to clipboard
| Challenge: | Recent advances in Automatic Speech Recognition (ASR) have been fueled by massive speech corpora, but extending coverage to diverse languages with limited resources remains a formidable challenge. |
| Approach: | They propose a pipeline that converts large-scale text corpora into synthetic speech using off-the-shelf text-to-speech (TTS) models. |
| Outcome: | The proposed pipeline generates 500,000 hours of synthetic speech in ten languages and achieves transcription error reductions of over 30%. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed model improves naturalness and prosody diversity with clear margins. |
| Approach: | They propose a cross-utterance conditional VAE to estimate posterior probability distribution of latent prosody features for each phoneme by conditioning on acoustic features, speaker information, and text features from past and future sentences. |
| Outcome: | The proposed model improves naturalness and prosody diversity with clear margins. |
Copied to clipboard
| Challenge: | 3D visual grounding aims to localize the desired objects in a 3D point cloud by a free-form language description. |
| Approach: | They propose a relation-aware framework which captures relative spatial relationships between objects and enhances object attributes. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three benchmarks . it captures relative spatial relationships between objects and enhances object attributes . |
Copied to clipboard
| Challenge: | a recent study defines a conversation target from the system side to proactively steer conversations toward predefined targets or accomplish specific system-side goals. |
| Approach: | They propose a dataset curation framework that automatically curations a large-scale personalized dialogue dataset using a role-playing approach. |
| Outcome: | The proposed dataset is of high quality and could contribute to exploring personalized target-oriented dialogue. |
Copied to clipboard
| Challenge: | a library to facilitate the development, use, and evaluation of large language models (LLMs) is presented. |
| Approach: | They propose a unified library to facilitate the development, use and evaluation of large language models (LLMs). |
| Outcome: | The proposed library is based on extensive experiments in a variety of evaluation settings. |
Copied to clipboard
| Challenge: | Recent reinforcement learning approaches have advanced reasoning in Large Language Models (LLMs), yet their adaptation to multimodal LLMs remains underexplored. |
| Approach: | They propose a reinforcement learning framework that eliminates KL penalties and rewards consistency . they propose GRPO-CARE, which outperforms standard GR PO, with a base reward for accuracy and an adaptive bonus for consistency. |
| Outcome: | The proposed framework outperforms standard GRPO on the most difficult evaluation level and reasoning consistency test benchmarks. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) fail under RDEI, leading to disrupted structure and evidence-unsupported hallucinations. |
| Approach: | They propose a backbone-agnostic, evidence-driven pipeline that treats off-the-shelf MLLMs as interchangeable components to improve stem consistency and figure consistency. |
| Outcome: | The proposed pipeline improves stem consistency by 1.01-3.18%, figure consistency by 0.50-49.16%, and refusal F1 by 1.06-10.88% across question types. |
Copied to clipboard
| Challenge: | Autoregressive language models are trained exclusively left-to-right, yet they are limited in their ability to factorize text. |
| Approach: | They propose a purely reverse autoregressive language model that factorizes text as a product of left-to-right conditionals. |
| Outcome: | The proposed model can be used to score forward outputs using reverse posterior estimates. |
Copied to clipboard
| Challenge: | Contrastive Language-Image Pre-training (CLIP) is a standard for cross-modal image-text representation learning. |
| Approach: | They propose a framework that enhances pre-trained CLIP models by exploiting challenging text-image pairs within existing datasets. |
| Outcome: | The proposed framework improves CLIP models by exploiting text-image pairs in training. |
Copied to clipboard
| Challenge: | Uncertainty quantification (UQ) in natural language generation tasks remains an open challenge . however, black-box uncertainty measures require investigating with the proliferation of LLMs served via APIs. |
| Approach: | They propose a conformal uncertainty measure and a method to transform heuristic uncertainty notions into rigorous prediction sets. |
| Outcome: | Empirical results show that the proposed method outperforms state-of-the-art methods and can provide reliable guarantees for open-ended NLG tasks. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) scales capacity via conditional computation, but lacks knowledge lookup primitive. |
| Approach: | They propose a conditional memory instantiated via Deep Sparse Embedding (DSE) they propose 'u-shaped scaling law' that identifies optimal balance between MoE experts and DSE memory . |
| Outcome: | The proposed model outperforms an iso-parameter and isoFLOPs MoE baseline across knowledge and reasoning benchmarks and is infrastructure-efficient. |
Copied to clipboard
| Challenge: | Metaphor identification is a core task in metaphor processing, which involves recognizing and analyzing metaphorical expressions in text. |
| Approach: | They propose a new formulation of metaphor identification as a relation extraction problem . they use a dataset to analyze metaphorical relations between two spans, a target and a source . |
| Outcome: | The proposed model can capture the properties of the target and source in Chinese sentences. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have a tendency to generate factually incorrect or purely fictional responses, a phenomenon known as hallucination. |
| Approach: | They propose to use remote RAG to protect user query from privacy leakage . they introduce (n,)-DistanceDP to characterize privacy leakages of user query . |
| Outcome: | The proposed solution can resist embedding inversion attacks while achieving no loss in retrieval under various settings. |
Copied to clipboard
| Challenge: | Existing approaches to generalize commonsense reasoning lack instantiated knowledge and require pre-built concept taxonomies and annotations. |
| Approach: | They propose a framework that iteratively performs contextualized conceptualization and instantiation over commonsense knowledge bases by instructing large language models to generate both types of knowledge with critic filtering. |
| Outcome: | Empirical results show that distilling CANDLE on student models provides benefits across three downstream tasks. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated ChatGPT's remarkable few-shot, even zero-shot learning abilities when compared to other models. |
| Approach: | They quantitatively evaluate the performance of ChatGPT on inter-sentential relations such as temporal relations, causal relations, and discourse relations. |
| Outcome: | The proposed model performs well on temporal relations, causal relations, and discourse relations. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impressive capabilities but need for task-specific prompt engineering can hinder their generalization. |
| Approach: | They propose a lightweight and versatile retriever that automatically retrieves prompts for a given zero-shot task input. |
| Outcome: | The proposed model is universally applicable across tasks and models . it mitigates hallucination problem in chatGPT, and it improves even the strongest LLMs. |
Copied to clipboard
| Challenge: | Experimental results show RASAT can leverage a variety of relational structures while inheriting the pretrained parameters from the T5 model. |
| Approach: | They propose a Transformer seq2seq architecture augmented with relation-aware self-attention that leverages relational structures while inheriting pretrained parameters from the T5 model. |
| Outcome: | The proposed model can leverage relational structures while inheriting pretrained parameters from the T5 model effectively. |
Copied to clipboard
| Challenge: | Multiagent debate (MAD) is a popular approach for large language models . however, the performance of LLMs is suboptimal in complex reasoning scenarios . |
| Approach: | They propose a sequential collaboration framework to enable agents to provide constructive assistance to peers by decomposing complex tasks into essential subtasks. |
| Outcome: | The proposed framework achieves the highest performance on 19 out of 23 tasks and lower costs compared to MAD. |
Copied to clipboard
| Challenge: | Existing methods to improve embeddings from Mixture-of-Experts models allocate a fixed number of experts uniformly across all layers and tokens, ignoring inter-layer and inter-token heterogeneity. |
| Approach: | They propose an Adaptive Expert Allocation framework that performs layer-wise and token-wise expert allocation to enhance embedding quality. |
| Outcome: | The proposed method improves embedding quality across multiple MoE models. |
Copied to clipboard
| Challenge: | Recent deep generative models have already shown encouraging * Equal contribution. |
| Approach: | They propose to use generic instruction-tuned LLMs as direct text-to-sequence generators to achieve this goal. |
| Outcome: | Recent studies show that reflection improves sequence quality and alignment while maintaining competitive foldability. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a promising tool for OR, but they face challenges when dealing with complex problems. |
| Approach: | They propose a framework that augments existing datasets and generates high-quality fine-tuning data tailored to OR. |
| Outcome: | The proposed framework augments existing datasets and generates high-quality fine-tuning data . it prevents error propagation and ensures the quality of the generated dataset . |
Copied to clipboard
| Challenge: | Recent work has leveraged probing-based approaches to study the separability of malicious and benign inputs in Large Language Models’ internal representations. |
| Approach: | They propose to use probing-based methods to study separability of malicious and benign inputs in LLMs' internal representations to detect harmful and benign content. |
| Outcome: | The proposed methods show that they learn superficial patterns rather than semantic harmfulness. |
Copied to clipboard
| Challenge: | Existing surveys focus on LLMs' specific utility for data annotation and synthesis. |
| Approach: | They propose to use large language models to generate annotations from raw data . they also propose to review learning strategies for models utilizing LLM-generated annotations . |
| Outcome: | The proposed models can be used to improve the efficacy of machine learning models by generating and labeling raw data with relevant information. |
Copied to clipboard
| Challenge: | Large language models (LLMs) excel at natural language tasks but face deployment bottlenecks due to computational demands. |
| Approach: | They propose a training-free framework that exploits activation and weight sparsity . they use a three-tier routing strategy that uses magnitude-based pruning . |
| Outcome: | Experiments on Llama and Mistral models show that DAWS outperforms activation-weight sparsity pruning methods. |
Copied to clipboard
| Challenge: | Existing methods for predicting future facts from time-evolving graphs rely on statistical co-occurrences and extensive path enumeration. |
| Approach: | They propose a Critic-Guided Rule Induction method which treats temporal rules as rule hypotheses to be examined and adopts a decoupled Generation-Discrimination pipeline to induce rules that are high-coverage and high-precision. |
| Outcome: | The proposed method outperforms strong baselines on three benchmarks and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing knowledge graphs focus on connecting intentions but lacks the ability to model the relationships between different intentions. |
| Approach: | They propose a framework to automatically generate an intention knowledge graph, capturing connections between user intentions. |
| Outcome: | The proposed model outperforms state-of-the-art methods and shows its utility. |
Copied to clipboard
| Challenge: | Current compression techniques entail structural pruning and a recovery phase that leverages the Low-Rank Adaptation algorithm. |
| Approach: | They propose a hierarchical rank allocation method that enables efficient fine-tuning of pruned LLMs according to layerwise specific recovery requirements. |
| Outcome: | The proposed algorithm outperforms state-of-the-art methods across pruning settings and LLM architectures with improvements ranging from 0.7% to 5.5%. |
Copied to clipboard
| Challenge: | Existing methods for learning textual network embeddings are noisy and sparse. |
| Approach: | They propose to use text-based attention parsing to learn context-aware network embeddings. |
| Outcome: | The proposed model outperforms state-of-the-art methods in a number of domains. |
Copied to clipboard
| Challenge: | Existing methods to improve knowledge base are incomplete and difficult to understand. |
| Approach: | They propose a novel QA method by leveraging text information to enhance the incomplete KB. |
| Outcome: | Extensive experiments on the WebQuestionsSP benchmark prove the effectiveness of the proposed model. |
Copied to clipboard
| Challenge: | Traditional contrastive learning methods treat negative feedback as equally hard or easy, ignoring informative semantic difficulty during training. |
| Approach: | They propose a framework leveraging Large Language Models to Activate interactions in Graph Contrastive Learning for Recommendation. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks on multiple benchmarks. |
Copied to clipboard
| Challenge: | Empirical results show that AFT-trained models achieve substantial gains with test-time scaling. |
| Approach: | They introduce a supervised fine-tuning paradigm where models synthesize multiple draft responses into a single, refined answer. |
| Outcome: | Empirical results show that AFT-trained models outperform baseline models while eliminating external guidance. |
Copied to clipboard
| Challenge: | Current speaker diarization systems consider only acoustic information, resulting in performance degradation when encountering adverse acustic environment. |
| Approach: | They propose methods to extract speaker-related information from conversational semantics in multi-party meetings. |
| Outcome: | The proposed method improves on AISHELL-4 and AliMeeting datasets on speakers diarization and speaker-turn detection. |
Copied to clipboard
| Challenge: | Current alignment approaches rely on refusal alignment to avoid harmful content . large language models are often overly cautious or overlook subtle harmful content. |
| Approach: | They propose a framework for fine-grained safe generation in Large Language Models that enables real-time, token-level harmfulness detection and redaction without loss in capability. |
| Outcome: | The proposed framework achieves over 90% in F1 score for detecting and redacting harmful content while preserving overall utility and informativeness of the model’s responses. |
Copied to clipboard
| Challenge: | Existing generative models lack the capacity for explicit and controllable reasoning, a key advantage of LLMs. |
| Approach: | They propose a framework that integrates dialogue, reasoning, and personalized recommendation. |
| Outcome: | Experiments across public benchmarks show state-of-the-art performance. |
Copied to clipboard
| Challenge: | Excessive exploitation can cause the model to become overconfident in its suboptimal solutions, thereby limiting its capabilities to explore novel reasoning strategies. |
| Approach: | They propose a method that dynamically down-weights extreme token-level updates via a Gaussian kernel and reduces the instability caused by the trade-off. |
| Outcome: | The proposed method improves downstream performance across reasoning benchmarks and stabilizes entropy as training progresses. |
Copied to clipboard
| Challenge: | Incorporating external context can enhance the response quality of Large Language Models (LLMs). however, real-world contexts often mix relevant information with disproportionate inappropriate content. |
| Approach: | They propose a Poisoned Context Testbed to pair queries with real-world contexts . they propose 'rw-Steering' to internalize inappropriate signals . |
| Outcome: | The proposed model improves response quality by 39.8% and reverses undesirable behavior curve. |
Copied to clipboard
| Challenge: | Post-trained LLMs typically compromise reliability with severe overconfidence, resulting in inaccurate responses. |
| Approach: | They propose a solution that feeds PoLLMs into the base LLM to get confidence. |
| Outcome: | The proposed solution reduces expected calibration error (ECE) by 42.90% compared to the best unsupervised baselines. |
Copied to clipboard
| Challenge: | Existing studies focus on causality existence, but ignore causal direction. |
| Approach: | They propose a new *identifying while learning* mode for the ECI task that takes care of the causal direction and updates events’ representations for boosting next round of causality identification. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on two public datasets. |
Copied to clipboard
| Challenge: | Existing data-centric paradigms equate quality with factuality or diversity and ignore the internal logical complexity of training samples. |
| Approach: | They propose a density-aware re-cognizing optimization strategy that prioritizes high-density logical samples to align training with the model's reasoning boundary. |
| Outcome: | The proposed metric outperforms existing methods and improves reasoning performance without increasing total data volume. |
Copied to clipboard
| Challenge: | Linux kernel device drivers are tightly coupled with hardware, making them difficult to execute and test without physical devices. |
| Approach: | They present a tool that generates QEMU-based virtual devices directly from Linux driver source code. |
| Outcome: | The proposed tool generates QEMU-based virtual devices directly from Linux driver source code. |
Copied to clipboard
| Challenge: | Existing approaches to optimize sequential recommendation systems rely on item ID sequences, but they lack collaborative knowledge and limited controllability. |
| Approach: | They propose a simple bi-tuning framework with collaborative information for controllable Large Language Model-based Sequential Recommendation (Laser) they incorporate learnable virtual tokens at prefix and suffix of input text to adapt LLMs with collaborative knowledge . |
| Outcome: | The proposed framework outperforms state-of-the-art recommendations on real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods to adapt Pre-trained Language Models to downstream tasks are limited by their inference APIs. |
| Approach: | They propose a multi-prompting decoding framework that query PLMs with multiple prompts . they propose to query Plms with optimal transport for hidden states and calibrated decoding for class scores . |
| Outcome: | The proposed framework achieves state-of-the-art results on multiple natural language understanding datasets under the few-shot setting. |
Copied to clipboard
| Challenge: | Current methods for instruction generation depend on privileged inputs such as semantic maps, landmark annotations, and panoramic views. |
| Approach: | They propose a task that generates coherent navigation instructions from egocentric visual observations. |
| Outcome: | The proposed task generates coherent navigation instructions from egocentric visual data . the proposed task improves performance over state-of-the-art methods in BLEU-4 and CIDEr scores . |
Copied to clipboard
| Challenge: | Existing question reformulation models are based on supervised question labels without considering feedback information from answers. |
| Approach: | They propose a question reformulation model that integrates conversational history information with reinforcement learning. |
| Outcome: | The proposed model is more effective in conversational machine comprehension with reinforcement learning. |
Copied to clipboard
| Challenge: | Existing abstractive summarization systems generate incorrect facts with respect to the source text. |
| Approach: | They propose a suite of two factual correction models that leverages question-answering knowledge to make corrections in system-generated summaries via span selection. |
| Outcome: | The proposed model improves factuality of news summarization without sacrificing summary quality. |
Copied to clipboard
| Challenge: | Existing secure code generation methods have limited generalizability to unseen test cases and poor robustness against the attacked model, leading to safety failures in code generation. |
| Approach: | They propose a generalizable and robust secure code generation method SecCoder by using in-context learning and the safe demonstration. |
| Outcome: | The proposed method achieves a significant security improvement of 7.20% on unseen test cases and better robustness against the attacked model. |
Copied to clipboard
| Challenge: | Existing methods for training contrastive learning based sentence embedding models are largely influenced by the quality of sentence pairs. |
| Approach: | They propose a framework that decomposes LLMs into three stages for training . they propose to refine the generated content at these stages to ensure only high-quality sentence pairs are utilized to train a base contrastive learning model. |
| Outcome: | The proposed framework surpasses ChatGPT and ChatGPP in terms of performance. |
Copied to clipboard
| Challenge: | Existing studies focus on singing voice synthesis and music generation independently. |
| Approach: | They propose a novel task called Text-to-Song synthesis which incorporates both vocal and accompaniment generation. |
| Outcome: | The proposed method can synthesize songs with comparable quality and style consistency. |
Copied to clipboard
| Challenge: | Recent advances in Chain-of-Thought prompting have facilitated significant breakthroughs for Large Language Models (LLMs) in complex reasoning tasks. |
| Approach: | They propose a hierarchical reasoning aggregation framework to address this problem . they propose dynamic sampling to adjust the number of reasoning chains . |
| Outcome: | The proposed framework outperforms existing ensemble methods on complex reasoning tasks. |
Copied to clipboard
| Challenge: | Generative AI has demonstrated unprecedented creativity in the field of computer vision, yet such phenomena have not been observed in the realm of literary creation. |
| Approach: | They propose a framework for unleashing the creativity of large language models (LLMs) they assign LLMs to different roles involved in real-world scenario, they write . |
| Outcome: | The proposed framework outperforms baselines in terms of coherence, relevance, interestingness and overall quality on automatically generated screenplays. |
Copied to clipboard
| Challenge: | a new framework for training data processing for Chinese medical language models is proposed . experimental results show that the framework significantly improves model accuracy . |
| Approach: | They propose a data processing framework for Chinese medical language models training and deployment . the framework is based on a question-oriented model training strategy and privacy preservation . |
| Outcome: | The proposed framework significantly improves model accuracy and reduces privacy leakage by 27%. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown outstanding breakthroughs in code generation. |
| Approach: | They propose a case-to-code induction task that exploits the expressiveness and correctness of programs by incorporating LLMs into their training. |
| Outcome: | The proposed task improves distribution case-to-code induction and various coding generation tasks. |
Copied to clipboard
| Challenge: | Existing approaches lack flexibility to address diverse and ever-evolving user queries in open domains. |
| Approach: | They propose to evaluate LLMs on open-domain knowledge that requires tools to solve diverse and ever-evolving user queries. |
| Outcome: | The proposed system outperforms baselines in the open domain task-solving benchmark. |
Copied to clipboard
| Challenge: | Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). |
| Approach: | They present AraMUS, the largest Arabic PLM with 11B parameters trained on 529GB of high-quality Arabic textual data. |
| Outcome: | The proposed model achieves state-of-the-art performance on a diverse set of Arabic classification and generative tasks. |
Copied to clipboard
| Challenge: | Existing pre-trained dialog models shed light on various downstream tasks in natural language processing (NLP). |
| Approach: | They propose a dialog pre-training framework that introduces latent variables into the enhanced encoder-decoder pre-train framework to increase relevance and diversity of responses. |
| Outcome: | The proposed model achieves state-of-the-art on personaChat, DailyDialog, and DSTC7-AVSD datasets. |
Copied to clipboard
| Challenge: | Recent work attributes performance degradation to an exponential decay in hidden-state memory. |
| Approach: | They propose a token filtering strategy that is training-free and attention-guided . they propose 'LAMB' to preserve critical tokens during inference . |
| Outcome: | The proposed token filtering improves long-context performance by 30.35% over state-of-the-art methods on benchmarks. |
Copied to clipboard
| Challenge: | Existing retrievers suffer from temporal-semantic misalignment and outdated-document interference . Existing frameworks suffer from both temporal validity and outdated factual versions . |
| Approach: | They propose a framework that mitigates temporal hallucinations by embedding heterogeneous temporal signals into the semantic space to ensure retrieval fidelity. |
| Outcome: | Experiments show that Re3 outperforms baselines by 9.7% in generation accuracy . the framework outperformed strongest baselines on challenging dynamic tasks . |
Copied to clipboard
| Challenge: | Early approaches focus on text-based reasoning, but they often follow a single task-specific reasoning pattern. |
| Approach: | They propose a generative multimodal reasoning paradigm that unifies diverse reasoning skills by generating intermediate images during the reasoning process. |
| Outcome: | The proposed model unifies diverse multimodal reasoning skills by generating intermediate images during the reasoning process. |
Copied to clipboard
| Challenge: | Existing methods for few-shot learning are insufficient to capture task variations in natural language domains. |
| Approach: | They propose an adaptive metric learning approach that automatically determines the best weighted combination from a set of metrics obtained from meta-training tasks for a newly seen few-shot task. |
| Outcome: | The proposed method performs favorably against state-of-the-art few shot learning algorithms on real-world sentiment analysis and dialog intent classification datasets. |
Copied to clipboard
| Challenge: | Existing knowledge graphs focus on the representation and reasoning of general factual knowledge, while there are significant deficiencies in the understanding and reasoning for emotional knowledge. |
| Approach: | They propose a commonsense knowledge graph that can be used to represent emotional knowledge by combining theories from psychology, cognitive science, and linguistics. |
| Outcome: | The proposed model surpasses GPT-4-Turbo in the emotion-related tasks. |
Copied to clipboard
| Challenge: | Existing methods for generating paraphrases with linguistic knowledge are often domain specific and hard to scale, or yield inferior results. |
| Approach: | They propose an end-to-end conditional generative architecture for generating paraphrases via adversarial training which does not depend on extra linguistic information. |
| Outcome: | The proposed method outperforms existing models on automatic metrics and human evaluations on four public datasets. |
Copied to clipboard
| Challenge: | Using large language models, large multimodal models struggle with basic tasks like reading time from a clock and planning a route using a road map. |
| Approach: | They propose a multimodal self-instruct that synthesizes massive abstract images and visual reasoning instructions. |
| Outcome: | The proposed model synthesizes 11,193 abstract images and reasoning instructions across eight visual scenarios. |
Copied to clipboard
| Challenge: | Existing methods to update deployed models are prone to overfit . however, non-parametric methods are liable to over-fit the retrieved examples . |
| Approach: | They propose to learn Kernel-Smoothed Translation with Example Retrieval (KSTER) this approach allows users to adapt models to emerging cases without retraining . |
| Outcome: | The proposed approach achieves 1.1 to 1.5 BLEU scores over existing methods without retraining . the proposed model is released on https://github.com/jiangqn/KSTER. |
Copied to clipboard
| Challenge: | Existing integrations of large language models and large multimodal models are limited . Existing platforms for developing embodied agents are limited and limited based on open-source software. |
| Approach: | They propose an open platform for developing embodied agents using LLMs and LMMs. |
| Outcome: | The proposed platform surpasses GPT-4V in embodied tasks with its model training on LEGENT data. |