Papers by Wei Zhao
Copied to clipboard
| Challenge: | Existing knowledge base population systems require a machine translation task to generate multiple facts, but the fact order is not considered. |
| Approach: | They propose a knowledge base population task that aims to discover facts about entities from texts and expand a KB with these facts. |
| Outcome: | The proposed networks achieve state-of-the-art (SoTA) performance on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing data augmentation methods for event extraction are costly and time-consuming. |
| Approach: | They propose a data augmentation framework that randomly masks out an adjunct sentence fragment and infills a variable-length text span with a fine-tuned infilling model. |
| Outcome: | The proposed framework can generate more diverse data while keeping the original structure unchanged . it can replace a fragment of arbitrary length in the text with another fragment of variable length . |
Copied to clipboard
| Challenge: | Existing studies focus on language-agnostic settings, neglecting the inherently multilingual nature of modern software development. |
| Approach: | They propose a proportion-dependent scaling law that prioritizes high-utility languages . they propose PLs to have varying effects during pre-training that affect model performance . |
| Outcome: | The proposed scaling law is based on 1000+ experiments across multiple languages and models. |
Copied to clipboard
| Challenge: | Personalized Large Language Models (PLLMs) aim to align outputs with individual user preferences . current methods of fine-tuning a separate module for each user are unscalable . |
| Approach: | They propose a Merge-then-Adapt framework for Personalized Large Language Models . they construct a shared Meta-LoRA bank and propose an Adaptive LoRA Fusion stage . |
| Outcome: | The proposed framework outperforms existing SOTA methods on the LaMP benchmark. |
Copied to clipboard
| Challenge: | MELLE is a novel language modeling approach for text-to-speech synthesis that generates continuous tokens from text . authors demonstrate that it reduces the need for vector quantization and improves model robustness . |
| Approach: | They propose to autoregressively generate continuous mel-spectrogram frames directly from text condition, bypassing vector quantization. |
| Outcome: | The proposed model achieves superior performance across multiple metrics and is more streamlined. |
Copied to clipboard
| Challenge: | Despite significant efforts in safety alignment, large language models (LLMs) such as GPT-4 and LLaMA 3 remain vulnerable to jailbreak attacks that can induce harmful behaviors. |
| Approach: | They propose a feature extraction method to extract sample-agnostic features from benign datasets in the form of adversarial suffixes and propose 'suffix maybe features' they show that adversarials generated from jailbreak attacks may contain meaningful features, i.e. appending the same suffix to different prompts results in responses exhibiting specific characteristics. |
| Outcome: | The proposed method extracts sample-agnostic features from benign datasets and shows that they may contain meaningful features. |
Copied to clipboard
| Challenge: | Topic segmentation aims to split automatic speech recognition transcriptions into segments that are bounded by thematic meanings. |
| Approach: | They propose a Maximum-Minimum Mutual information paradigm for linear topic segmentation without using any parallel data. |
| Outcome: | The proposed paradigm outperforms the state-of-the-art methods by a significant margin. |
Copied to clipboard
| Challenge: | Existing methods for zero-shot relation extraction lack explicit modeling of matching pattern . et al. (2018) show that our method achieves higher matching accuracy and faster inference speed . |
| Approach: | They propose a fine-grained semantic matching method tailored for zero-shot relation extraction . they decompose sentence-level similarity score into entity matching score and context matching score . |
| Outcome: | The proposed method achieves higher matching accuracy and faster inference speed than state-of-the-art methods. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly integrated into our daily lives, raising ethical concerns, especially about perpetuating stereotypes. |
| Approach: | They propose a method that incorporates a neutral word semantics-based loss function to alleviate the deterioration of the LMS during debiasing. |
| Outcome: | The proposed method alleviates the deterioration of the Language Modeling Score (LMS) by incorporating a neutral word semantics-based loss function. |
Copied to clipboard
| Challenge: | a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities . |
| Approach: | They present a comparative analysis to identify and distinguish LLM activities from human activities. |
| Outcome: | The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities. |
Copied to clipboard
| Challenge: | Lowrank adaptation and its variants introduce significant latency in multi-tenant settings, hindering their applications in the industry. |
| Approach: | They propose a framework to fine-tune LoRA modules on a large-scale instruction tuning dataset. |
| Outcome: | The proposed framework outperforms existing PEFT methods and significantly reduces inference latency. |
Copied to clipboard
| Challenge: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Approach: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
| Outcome: | This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process. |
Copied to clipboard
| Challenge: | Existing methods for analyzing and utilizing toxic samples are limited . current methods fail to fully harness their potential . |
| Approach: | They propose a diverse detoxification framework that leverages toxic samples' diversity . they propose MPSG strategy and SC-DPO approach to elicit personalized toxic responses . |
| Outcome: | The proposed framework could be used to optimize large language models for user safety . it incorporates two components: MPSG strategy and SC-DPO approach . |
Copied to clipboard
| Challenge: | Existing approaches to simultaneous machine translation require a robust read/write policy . a standalone multi-path wait-k model performs competitively with adaptive policies . |
| Approach: | They propose a more flexible approach by decoupling the adaptive policy model from the translation model. |
| Outcome: | The proposed approach outperforms baseline approaches in translation tasks. |
Copied to clipboard
| Challenge: | Existing studies on cognitive distortion have limited generalizability and performance of models in large-scale and cross-linguistic contexts. |
| Approach: | They propose a multi-task learning model based on teacher student architecture solution which improves generalization performance. |
| Outcome: | The proposed model improves generalizability and interpretability of the proposed model. |
Copied to clipboard
| Challenge: | Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos. |
| Approach: | They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues. |
| Outcome: | The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor . |
Copied to clipboard
| Challenge: | Parameter-Efficient Fine-Tuning (PEFT) methods have gained popularity for adapting pre-trained Large Language Models (LLMs) to downstream tasks. |
| Approach: | They propose a method to optimize the importance of full layers with layer-wise importance scoring by leveraging the estimated importance scores. |
| Outcome: | The proposed method is compatible with PEFT methods that operate on a per-layer basis and achieves better performance. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown nearly saturated performance on many NLP tasks. |
| Approach: | They construct multiple sensitive factors time QA which encompasses three temporal factors . they test current mainstream LLMs with different parameter sizes . |
| Outcome: | The proposed model incorporates three temporal factors with 2,853 samples . the results show that LLMs fall behind smaller models on these factors . |
Copied to clipboard
| Challenge: | Existing studies on hallucination focus on text or vision, while few audio-oriented studies are limited in scale, modality coverage, and diagnostic depth. |
| Approach: | They propose a large-scale benchmark for evaluating hallucinations across speech, sound, and music. |
| Outcome: | The proposed model improves hallucination rate, yes/no bias, error-type analysis, and refusal rate. |
Copied to clipboard
| Challenge: | DiscoScore is a parametrized discourse metric that uses BERT to model discourse coherence . it is weak when operated at system level, and is therefore not reliable in a way to spot improvements . |
| Approach: | They propose a parametrized discourse metric which uses BERT to model discourse coherence from different perspectives. |
| Outcome: | The proposed model outperforms existing models on document-level machine translation and summarization. |
Copied to clipboard
| Challenge: | Existing work describes paragraph-level counter-argument generation task as paragraph-based . however, sentence-level generation can be quite different due to its unique constraints and brevity-focused challenges. |
| Approach: | They propose a benchmark framework for sentence-level counter-argument generation . they use an annotated debate forum dataset to generate high-quality counter-argments . |
| Outcome: | The proposed framework and evaluator are competitive in counter-argument generation tasks. |
Copied to clipboard
| Challenge: | Existing literature observes bias in question answering (QA) models, but there is no method to mitigate it. |
| Approach: | They propose an approach to mitigate the bias of question answering models by observing the influence of a query instance on another instance. |
| Outcome: | The proposed method reduces bias level in all 9 bias categories while maintaining comparable QA accuracy. |
Copied to clipboard
| Challenge: | Existing studies focus on specialized agents designed for particular tasks. |
| Approach: | They propose to scale annotated interaction trajectories and fine-tune LLMs on AgentBank to get a series of agent models, Samoyed. |
| Outcome: | The proposed model can scale to get generalized agent capabilities. |
Copied to clipboard
| Challenge: | Existing approaches for improving the correctness of LLMs include prompt engineering, retrieval methods, and a generative model. |
| Approach: | They propose a method that generates a risk score to estimate the probability of error in an LLM by integrating multiple sources of information. |
| Outcome: | The proposed method is well correlated with the true LLM error rate, thus facilitating error correction. |
Copied to clipboard
| Challenge: | Existing studies on graph learning on text-attributed graphs have been limited by memory cost and underutilization of relationships between nodes and words. |
| Approach: | They propose a Node Representation Update Pre-training Architecture based on Co-modeling text and graph to learn representations of papers and words simultaneously. |
| Outcome: | The proposed model outperforms baselines on the ogbn-arxiv benchmark dataset. |
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is essential for dialogue systems to identify the emotions expressed by speakers. |
| Approach: | They propose a method that incorporates both belief and desire to accurately identify emotions by extracting emotion-eliciting events from utterances and construct graphs that represent beliefs and desires in conversations. |
| Outcome: | The proposed model outperforms existing models on four popular ERC datasets and validates its performance with multiple state-of-the-art models. |
Copied to clipboard
| Challenge: | Reasoning-Intensive Retrieval (RIR) targets retrieval settings where relevance is mediated by latent inferential links between a query and supporting evidence, rather than semantic similarity. |
| Approach: | They propose a taxonomy that categorizes methods based on where and how reasoning is integrated into the retrieval pipeline. |
| Outcome: | The proposed method framework provides a detailed analysis of the current landscape and its trade-offs and practical applications. |
Copied to clipboard
| Challenge: | Currently, vision-language models excel in many downstream tasks but struggle with spatial reasoning, which is crucial for navigation and interaction with physical environments. |
| Approach: | They propose a framework that generates synthetic data to provide targeted supervision for VLMs across these basic spatial capabilities. |
| Outcome: | The proposed framework disentangles 2D spatial reasoning into three core components: direction comprehension, distance estimation, and localization. |
Copied to clipboard
| Challenge: | Existing frameworks for detecting fake news videos are limited . a new approach is proposed to integrate neighborhood information of new videos . |
| Approach: | They propose a framework for automatically detecting fake news videos . it integrates neighborhood relationship of new videos belonging to same event . |
| Outcome: | The proposed framework improves performance of existing detectors and graph aggregation and debunking rectification modules. |
Copied to clipboard
| Challenge: | Evaluation in NLP is usually done by comparing the scores of competing systems . averaging scores independently and declaring the best system is difficult . |
| Approach: | They examine the use of averages to aggregate evaluation scores into a final number . they argue that the average ignores the pairing arising from the fact that systems are evaluated on the same test instances. |
| Outcome: | The proposed method ignores the pairing arising from the fact that systems are evaluated on the same test instances. |
Copied to clipboard
| Challenge: | Existing approaches focus on improving the quality of correct training data, neglecting the value contained in error data, thereby hindering the model’s reflective ability. |
| Approach: | They propose to improve LLM's reasoning ability by learning from error data and a grounded mistake augmentation method to collect representative errors. |
| Outcome: | The proposed model achieves significant performance improvements over other strong models with less than 90k data. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate outputs that stray from user input or contravene established knowledge. |
| Approach: | They propose a new phenomenon, Authority Bias, where LLMs favor one knowledge source over the other . they propose atomic information that generates conflicts and a Conflict Detection Enhanced Query framework . |
| Outcome: | The proposed framework reduces Authority bias in large language models . it detects conflicts, performs credibility assessment on conflicting paragraphs, and detects perturbed text . |
Copied to clipboard
| Challenge: | Existing methods for generating open-ended rubrics suffer from scalability bottlenecks and coarse criteria resulting in a supervision ceiling effect. |
| Approach: | They propose a framework for automated Coarse-to-Fine Rubric Generation . their framework uses principle-guided synthesis, multi-model aggregation, difficulty evolution . |
| Outcome: | The proposed framework produces comprehensive and highly discriminative criteria capable of capturing the subtle nuances. |
Copied to clipboard
| Challenge: | Currently, retrieval-based dialogues are performed in shallow ways . a recent study investigated the problem of context-response matching in open-domain . |
| Approach: | They propose a model that lets utterance-response interaction go deep by stacking interaction blocks. |
| Outcome: | The proposed model outperforms state-of-the-art methods on three benchmark data sets. |
Copied to clipboard
| Challenge: | Multi-agent LLMs are rapidly moving from prototype to real-world use . network topology is a first-order security parameter in multi-aggent systems . |
| Approach: | They propose a framework for comparing topology-conditioned memory leakage in multi-agent LLM systems. |
| Outcome: | The proposed framework evaluates topology-conditioned memory leakage in multi-agent LLM systems. |
Copied to clipboard
| Challenge: | Large language models (LLMs) demonstrate impressive multilingual capability, but their performance varies substantially across different languages. |
| Approach: | They propose a generic template prompt that stimulates cross-lingual and logical reasoning skills to enhance task performance across languages. |
| Outcome: | The proposed method improves multilingual capability across languages and covers high-resource and low-resourced languages. |
Copied to clipboard
| Challenge: | Image-to-text tasks such as captioning and controllable image descriptions have received extensive attention for decades. |
| Approach: | They propose a new perspective for image-to-text to generate spatial descriptions by combining two objects in an image. |
| Outcome: | The proposed model is awe-inspiring and human-like, and the proposed end-to-end architecture is the better choice for their integration. |
Copied to clipboard
| Challenge: | Document images are characterized by higher resolutions, denser content, and more complex structural layouts. |
| Approach: | They propose a 1.2B-parameter document parsing vision-language model that decouples layout analysis from local content recognition. |
| Outcome: | The proposed model surpasses general-purpose and domain-specific models on multiple benchmarks while maintaining significantly lower computational overhead. |
Copied to clipboard
| Challenge: | Language model hallucinations and limited availability of labeled datasets often result in misaligned formulations, code errors and feasibility failures. |
| Approach: | They propose a Monte Carlo Tree Search framework that automates optimization problems from natural language descriptions with efficiency and reliability. |
| Outcome: | The proposed framework achieves state-of-the-art solution accuracy and reduces token usage. |
Copied to clipboard
| Challenge: | Existing works store a small number of typical samples to re-train the model for alleviating forgetting. |
| Approach: | They propose a continual relation extraction model that uses memory-insensitive relation prototypes and memory augmentation to overcome the overfitting problem. |
| Outcome: | The proposed model outperforms existing models on analogous relations and overcomes overfitting problem. |
Copied to clipboard
| Challenge: | Existing methods for distantly supervised relation extraction suffer from low quality of test set, which leads to considerable biased performance evaluation. |
| Approach: | They propose a method to evaluate distantly supervised relation extraction using noisy test sets and manual annotations. |
| Outcome: | Experiments on a widely used benchmark show that the proposed method can yield approximately unbiased evaluations for distantly supervised relation extractors. |
Copied to clipboard
| Challenge: | Existing methods to attack pre-trained language models rely on manual selection of triggers and backdoor representations. |
| Approach: | They propose a backdoor attack method that turns manual selection into automatic optimization . they propose to use poisoned contrastive learning to learn more uniform backdoor representations . |
| Outcome: | The proposed method achieves better attack performance on text classification tasks compared to manual methods. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning. |
| Approach: | They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations. |
| Outcome: | The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion. |
Copied to clipboard
| Challenge: | Low-rank adaptation and its mixture-of-experts (MOE) methods are highly effective but introduce significant latency in multi-tenant settings due to the LoRA modules and MOE routers added to multiple linear modules. |
| Approach: | They propose a low-rank adaptation variant that considers each LoRA module as an expert and employs a prompt-aware routing mechanism. |
| Outcome: | Extensive analysis on commonsense reasoning tasks and math reasoning tasks show that MiLoRA outperforms strong PEFT baselines with comparable tunable parameter budgets. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens. |
| Approach: | They propose a training-free pruning framework that prunes multimodal tokens without a trained pruning method. |
| Outcome: | The proposed pruning framework outperforms existing token pruning methods and generalizes across diverse MLLMs. |
Copied to clipboard
| Challenge: | Existing approaches to generalize deep neural networks are datahungry and generalize poorly from small datasets. |
| Approach: | They propose an agreement score to evaluate routing processes at instance-level and an adaptive optimizer to enhance routing. |
| Outcome: | The proposed approach improves on two NLP tasks and in low-resource settings with few training instances. |
Copied to clipboard
| Challenge: | a recent study shows that sparse activation techniques can reduce inference performance without sacrificing performance. |
| Approach: | They propose to sparsify a pre-trained dense large language model into a mixture-of-experts architecture for faster inference. |
| Outcome: | The proposed approach is more efficient than one-shot sparsification techniques . it achieves 97% performance retention on downstream tasks with only 50% of parameters activated . |
Copied to clipboard
| Challenge: | Existing controllable dialogue generation models focus on single attribute and lack generalization capability to out-of-distribution multiple attribute combinations. |
| Approach: | They propose a compositional generalization model that learns from seen attributes and generalizes to unseen combinations. |
| Outcome: | The proposed model can learn from seen attribute values and generalize to unseen combinations. |
Copied to clipboard
| Challenge: | Existing studies evaluate the efficiency of LLM-based rerankers using proxy metrics such as latency and the number of forward passes. |
| Approach: | They propose to use a large language model to evaluate the efficiency of LLM-based rerankers . they propose to measure ranking quality and query processing efficiency using an interpretable FLOPs estimator . |
| Outcome: | The proposed metrics evaluate LLM-based rerankers with different architectures without running any experiments. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. |
| Approach: | They propose to use a model averaging technique to find the most powerful alignment-forging Pareto front among RLHF algorithms. |
| Outcome: | The proposed method achieves the strongest alignment-forging Pareto front among competing methods. |
Copied to clipboard
| Challenge: | Text-Centric Visual Question Answering (TEC-VQA) is a text-centric visual task understanding tool. |
| Approach: | They introduce a benchmark that features human expert annotations across 9 languages . they prioritize the text in question-answer pairs while disregarding visual text in images . |
| Outcome: | The proposed benchmarks prioritize the text in question-answer pairs while disregarding visual text in images. |
Copied to clipboard
| Challenge: | Existing methods for aligning LVLMs rely on external datasets, human annotations or complex post-processing. |
| Approach: | They propose a method that generates a debiased self-judgment score for LVLMs . this self-evaluation metric is created internally by the model without external resources . |
| Outcome: | The proposed approach outperforms existing methods in reducing hallucinations and safety concerns. |
Copied to clipboard
| Challenge: | Existing knowledge injection methods fail to understand the semantics of tweets . |
| Approach: | They propose a method to flexibly inject knowledge into a pre-trained language model and adaptively expand tweets context. |
| Outcome: | The proposed method is based on two training stages to flexibly inject knowledge into the pre-trained language model and adaptively expand tweets context. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for literature prioritize mechanical accuracy over artistic expression . this bias could result in an irreversible decline in translation quality and cultural authenticity . |
| Approach: | They propose a novel, reference-free, LLM-based question-answering framework for literary translation evaluation. |
| Outcome: | a novel, reference-free, LLM-based question-answering framework is developed for literary translation evaluation. |
Copied to clipboard
| Challenge: | Existing prompt-tuning methods focus on language branch or learn vision-language interaction in a shallow mechanism. |
| Approach: | They propose a Deeply coupled Cross-modal Prompt learning method based on CLIP to facilitate the interplay between vision and language with a Cross-Modal Prompting Attention mechanism. |
| Outcome: | The proposed method enables the interplay between vision and language with a Cross-Modal Prompt Attention mechanism. |
Copied to clipboard
| Challenge: | Existing approaches to end-to-end speech translation (E2E) models only allow one way knowledge transfer, which is limited by the performance of the teacher model. |
| Approach: | They propose a one-way knowledge transfer paradigm where the MT and ST models are collaboratively trained and considered as peers rather than teacher/student. |
| Outcome: | The proposed model improves the performance of end-to-end speech translation (ST) task by combining knowledge from two models with peer models. |
Copied to clipboard
| Challenge: | Empirical results show that MoMs consistently outperform vanilla transformers . |
| Approach: | They propose an architecture that allows for a mixture-of-modules computation that uses a finite set of modules defined by multi-head attention and feed-forward networks. |
| Outcome: | The proposed architecture outperforms vanilla Transformers and their variants in multiple ways. |
Copied to clipboard
| Challenge: | Existing approaches to test-time scaling are limited due to the quality of candidate responses. |
| Approach: | They propose a new metric to quantify the relative improvement of self-refinement beyond majority voting. |
| Outcome: | The proposed method achieves state-of-the-art performance across five benchmarks over other methods. |
Copied to clipboard
| Challenge: | Temporal Knowledge Graphs (TKGs) store dynamic facts in the real world. |
| Approach: | They propose a Spatial-Temporal Knowledge Adapter which integrates the evolving graph encoder and the LLM to facilitate TKG reasoning. |
| Outcome: | The proposed method outperforms state-of-the-art methods on benchmark datasets and exhibits strong generalization capabilities in cross-dataset task. |
Copied to clipboard
| Challenge: | Social media is a key platform for emotional expression, yet deep learning lacks flexibility and interpretability. |
| Approach: | They propose to use Chinese social media to train interpretable mental health instruction datasets to test models' ability to explain their decisions. |
| Outcome: | The proposed models outperform deep learning and LLMs on three mental health downstream tasks and demonstrate their potential for clinical applications. |
Copied to clipboard
| Challenge: | Existing methods for learning a robust matching model from noisy training data are retrieval-based or generation-based. |
| Approach: | They propose a general co-teaching framework that learns matching models from noisy training data. |
| Outcome: | The proposed learning framework can improve existing models on two public data sets. |
Copied to clipboard
| Challenge: | Generative retrieval (GR) is an emerging search paradigm for food delivery search. |
| Approach: | They propose a method that harnesses the advanced query understanding capabilities of large language models to enhance the retrieval of results for complex and long-tail queries in food delivery search scenarios. |
| Outcome: | The proposed method increases the number of online orders by 0.68% for complex search intents. |
Copied to clipboard
| Challenge: | Existing offline approaches to improve an LLM-based customer support system rely on batch annotations. |
| Approach: | They propose an agent-in-the-loop framework that integrates four key types of annotations directly into live customer operations: (1) pairwise response preferences, (2) agent adoption and rationales, (3) knowledge relevance checks, and (4) identification of missing knowledge. |
| Outcome: | The proposed framework reduces retraining cycles from months to weeks by integrating four key types of annotations directly into live customer operations. |
Copied to clipboard
| Challenge: | Emotion-cause pair extraction aims to extract all emotion clauses coupled with their cause clauses from a given document. |
| Approach: | They propose a one-step neural approach which emphasizes inter-clause modeling to perform end-to-end extraction. |
| Outcome: | The proposed method outperforms existing methods in the extraction of emotion-cause pairs . it emphasizes inter-clause modeling to perform end-to-end extraction . |
Copied to clipboard
| Challenge: | Existing visual dialog methods use RNN to encode the dialog history as a vector representation . a new method for video dialog is proposed, which progressively updates query information based on dialog history and video content until the agent think the information is sufficient and unambiguous. |
| Approach: | They propose a method which progressively updates query information based on dialog history and video content until the agent thinks it is sufficient and unambiguous. |
| Outcome: | The proposed method can be used to infer video dialog answers on large-scale datasets. |
Copied to clipboard
| Challenge: | Word embeddings trained on human-generated corpora inherit strong gender stereotypes . prior studies show such embeddables exhibit social biases, such as gender stereotype . |
| Approach: | They propose a method to preserve gender information in certain dimensions of word vectors . they propose GN-GloVe, which is a gender-neutral variant of the word embedding model . |
| Outcome: | The proposed method preserves gender information in certain dimensions of word vectors while compelling other dimensions to be free of gender influence. |
Copied to clipboard
| Challenge: | OpenAI introduces deliberative alignment (DA) to enhance safety of its o-series models, but effectiveness of this approach in open-source LLMs is understudied. |
| Approach: | They propose a case-augmented deliberative alignment method for large language models . they propose to use reinforcement learning on self-generated safety reasoning chains . |
| Outcome: | The proposed method avoids narrowly enumerated rules and allows broader adaptability. |
Copied to clipboard
| Challenge: | Recent studies have shown that powerful pre-trained language models can be fooled by small perturbations or intentional attacks. |
| Approach: | They propose a framework for fine-tuning PLMs using a masked language model and Gaussian noise to augment semantically relevant examples with sufficient diversity. |
| Outcome: | The proposed framework improves the robustness of pre-trained language models and alleviates performance degradation under adversarial attacks. |
Copied to clipboard
| Challenge: | Pro-drop (‘pronoun-dropping’) language requires NMT systems to recover omitted pronouns, but this task lacks sufficient datasets for benchmarking . |
| Approach: | They propose a benchmarking method that leverages the semantic embedding of dropped pronouns to augment training pairs to alleviate the negative impact introduced by pro-drop . |
| Outcome: | The proposed method outperforms existing methods regarding omitted pronoun retrieval and overall translation quality on four Chinese-English translation corpora. |
Copied to clipboard
| Challenge: | Recent research has focused on literary machine translation (MT) but evaluation of literary MT remains an open problem. |
| Approach: | They propose a paragraph-level parallel corpus containing verified human translations and 13k evaluated sentences across four language pairs. |
| Outcome: | The proposed corpus compares human evaluations with students and professionals . it shows that the adequacy of human evaluation is controlled by two factors . |
Copied to clipboard
| Challenge: | Existing paraphrase-based watermark removal methods struggle to balance efficacy with text quality. |
| Approach: | They propose a training-free evolutionary framework that models watermark removal as a constrained multi-objective optimization problem by using a Pseudo-Log-Likelihood-guided mutation to precisely target and modify watermark-carrying tokens. |
| Outcome: | The proposed method outperforms baseline methods on a Qwen3 series watermark scheme while maintaining high semantic fidelity. |
Copied to clipboard
| Challenge: | Multimodal reasoning is a key capability for large vision-language models . however, the vanilla Chain-of-Thought method fails to address critical steps in multi-step reasoning tasks. |
| Approach: | They propose a bi-modal Behavioral Alignment method to augment multimodal reasoning . they use domain-specific language to integrate multimodal information into a precise alternative form . |
| Outcome: | The proposed method significantly improves GPT-4V(ision) on geometry problem solving, chess positional advantage prediction and molecular property prediction. |
Copied to clipboard
| Challenge: | Existing models for large vision language models do not fully reflect their knowledge capacity and reliability, resulting in erroneous outputs that do not align with the image content or provide answers lacking knowledge evidence. |
| Approach: | They propose a Chinese-based benchmark for visual factuality across 8 major topics and 56 subtopics and a multi-hop question construction. |
| Outcome: | The proposed model decouples visual factuality into two parts: seeing the world and discovering knowledge. |
Copied to clipboard
| Challenge: | Existing methods to make comments on articles are based on human-annotated subsets, but they are not suitable for online forums. |
| Approach: | They propose to use a large-scale Chinese corpus with millions of real comments and a human-annotated subset characterizing the comments’ varying quality to generalize a broad set of popular reference-based metrics. |
| Outcome: | The proposed model incorporates human-annotated subset characterizing the comments’ varying quality and shows that it is more accurate than previous models. |
Copied to clipboard
| Challenge: | Recent advances in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. |
| Approach: | They propose to learn straight flow for fast simulation by using flashAudio with rectified flows and immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. |
| Outcome: | The proposed method can learn straight flow for fast simulations and reduce noise distribution. |
Copied to clipboard
| Challenge: | Existing approaches to perform large-scale query-passage retrieval are term-based, but they lose interaction between query-pastage pairs. |
| Approach: | They propose to fuse query (passage) information into query representations via graph neural networks that are constructed by queries and their top retrieved passages. |
| Outcome: | The proposed model outperforms existing models on MSMARCO, Natural Questions and TriviaQA datasets and achieves the new state-of-the-art on these datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly used for creative tasks such as literary translation. |
| Approach: | They propose a paired-task framework that assesses translational creativity using Units of Creative Potential (UCPs) they benchmark 23 models and four creativity-oriented prompts to assess translational comprehension . |
| Outcome: | The proposed framework compares 23 models and four creativity-oriented prompts on literary excerpts from 11 books. |
Copied to clipboard
| Challenge: | Existing methods for offsite-tuning of large language models require high computational costs and lack theoretical analysis. |
| Approach: | They propose an offsite-tuning approach that selectively applies compression techniques such as rank compression and channel pruning to preserve the gradients of fine-tuned adapters while ensuring privacy. |
| Outcome: | The proposed method surpasses existing OT methods in privacy protection and model performance. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but the complexity of emerging tasks and higher performance demands highlight the need for continuous improvement. |
| Approach: | They propose a method that refines evaluation results and characterizes model profiles at the knowledge component level. |
| Outcome: | The proposed method improves performance across multiple benchmarks and academic exams. |
Copied to clipboard
| Challenge: | Existing methods for multitask learning fail to match input semantics with expert capabilities, leading to weak expert specialization. |
| Approach: | They propose a parameter-efficient mixture-of-experts framework for task-adaptive learning that aligns textual semantics with the most suitable experts for precise routing. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods and holds excellent task generalization capabilities. |
Copied to clipboard
| Challenge: | Existing researches focus on sentence matching but the interaction of opinions requires reasoning of knowledge, which is beyond textual information. |
| Approach: | They propose to leverage external knowledge to enhance the identification of interactive argument pairs by analyzing the discussion thread of the target topic in an online forum. |
| Outcome: | The proposed model achieves state-of-the-art in the benchmark dataset. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning. |
| Approach: | They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios. |
| Outcome: | The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics. |
Copied to clipboard
| Challenge: | Large language models struggle to process lengthy inputs due to limited length generalization and attention’s quadratic computational demands. |
| Approach: | They propose a training-free framework that allows each head to attend to important context chunks instead of allowing each head a full sentence . |
| Outcome: | The proposed framework unlocks multi-head attention's untapped potential by allowing each head to attend to important context chunks instead of the full sentence. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) are becoming tool-interactive agents . quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data . et al.: a new approach to scale trajectory diversity improves tool-use generalization . |
| Approach: | They propose a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. |
| Outcome: | Experiments on general tool-use benchmarks and code agent tasks show that TDScaling improves tool-user generalization and inherent coding proficiency. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been successful in understanding language and processing text, but their cost prohibits their practical applications. |
| Approach: | They propose a multi-agent collaboration method that breaks down lengthy documents into smaller, more manageable chunks and organizes the member agents to read their assigned chunks. |
| Outcome: | The proposed method achieves 16.42% and 1.63% accuracy gains over existing models on single-hop and multi-hop QA settings. |
Copied to clipboard
| Challenge: | Existing benchmarks evaluate multi-session memory in text-only conversations or assess multimodal understanding within localized contexts. |
| Approach: | They propose a benchmark for evaluating multimodal long-term conversational memory in MLLM agents. |
| Outcome: | The proposed framework assesses key memory capabilities along three functional dimensions: memory extraction and test-time adaptation, memory reasoning, and memory knowledge management. |
Copied to clipboard
| Challenge: | Existing studies focus on language-based premises and deduce valid conclusions from visual observations. |
| Approach: | They propose a rule-based deductive reasoning task that uses video to deduce the correct future event . they use commonsense knowledge to annotate video and a strong baseline to conduct reasoning . |
| Outcome: | Empirical studies validate the rationality of ARTNet in deductive reasoning upon visual observations . ART is a method that rigorously follows a set of explicit constraints to deduce valid conclusions from empirical facts . |
Copied to clipboard
| Challenge: | Emotional Intelligence (EI) is a key concept in the field of human intelligence. |
| Approach: | They propose a method to enhance EI of large language models by naive fine-tuning on EI-related tasks. |
| Outcome: | The proposed method improves EI of two LLM-based assistants without compromising GI. |
Copied to clipboard
| Challenge: | Existing safety benchmarks focus on explicitly harmful content, but ignore context-dependent expressions such as dogwhistles. |
| Approach: | They propose a benchmark for evaluating LLM safety under dogwhistle-driven prompts . their findings expose a blind spot in current safety evaluation practices . |
| Outcome: | The proposed benchmark compared safety performance with toxic terms using dogwhistle-driven prompts. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). |
| Approach: | a new study proposes a domain-informed self-consistency policy optimization extension to GRPO that addresses inter-group imbalance. |
| Outcome: | a new extension of GRPO addresses inter-group imbalance with two key innovations . the proposed method outperforms existing GR PO variants by 5% on Qwen3 models . |
Copied to clipboard
| Challenge: | Existing studies focus on constructing a matching model with sophisticated neural architectures, but do little to how to effectively learn such architectures from data. |
| Approach: | They propose to sample negative examples to automatically construct a training set for effective model learning in retrieval-based dialogue systems by using four sampling strategies. |
| Outcome: | The proposed learning method improves the performance of matching models on two benchmarks with three matching models. |
Copied to clipboard
| Challenge: | Existing studies fail to provide comprehensive service satisfaction analysis . Existing models fail to include satisfaction polarity classification and sentimental utterance identification . |
| Approach: | They propose a model that predicts customer sentiments and aggregates them into service satisfaction polarity. |
| Outcome: | The proposed model predicts customer sentiments and aggregates them into service satisfaction polarity and reasoning clues. |
Copied to clipboard
| Challenge: | Existing approaches to VideoQA often fail when complex reasoning or temporal relationships are involved. |
| Approach: | They propose a method that leverages reasoning processes generated by Multimodal Large Language Models to improve VideoQA models. |
| Outcome: | The proposed method improves VideoQA models on three benchmarks. |
Copied to clipboard
| Challenge: | Existing QR systems that reformulate defective user queries are limited in English due to the scarcity of non-English QR labels. |
| Approach: | They propose a query reformulation method which reformulates defective user queries to improve non-English QR performance. |
| Outcome: | The proposed framework improves non-English QR performance by leveraging abundant reformulation resources in English. |
Copied to clipboard
| Challenge: | Existing methods to predict future traffic flows capture spatio-temporal dependencies, but they fail to adapt to test-time environmental changes. |
| Approach: | They propose to use large language models to help traffic flow forecasting by capturing spatio-temporal dependencies and using a large language model to select the most likely result. |
| Outcome: | The proposed method is based on large language models (LLMs) and an LLM-based selector. |
Copied to clipboard
| Challenge: | Existing Large Language Models struggle to maintain emotionally consistent and psychologically plausible character personalities. |
| Approach: | They propose a framework that enhances LLMs’ ability to generate authentic personality-driven dialogues through stochastic emotional transitions. |
| Outcome: | The proposed framework achieves 93.3% emotional accuracy on the RAPD dataset and significantly outperforms existing approaches. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) often hallucinate due to fragile, linear reasoning and weak visual grounding. |
| Approach: | They propose a framework that reformulates reasoning as a hierarchical search with self-verification and replaces linear Chain-of-Thought with a tree-search policy capable of backtracking to correct logical errors. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on hallucination and safety benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in Instruction-fine-tuned Vision and Language Models (IVLMs) have prompted some studies to analyze the reasoning capabilities of IVLMs. |
| Approach: | They introduce a vision and language task for Inductive Visual Reasoning that uses common attributes across visual scenes to find common answers. |
| Outcome: | The proposed model can archive with 48% accuracy on the FTC, compared with state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing methods to extract relations from distant supervision contain low-quality instances with noisy words and overlapped relations. |
| Approach: | They propose a Regularized Attentive Capsule Network to better identify overlapped relations in informal sentences . they embed multi-head attention into the capsule network as the low-level capsules . |
| Outcome: | Extensive experiments show that the proposed model improves relation extraction. |
Copied to clipboard
| Challenge: | Current neural architecture search methods suffer from huge computational cost. |
| Approach: | They propose a reversible recursive backpropagation algorithm that uses the last layer to store the outputs of the network. |
| Outcome: | The proposed algorithm outperforms standard Transformers on three sequence-to-sequence datasets. |
Copied to clipboard
| Challenge: | Empirical results indicate that pre-trained language models can significantly outperform state-of-the-art methods in both automatic evaluation and human judgment. |
| Approach: | They propose to equip a pre-trained language model with a knowledge selection module to generate knowledge-grounded dialogues. |
| Outcome: | The proposed model outperforms state-of-the-art methods in evaluation and human judgment. |
Copied to clipboard
| Challenge: | Existing detectors for Large Language Models (LLMs) struggle to generalize in open-world settings. |
| Approach: | They propose a framework to detect LLM-generated text with exceptional generalization to unseen domains by reinforcing LLMs’ inherent rewriting tendencies. |
| Outcome: | The proposed framework outperforms state-of-the-art detection methods by 23.04% in AUROC, 35.10% for out-of distribution tests, and 48.66% under adversarial attacks. |
Copied to clipboard
| Challenge: | Existing benchmarks assess integrated and agent-oriented scientific reasoning in isolation . Existing systems assess integrated reasoning in isolated tasks . |
| Approach: | They propose a benchmark to evaluate integrated and agent-oriented scientific reasoning over research papers. |
| Outcome: | The proposed benchmark evaluates integrated and agent-oriented scientific reasoning over scientific papers. |
Copied to clipboard
| Challenge: | Existing evaluation methods suffer from cognitive dimensional simplification and methodological unreliability due to the ”LLM-as-a-Judge” approach. |
| Approach: | They propose a six-tiered benchmark that evaluates ASG systems by prioritizing deterministic algorithms and introducing a GRADE approach for abstract abilities. |
| Outcome: | The proposed method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research. |
Copied to clipboard
| Challenge: | Extensive evaluations of large language models (LLMs) are conducted on a wide range of models, revealing a notable cultural-linguistic synergy phenomenon, where models exhibit better performance when questions are culturally aligned with the language. |
| Approach: | They propose a Dual Evaluation Framework to comprehensively assess the multilingual capabilities of large language models by decomposing evaluation along dimensions of linguistic medium and cultural context. |
| Outcome: | The proposed framework allows for a nuanced analysis of LLMs’ ability to process questions within both native and cross-cultural contexts cross-lingually. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge. |
| Approach: | They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth. |
| Outcome: | The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation. |
Copied to clipboard
| Challenge: | Existing methods for enhancing Large Language Models (LLMs) struggle with novelty and Reinforcement Learning from human feedback (RLHF) is costly. |
| Approach: | They propose to use a Reward Model (RM) and a principle-guided LLM-as-a-Judge to enhance creative output over baselines. |
| Outcome: | The proposed approach significantly enhances creative output over baselines, but the principle-guided LLM-as-a-Judge yields superior generation quality. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive performance across various mathematical reasoning benchmarks. |
| Approach: | They introduce an adversarial grade school math dataset and explore whether LLMs can be more robust when questions are slightly changed. |
| Outcome: | The proposed method generates and verifies each intermediate thought based on its reasoning goal and calculation result. |
Copied to clipboard
| Challenge: | Existing approaches to extract rich correlations between entities and relations are not fully exploited by existing methods. |
| Approach: | They propose to unify entities and relations by jointly encoding them within a concatenated natural language sequence and unify the modeling of interactions with a proposed Interaction Map. |
| Outcome: | The proposed method is more efficient and efficient than existing methods and can be scaled up to 2021. |
Copied to clipboard
| Challenge: | Existing approaches to solve non-deterministic reasoning problems in large language models are limited by their complexity and lack of a clear understanding of the problem. |
| Approach: | They propose a method to diagnose and correct non-deterministic reasoning behaviors in large language models. |
| Outcome: | The proposed method outperforms baselines and WebQSP benchmarks on the widely used WebQ SP and CWQ benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have seen remarkable progress in providing general instruction-following ability, but struggle with critical problems when required to provide a detailed and accurate response to a visual instruction. |
| Approach: | They propose to enhance the mapping process by using retrieval-augmented tag tokens, which contain rich object-aware information such as object names and attributes. |
| Outcome: | The proposed model outperforms baselines that share the same language model and training data on 12 benchmarks and shows zero-shot capability when provided with specific datastores. |
Copied to clipboard
| Challenge: | Existing named entity recognition systems require large scale labeled data to perform, while annotation of NER data is laborious and time-consuming. |
| Approach: | They propose to adjust an existing named entity recognition system to recognize entity types not defined in the system. |
| Outcome: | The proposed method can be quickly adjusted to a named entity recognition system. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on scalar rewards to capture user preferences. |
| Approach: | They propose a framework that integrates multi-objective reward modeling to represent diverse preference profiles. |
| Outcome: | The proposed method improves performance across reward objectives and targets. |
Copied to clipboard
| Challenge: | Document-level context is crucial for speech translation due to noise from ASR . incorporating document-level contextual information into ST remains a challenge . |
| Approach: | They develop an online framework that integrates document-level context into machine translation . they use document-based modules to integrate document- level context into ST . |
| Outcome: | The proposed framework outperforms baselines in sentence and discourse metrics . it can correct ASR transcription errors and improve translation performance . |
Copied to clipboard
| Challenge: | Recent studies have discussed its capability to assist language models for various applications. |
| Approach: | They propose a structure to organize arguments using the **Hi**erarchical **Ar**gumentation **G**raph (Hi-ArG) and propose two approaches to exploit Hi-AarG, including a text-graph multi-modal model GreaseArR and a framework augmented with graph information. |
| Outcome: | The proposed structure supersedes existing language models on two argumentation tasks while incorporating graph information during further training improves vanilla language models. |
Copied to clipboard
| Challenge: | Existing evaluation methods for document summarization require human annotations and annotations. |
| Approach: | They propose a method which measures the quality of a summary by measuring its semantic similarity with a pseudo reference summary, using contextualized embeddings and soft token alignment techniques. |
| Outcome: | The proposed method correlates better with human ratings by 18- 39% compared to the state-of-the-art evaluation metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have exhibited remarkable proficiency across a wide array of NLP tasks. |
| Approach: | They propose a method for pruning large language models using general or task-specific weights to extract a compressed, task-agnostic LLM. |
| Outcome: | The proposed method extracts a compressed, domain-specific, and task- agnostic LLM by identifying LLM weights that are pivotal for general capabilities, like linguistic capability and multi-task solving, and domain- specific knowledge. |
Copied to clipboard
| Challenge: | Existing methods for commonsense question generation produce shallow questions that can be answered by simple word matching. |
| Approach: | They propose a task of commonsense question generation that aims to yield deep-level questions from the text. |
| Outcome: | The proposed model can yield deep-level and to-the-point questions from the text. |
Copied to clipboard
| Challenge: | Existing evaluations of persona-emotion decoupling frameworks do not isolate which component caused the failure . current evaluations do little to isolate the cause of the failure, making fixes slow and ad hoc . |
| Approach: | They propose a diagnostic evaluation framework that decomposes persona expression into two observable routes. |
| Outcome: | The proposed framework decomposes persona expression into two observable routes . it can be used to perform route-comparable, reference-based analyses of separability, drift, failures and coupling . |
Copied to clipboard
| Challenge: | Existing methods for evaluation of dialog systems are expensive and not scalable . a framework for estimating human evaluation scores is proposed to bridge this gap . |
| Approach: | They propose a framework for estimating human evaluation scores based on off-policy evaluation . they use language quality metrics for single-turn response generation given a fixed context . |
| Outcome: | The proposed framework outperforms existing methods in terms of correlation with human evaluation scores. |
Copied to clipboard
| Challenge: | Existing benchmarks for automated grading of student work fail to evaluate real student responses . existing models fail to assess real student work, especially on cognitively demanding tasks . |
| Approach: | They propose a multimodal benchmark for rubric-aligned evaluation of real Chinese K-12 student answers. |
| Outcome: | The proposed model improves performance and interpretability of existing models on EduMARS . existing models fail to perform on real-world, cognitively demanding tasks, authors say . |
Copied to clipboard
| Challenge: | Recent studies on video large language models focus on model architectures and training datasets . interaction format between user and model is unsatisfactory for time-sensitive tasks . |
| Approach: | They propose a video-text duet interaction format that allows for continuous playback of the video . when a text message ends, the video continues to play, similar to the alternative of two performers in a duet. |
| Outcome: | The proposed format improves performance on time-sensitive tasks with minimal training efforts. |
Copied to clipboard
| Challenge: | Existing studies on integrating online community to solve social problems have not fully utilized these three components and the relationship among them. |
| Approach: | They propose a framework that simultaneously considers communities, users, and texts and can easily connect with a variety of downstream tasks related to social media. |
| Outcome: | The proposed model can be used to perform violation detection, sentiment analysis, and community recommendation across multiple tasks. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated remarkable performance on few-shot Named Entity Recognition tasks due to the high cost of obtaining high-quality labeled data. |
| Approach: | They propose to decompose the task into entity span detection and entity type classification using a type-independent entity span detector and then classify the detected spans based on their types. |
| Outcome: | The proposed method consistently yields improvements over two baseline approaches. |
Copied to clipboard
| Challenge: | Influence functions are important for quantifying the impact of individual training data points on a model’s predictions. |
| Approach: | They conduct a systematic study to address a key question: do influence functions work on large language models? |
| Outcome: | The influence functions perform poorly across multiple tasks and are therefore unsuitable for large language models. |
Copied to clipboard
| Challenge: | Existing supervised relation extraction methods can still misclassify unknown relations into known relations due to the lack of supervision signals. |
| Approach: | They propose a method that regularizes the model by dynamically synthesizing negative instances that can provide the missing supervision signals. |
| Outcome: | The proposed method achieves SOTA unknown relation detection without compromising the classification of known relations. |
Copied to clipboard
| Challenge: | Earlier efforts in text modeling have achieved limited success on word meanings . convolutional neural networks (CNNs) are used to model higher level concepts and facts in texts . |
| Approach: | They propose three strategies to stabilize dynamic routing process to alleviate disturbance of noise capsules. |
| Outcome: | The proposed methods achieve state-of-the-art on 4 out of 6 datasets . they show that capsule networks exhibit significant improvement over baseline methods . |
Copied to clipboard
| Challenge: | Neural machine translation systems have become state-of-the-art approaches for Grammatical Error Correction (GEC) task. |
| Approach: | They propose a copy-augmented architecture for the Grammatical Error Correction task by copying unchanged words from the source sentence to the target sentence. |
| Outcome: | The proposed architecture outperforms all recently published state-of-the-art results by a large margin. |
Copied to clipboard
| Challenge: | Existing studies focus on contrastive learning on the instance level without discriminating the contribution of each word. |
| Approach: | They propose a hierarchical contrastive learning mechanism which can unify semantic meaning in the input text. |
| Outcome: | The proposed model outperforms baselines on storytelling, paraphrasing, dialogue generation, and storytelling tasks. |
Copied to clipboard
| Challenge: | Existing methods for jailbreaking large language models rely on laborious human engineering and whitebox access to model internals. |
| Approach: | They propose a method that instructs large language models to deviate from prior context and generate harmful outputs by instructing them to deviat from previous attacks. |
| Outcome: | The proposed method achieves a 62.83% higher success rate in compromising ten leading chatbots, while using only 12.9% of the queries. |
Copied to clipboard
| Challenge: | Grammatical Error Correction (GEC) systems perform well in academic benchmarks, but in practical applications they may not correct errors when users perform irrelevant modifications. |
| Approach: | They propose a benchmark to evaluate the context robustness of Grammatical Error Correction systems. |
| Outcome: | The proposed method improves the accuracy of errors corrected by human annotations. |
Copied to clipboard
| Challenge: | Existing studies have shown that reflection can enhance performance, but our investigation reveals an undesirable pattern in reflection framework: effective self-reflection occurs primarily at the beginning of iterations, with subsequent attempts failing to produce further improvements. |
| Approach: | They propose a framework that generates task-adaptive reflection advice using an external open-source small language model. |
| Outcome: | The proposed framework generates task-adaptive and diverse reflection advice in MiniWoB++ and Alfworld environments. |
Copied to clipboard
| Challenge: | Reinforcement learning from human feedback (RLHF) is the primary method for aligning large language models with human preferences. |
| Approach: | They propose to train an Absolute-Rating Multi-Objective Reward Model with multi-dimensional absolute-rating data. |
| Outcome: | The proposed model outperforms the LLM-as-a-judge method on RewardBench . it achieves state-of-the-art performance on the benchmark . |
Copied to clipboard
| Challenge: | Existing methods for hierarchical text classification are lacking in the field of natural language processing. |
| Approach: | They propose a hierarchy-aware T5 model with path-adaptive attention mechanism to exploit hierarchical dependency across different levels. |
| Outcome: | The proposed model outperforms state-of-the-art models especially in Macro-F1 and low Macro. |
Copied to clipboard
| Challenge: | Prior work limits search depth to reduce cost, but this often leads to underexploration of complex questions. |
| Approach: | They propose a reinforcement learning framework that evaluates each search step via self-generated intermediate answers. |
| Outcome: | Extensive experiments on multiple benchmarks show that AutoSearch achieves a superior accuracy-efficiency trade-off, alleviating over-searching while preserving search quality. |
Copied to clipboard
| Challenge: | Existing studies have failed to scale up the pre-training process by putting aside unlabeled data . et al., 2019: multi-party dialogues are more difficult for models to understand since they involve multiple interlocutors resulting in interweaving reply-to relations and information flows. |
| Approach: | They propose to treat discourse structures as latent variables and jointly infer them to pre-train a model that understands the discourse structure of multi-party dialogues. |
| Outcome: | The proposed model outperforms baselines and achieves state-of-the-art results on multiple downstream tasks. |
Copied to clipboard
| Challenge: | Existing defense methods focus on detecting harmful prompts or reducing the likelihood of harmful responses. |
| Approach: | They propose a layer-specific editing method to align LLMs to harmful prompts by supervised fine-tuning and reinforcement learning. |
| Outcome: | The proposed method improves the performance of large language models against jailbreak attacks while maintaining performance on benign prompts. |
Copied to clipboard
| Challenge: | Existing methods for Chinese spelling error correction focus on local contextual information, thus misleading the user and reducing performance. |
| Approach: | They propose a global attention decoder that learns the global relationship of correct input characters and candidates of potential error characters. |
| Outcome: | The proposed method outperforms all competitor models by a large margin of up to 6.2% on three human-annotated datasets. |
Copied to clipboard
| Challenge: | Multiple-choice question datasets like Massive Multitask Language Understanding (MMLU) have inevitably led to benchmark contamination, resulting in unreliable evaluation. |
| Approach: | They propose a contamination-free MCQ benchmark called MMLU-CF which reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. |
| Outcome: | The proposed MMLU-CF reassesses LLMs’ understanding of world knowledge by averting both unintentional and malicious data contamination. |
Copied to clipboard
| Challenge: | Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world, but they currently require large pretraining corpora or access to typologically similar languages. |
| Approach: | They propose to remove language identity signals from multilingual embeddings by re-aligning vector spaces of target languages to a pivot source language and removing language-specific means and variances. |
| Outcome: | The proposed approaches reduce cross-lingual transfer gap by 8.9 points (m-BERT) and 18.2 points (XLM-R) on average across all tasks and languages. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) are powerful but still suffer from inefficient and off-target reasoning. |
| Approach: | They propose a training-free framework that automatically optimizes Large Reasoning Models' reasoning by generating think-prefixes that evolve driven by a taxonomy of reasoning behaviors. |
| Outcome: | The proposed framework significantly improves accuracy-length trade-off for efficient reasoning, drastically improves safety and improves instruction following. |
Copied to clipboard
| Challenge: | Existing methods for concept-level grounding and instruction-level reasoning use coarse representations and iterative mask filtering. |
| Approach: | They propose an instruction-following extension of the Segment Anything Model 3 family that unifies concept-level grounding and instruction-level reasoning within a single segmentation framework. |
| Outcome: | Experiments show that SAM3-I achieves appealing performance across referring and reasoning-based segmentation while maintaining its strong concept recall ability. |
Copied to clipboard
| Challenge: | Existing multimodal sentiment analysis methods are limited to textual data and cannot handle multimodal scenarios. |
| Approach: | They propose a transfer learning framework that allows cross-lingual and cross-modal alignments and a language family disentanglement module that enhances the sharing of language universals within families. |
| Outcome: | The proposed method is superior to existing methods and can handle low-resource languages. |
Copied to clipboard
| Challenge: | Continual pre-training (CPT) has been an important approach for adapting language models to specific domains or tasks. |
| Approach: | They propose a Continual pre-training method that can greatly improve Chinese language ability and scientific reasoning ability of LLMs. |
| Outcome: | The proposed method can greatly improve Chinese language ability and scientific reasoning ability of LLMs. |
Copied to clipboard
| Challenge: | Recent approaches to enhance agent performance focus on outcome rewards, which may lead to errors or suboptimal actions due to the absence of process supervision signals. |
| Approach: | They propose a step-level framework that provides detailed step-by-step guidance to enhance agent training by using Monte Carlo methods. |
| Outcome: | The proposed framework outperforms strong baselines on three tasks and shows that it is effective in augmenting efficiency and its applicability to diverse models. |
Copied to clipboard
| Challenge: | LSLMs have impressive conversational generation abilities, but consistently fall short of traditional pipeline systems on semantic understanding benchmarks. |
| Approach: | They propose to analyze the performance gap between speech and text inputs through a systematic experiment . they find that representation similarity is strongly correlated with the modality gap . |
| Outcome: | The proposed models improve the accuracy of speech inputs and their semantic understanding benchmarks. |
Copied to clipboard
| Challenge: | Existing dictionaries do not capture the full range of polysemous and homonymous words corresponding to different signs across contexts. |
| Approach: | They analyze 1,404 word use–to–sign ID mappings from German and German Sign Language . they identify three correspondence types: Type 1 (one-to-many), Type 2 (many-to-1), and Type 3 (one to one) |
| Outcome: | The proposed method outperforms existing methods using Exact Match and Semantic Similarity. |
Copied to clipboard
| Challenge: | Concatenating large language models are adapted to context-aware neural machine translation in a concatenated way . a recent paradigm shift has been witnessed in discourse-related challenges such as zero pronoun translation . |
| Approach: | They propose an alternative adaptation approach to make large language models discriminately model and utilize inter- and intra-sentence contexts. |
| Outcome: | The proposed approach outperforms concatenation mode and improves performance in discourse modeling. |
Copied to clipboard
| Challenge: | Recent development of large language models (LLMs) have boosted interest on dialogue agents . however, research on these tasks is limited by the insufficiency of public datasets . stance detection and debate summarization are key for engaging argumentative dialogues - but are not available for non-English languages. |
| Approach: | They propose to use ORCHID to benchmark stance detection and debate summarization in Chinese debates. |
| Outcome: | The proposed task is based on 1,218 real-world debates conducted in Chinese on 476 unique topics. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) predicts sentiment polarity towards a specific aspect in a sentence. |
| Approach: | They propose to use a dynamic aspect-oriented semantics-based method to learn ABSA. |
| Outcome: | The proposed method can learn dynamic aspect-oriented semantics for ABSA on three benchmark datasets. |
Copied to clipboard
| Challenge: | Prompt tuning is effective in extracting knowledge from foundation models, but its effectiveness is uncertain. |
| Approach: | They propose a parametric prompt tuning strategy that dynamically determines different factors of prompts based on specific tasks or instances. |
| Outcome: | The proposed approach improves performance across a wide range of tasks including NLP, vision recognition, and vision-language tasks. |
Copied to clipboard
| Challenge: | Unsupervised bilingual lexicon induction models fail on low-resource language pairs due to insufficient initialization. |
| Approach: | They propose a method to learn cross-lingual features from monolingual corpora for low-resource UBLI by integrating cross-linguistic representations with pre-trained word embeddings in a fully unsupervised initialization. |
| Outcome: | The proposed method outperforms state-of-the-art methods on low-resource language pairs and improves representational ability and robustness of existing embedding models. |
Copied to clipboard
| Challenge: | Abstractive summarization is a task that generates short and concise summaries of user generated reviews. |
| Approach: | They propose an interactive attention mechanism to learn the representations of context and aspect words within reviews, acted as an encoder. |
| Outcome: | The proposed model achieves impressive results compared to other strong competitors on a real-life dataset. |
Copied to clipboard
| Challenge: | Current approaches to commonsense reasoning are limited due to limited answer scope. |
| Approach: | They propose to solve a commonsense question without a pre-defined answer scope . they leverage pre-trained language models to iteratively retrieve reasoning paths on the external knowledge base . |
| Outcome: | The proposed method achieves better performance on two commonsense benchmark datasets. |
Copied to clipboard
| Challenge: | Unsupervised sentence representation learning is one of the fundamental problems in natural language processing . contrastive learning methods fail to capture fine-grained ranking information among the sentences . |
| Approach: | They propose a novel approach for unsupervised sentence representation learning that integrates ranking consistency and ranking distillation with contrastive learning into a unified framework. |
| Outcome: | The proposed approach performs better over state-of-the-art models on STS and TR tasks. |
Copied to clipboard
| Challenge: | Existing work on pre-training models have shown that it is important to use a framework to deploy various pre- training models efficiently. |
| Approach: | They propose an assemble-on-demand pre-training toolkit that assembles pre-trained models on demand and encapsulates them with rich modules. |
| Outcome: | The proposed framework can reproduce state-of-the-art models or develop models that remain unexplored. |
Copied to clipboard
| Challenge: | Small language models (SLMs) are promising for real-world deployment but struggle with high-stakes legal reasoning tasks. |
| Approach: | They propose a diagnostic-driven synthesis framework that extracts and refines reasoning trajectories from a capable teacher via fine-grained prompting and a self-reflective verification is employed to adaptively select the most effective data for the SLM student. |
| Outcome: | The proposed framework extracts and refines reasoning trajectories from a capable teacher via fine-grained prompting, then a self-reflective verification is employed to adaptively select the most effective data for the student. |
Copied to clipboard
| Challenge: | Existing methods rely on textual similarities between NL and KG to build relation links. |
| Approach: | They propose an implicit relation linking method called ImRL which links relation phrases in NL to relation paths in KG. |
| Outcome: | The proposed method significantly outperforms state-of-the-art methods on two benchmarks and a newly-created datasets. |
Copied to clipboard
| Challenge: | Existing work uses sentences within the same batch as negatives, which suffers from easy negatives. |
| Approach: | They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment . |
| Outcome: | The proposed model achieves state-of-the-art on several tasks. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for retrievers are narrow and evaluate them in isolation . existing evaluation benchmarking frameworks focus on evaluating retrievers in isolation, obscuring their value in real-world applications. |
| Approach: | They propose an evaluation framework that evaluates retrievers in agentic search systems . they provide expert-annotated reasoning aspects, positive documents, a reference response and evaluation rubrics . |
| Outcome: | The proposed framework assesses retrievers in agentic search systems. |
Copied to clipboard
| Challenge: | Multi-modal large language models have been used for processing and understanding information from diverse modalities. |
| Approach: | They propose to evaluate the audio-visual capabilities of multi-modal large language models . they focus on effectiveness, efficiency, generalizability, and robustness . |
| Outcome: | The proposed models exhibit strong zero-shot and few-shot generalization abilities . their success relies heavily on the vision modality, which impairs performance when visual input is corrupted or missing. |
Copied to clipboard
| Challenge: | Multi-sentence compression aims to generate a grammatical but reduced compression from multiple input sentences while retaining key information. |
| Approach: | They propose a neural rewriter for multi-sentence compression that does not need any parallel corpus. |
| Outcome: | Empirical studies show that the proposed approach achieves comparable results upon automatic evaluation and improves the grammaticality of compression based on human evaluation. |
Copied to clipboard
| Challenge: | Existing shortening methods for long reasoning models rely on additional supervision or multi-stage post-training. |
| Approach: | They propose a lazy length penalty that imposes length pressure on models without extra training stages. |
| Outcome: | The proposed method significantly reduces response length without extra training stages while maintaining or improving performance. |
Copied to clipboard
| Challenge: | Existing methods for aspect sentiment triplet extraction focus on the single interactions between an aspect and an opinion. |
| Approach: | They propose a multi-overlap triplet extraction method which decodes the complex relations between multiple aspects and opinions by learning their cooperative interactions. |
| Outcome: | The proposed method outperforms baselines, especially multi-overlap triplets. |
Copied to clipboard
| Challenge: | Existing methods for domain adaptation of abstractive dialogue summarization lack generalization ability on new domains. |
| Approach: | They propose a domain-oriented prefix-tuning model that uses a prefix module to alleviate domain entanglement and discrete prompts to guide the model to focus on key contents of dialogues. |
| Outcome: | The proposed model can be generalized to two multi-domain dialogue summarization datasets. |
Copied to clipboard
| Challenge: | Existing methods generate DocIDs based on textual content, which may result in weak semantic connections for similar documents due to variations in expression. |
| Approach: | They propose a new retrieval paradigm that generates unique document identifiers . they propose to use queries as a bridge to connect documents with varying relevance levels . |
| Outcome: | The proposed approach outperforms existing methods on multilingual e-commerce search datasets. |
Copied to clipboard
| Challenge: | Text-based methods lag behind graph embedding-based approaches for knowledge graph completion (KGC) |
| Approach: | They propose three types of negatives to improve contrastive learning to improve learning efficiency. |
| Outcome: | The proposed model outperforms embedding-based methods on several benchmark datasets. |
Copied to clipboard
| Challenge: | Existing Retrieval-Augmented Generation systems treat structure as a physical navigational skeleton rather than intrinsic semantic knowledge. |
| Approach: | They propose a framework that redefining hierarchy as intrinsic semantics and uses snippets to enrich hierarchical lineage. |
| Outcome: | The proposed framework outperforms state-of-the-art hierarchical and graph-based benchmarks on FinTierQA Gold. |
Copied to clipboard
| Challenge: | PromptCoT is a novel approach for synthesizing Olympiad-level math problems . it integrates rationale generation and mathematical concepts to generate complex problems based on concepts and rationale behind problem construction. |
| Approach: | They propose a method for automatically generating high-quality Olympiad-level math problems . they use mathematical concepts and the rationale behind problem construction to synthesize complex problems based on mathematical concepts . |
| Outcome: | The proposed method outperforms existing problem generation methods on benchmarks including GSM8K, MATH-500, and AIME2024. |
Copied to clipboard
| Challenge: | Effidit is a digital writing assistant that provides three modules to help users write faster and more efficiently. |
| Approach: | They present Effidit, a digital writing assistant that provides three modules to help users write higher-quality text more efficiently. |
| Outcome: | Effidit expands the capabilities of a typical writing assistant by providing three modules . Effit can help users create their own text faster and more efficiently . |
Copied to clipboard
| Challenge: | Evaluation metrics are a key ingredient for progress of text generation systems . a class of novel evaluation metrics based on BERT and its variants has been explored . |
| Approach: | They propose to disentangle BERT-based evaluation metrics along linguistic factors . they show they are sensitive to lexical overlap, just like BLEU and ROUGE . |
| Outcome: | The proposed metrics capture all aspects but are sensitive to lexical overlap, just like BLEU and ROUGE, the authors show . |
Copied to clipboard
| Challenge: | Long-context modeling is crucial for next-generation language models, but high computational cost of standard attention mechanisms poses significant computational challenges. |
| Approach: | They propose a natively trained Sparse Attention mechanism that integrates algorithms with hardware-aligned optimizations to achieve efficient long-context modeling. |
| Outcome: | The proposed model maintains or exceeds Full Attention models across general benchmarks, long-context tasks, and instruction-based reasoning. |
Copied to clipboard
| Challenge: | Multi-hop question answering is a challenging task that requires capturing information from multiple positions in multiple documents. |
| Approach: | They propose a framework for integrating text-based and triple-based paradigms that incorporates structured knowledge into large-scale question answering. |
| Outcome: | The proposed framework improves multi-hop question answering by incorporating structured knowledge into the models. |
Copied to clipboard
| Challenge: | Large language models face significant challenges in handling long-context tasks because of their limited effective context window size during pretraining, which restricts their ability to generalize over extended sequences. |
| Approach: | They propose a training strategy for extending the context window of LLMs including impactful token analysis, position index transformation, and training optimization strategies. |
| Outcome: | Experiments on three types of LLMs show that LongRecipe can utilize long sequences while requiring only 30% of the target context window size. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to an increase in synthetic content generation . the ability to detect LLMs-generated content has become of paramount importance . |
| Approach: | They propose to provide a detailed overview of existing detection strategies and benchmarks, scrutinizing their differences and advocating for more adaptable and robust models to enhance detection accuracy. |
| Outcome: | The proposed model will be able to detect human-written content in real time. |
Copied to clipboard
| Challenge: | Synthetic data generation is an increasingly popular way of training models without the need for large, manually labeled datasets. |
| Approach: | They propose a framework that aligns open-source small models to efficiently generate large-scale embedding data. |
| Outcome: | The proposed framework outperforms state-of-the-art embedding models by using only 1/10 of the GPT API calls. |
Copied to clipboard
| Challenge: | Existing studies focus on specific aspects or applications, but this study provides a comprehensive overview of Protein-specific large language models. |
| Approach: | This paper proposes a structured taxonomy of state-of-the-art ProteinLLMs . they analyze how they leverage large-scale protein sequence data for improved accuracy . |
| Outcome: | The proposed model covers their architectures, training datasets, evaluation metrics, and diverse applications. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) produce outputs that deviate from factual reality, especially in sensitive applications such as medical consultation and legal advice. |
| Approach: | They propose a Siamese network-based model that leverages LLMs’ inner states for factual detection. |
| Outcome: | The proposed model achieves over 96% accuracy on a custom-collected factual detection dataset. |
Copied to clipboard
| Challenge: | Existing RAG methods lack fine-grained control over query and source sides, resulting in noisy retrieval and shallow reasoning. |
| Approach: | They propose an agentic RAG framework that integrates information sieving via LLM-as-a-knowledge-router. |
| Outcome: | Experiments on multi-hop QA tasks across heterogeneous sources demonstrate improved reasoning depth, retrieval precision, and interpretability over conventional approaches. |
Copied to clipboard
| Challenge: | Existing methods for OOD detection are based on labeled in-domain data . detecting out-of-domain (OOD) or unknown intents is challenging . |
| Approach: | They propose a novel reassigned contrastive learning method to discriminate IND intents for over-confident OOD and an adaptive class-dependent local threshold mechanism to separate similar IND and OOD intents. |
| Outcome: | The proposed method is effective for both aspects of overconfidence issues. |
Copied to clipboard
| Challenge: | a standard evaluation setup for supervised machine learning tasks does not hold for natural language generation tasks. |
| Approach: | They propose to use reference-free machine translation evaluation to compare source texts to system translations to find key limitations. |
| Outcome: | The proposed metrics perform poorly as semantic encoders for reference-free machine translation evaluation. |
Copied to clipboard
| Challenge: | Existing approaches to open-domain dialogue generation ignore the nature of 1-to-1 mapping that there may exist multiple valid responses corresponding to the same query. |
| Approach: | They propose to model open-domain dialogue generation using 1-to-1 mapping . they first extract common features of different responses and then combine them with distinctive features to generate multiple diverse and appropriate responses. |
| Outcome: | The proposed model outperforms existing models on automatic and human evaluations. |
Copied to clipboard
| Challenge: | Existing studies on in-context learning have focused on quantifying the uncertainty associated with the model's response, but they neglect the complexity of the LLM and the uniqueness of in-constitut learning. |
| Approach: | They propose a method to quantify the uncertainty associated with in-context learning and propose corresponding estimation method to quantify both types of uncertainties. |
| Outcome: | The proposed method offers an unsupervised way to understand the prediction of in-context learning in a plug-and-play fashion. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) is an effective solution to supplement necessary knowledge to large language models. |
| Approach: | They propose a "generate-then-read" pipeline to replace retrieval stage with generation from the LLM itself. |
| Outcome: | The proposed framework outperforms single models in the base and chat versions and addresses safety and helpfulness post-adaptation challenges. |
Copied to clipboard
| Challenge: | Existing models focus on synthesizing a dialogue with proper knowledge, but neglect that the same knowledge could be expressed differently even under the same context. |
| Approach: | They propose a model that ground dialogue generation by extra knowledge by analyzing the structure of the response and the content style of each part. |
| Outcome: | The proposed model can learn the structure style defined by a few examples and generate responses in desired content style. |
Copied to clipboard
| Challenge: | Large language models have demonstrated exceptional performance across a wide range of tasks . however, selecting the optimal LLM to respond to a user query often necessitates a delicate balance between performance and cost. |
| Approach: | They propose a multi-LLM routing framework that efficiently routes user queries to the most suitable LLM. |
| Outcome: | The proposed framework outperforms baseline methods in terms of effectiveness and interpretability. |
Copied to clipboard
| Challenge: | Existing question answering systems use a retriever-reader framework to answer multi-hop questions . existing models lack retrieval, selector, and reasoner capabilities . |
| Approach: | They propose a three-stage text tableQA framework which comprises of retriever, selector, and reasoner. |
| Outcome: | The proposed framework outperforms baseline methods in the few-shot setting and ranks first on the HybridQA leaderboard. |
Copied to clipboard
| Challenge: | Existing approaches to detect semantic change using contextualized embeddings are underperforming . a graph-based clustering approach captures nuanced changes in word senses across time and languages . |
| Approach: | They propose a graph-based clustering approach to capture nuanced changes in word senses across time and languages. |
| Outcome: | The proposed approach surpasses previous approaches in the SemEval2020 binary classification task across four languages. |
Copied to clipboard
| Challenge: | Existing methods for parameter-efficient fine-tuning excel in the context of single-backbone multi-tenant applications. |
| Approach: | They propose to integrate a lightweight vector generator within each Transformer layer to improve prompt-aware representation adjustment. |
| Outcome: | The proposed method surpasses current benchmarks in terms of performance despite having a similar number of adjustable parameters. |
Copied to clipboard
| Challenge: | Existing methods for Open Relation Extraction (OpenRE) use a two-stage pipeline, which learns relation representations and assignments in the first stage, then manually labels relation for each cluster. |
| Approach: | They propose a method that performs relation learning and relation labeling simultaneously without a significant increase in human effort. |
| Outcome: | The proposed method improves existing SOTA methods by 13.8% and 10.6% on two datasets. |
Copied to clipboard
| Challenge: | Despite the impressive multilingual capabilities demonstrated by LLMs, the understanding of how these abilities develop and function remains nascent. |
| Approach: | They propose a novel detection method to pinpoint language-specific neurons within LLMs by selectively activating or deactivating these neurons. |
| Outcome: | The proposed method can “steer” the output language of LLMs by selectively activating or deactivating language-specific neurons. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) show potential for graph extraction, but often yield ill-formed structures or misinterpret logical constructs such as gateways. |
| Approach: | They propose a framework that treats procedural graph extraction as a multi-round reasoning process with structural and logical refinement agents. |
| Outcome: | The proposed framework achieves significant improvements in structural correctness and logical consistency over strong baselines. |
Copied to clipboard
| Challenge: | Existing ensemble methods for ensembling large language models rely on fixed weighting strategies that fail to adapt to dynamic, context-dependent characteristics of LLMs. |
| Approach: | They propose a framework that reformulates LLM ensemble through a Markov Decision Process. |
| Outcome: | The proposed framework outperforms existing methods by 3.3% on a diverse set of tasks while achieving lower time latency. |
Copied to clipboard
| Challenge: | Existing methods for aspect-sentiment analysis ignore internal correlations between aspect extraction and sentiment classification. |
| Approach: | They propose a hierarchical interactive network to model two-way interactions between two tasks appropriately using shallow-level and deep-level inputs. |
| Outcome: | Extensive experiments on three real-world datasets demonstrate that the proposed model outperforms existing methods. |
Copied to clipboard
| Challenge: | In this paper, we introduce SCALE, a collaborative framework that connects a compact Specialized Translation Model (STM) and a general-purpose Large Language Model (LLM) as one unified translation engine. |
| Approach: | They propose a collaborative framework that connects a Specialized Translation Model (STM) and a general-purpose Large Language Model (LLM) as one unified translation engine. |
| Outcome: | The proposed framework outperforms both LLMs and supervised models in high-resource or challenging low-resourced settings. |
Copied to clipboard
| Challenge: | Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. |
| Approach: | They propose a Monte Carlo Tree Search (MCTS) based Prompt Auto-generation (MPA) method to generate adversarial suffixes for valid jailbreak attacks. |
| Outcome: | The proposed method outperforms existing methods on open-source and closed-source models and shows that it can generate harmful responses. |
Copied to clipboard
| Challenge: | a new approach to customer support is proposed to integrate large language models with a framework designed to navigate the complexities of Airbnb customer support operations. |
| Approach: | They propose a method for integrating Large Language Models with a framework designed to navigate the complexities of Airbnb customer support operations. |
| Outcome: | The proposed approach is cost-effective and improves customer support performance . it also allows human agents to focus on more complex issues, the authors show . |
Copied to clipboard
| Challenge: | Existing methods for estimating uncertainty in large language models (LLMs) focus on final-step outputs, which fail to account for cumulative uncertainty over multi-step decision-making process and dynamic interactions between agents and their environments. |
| Approach: | They propose a framework that propagates uncertainty through each step of an LLM-based agent’s reasoning process. |
| Outcome: | Extensive experiments on benchmark datasets show that the proposed framework outperforms state-of-the-art methods by 20%. |
Copied to clipboard
| Challenge: | Legal question answering (LQA) aims to bridge the gap between limited availability of legal professionals and the extensive volume of legal issues. |
| Approach: | They propose a legal knowledge retriever and a hierarchical legal knowledge integration framework to address multiple user-specific circumstances. |
| Outcome: | The proposed framework outperforms baselines on the legal community question-answering dataset. |
Copied to clipboard
| Challenge: | Existing models for language analysis are inadequate for specialized domains like psychology. |
| Approach: | They have enriched a Chinese social media database with psychological lexicons to enhance its applicability to psychological text analysis. |
| Outcome: | The proposed model performed better on six public datasets and provided relevant predictions given the masked sentences. |
Copied to clipboard
| Challenge: | Existing methods that rely on limited demos and out-of-demonstration (OOD) queries fail when faced with out- of-demotion queries. |
| Approach: | They propose a query-aware prompting method that elicits the inherent generalizability of large language models by query-based demo generation. |
| Outcome: | The proposed method outperforms state-of-the-art methods in the OOD setting and two public math benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable linguistic comprehension and generation capability, but when applied to specialized industries, they face challenges such as hallucination, insufficient domain knowledge, and failing to incorporate the latest domain knowledge. |
| Approach: | They propose a paradigm for fine-tuning LLMs that effectively injects domain-specific knowledge while safeguarding inference-time data privacy. |
| Outcome: | The proposed model protects private data while enhancing the model's knowledge. |
Copied to clipboard
| Challenge: | Existing studies show that pretraining sequence composition strategy can lead to distracting information from previous documents. |
| Approach: | They propose to use a sequence construction method to concatenate documents into fixed-length sequences to compute the likelihood of each token given its context. |
| Outcome: | The proposed method can improve in-context learning, knowledge memorisation and context utilisation without sacrificing efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit potential artificial generic intelligence, however, their usage is costly with high response latency. |
| Approach: | They develop a dynamic contextual-bandit-based routing system for query-LLM assignment that leverages query tags to enhance query embeddings. |
| Outcome: | The proposed model maximizes response quality and minimizes cost and latency. |
Copied to clipboard
| Challenge: | Existing table benchmarks lack the capacity to adequately assess the practical application of table reasoning in industrial applications. |
| Approach: | They propose a bilingual table-to-report task and a table-based benchmark to assess the quality of table reasoning. |
| Outcome: | The proposed task is based on a bilingual benchmark with 457 industrial tables and evaluation criteria to measure the quality of report generation. |
Copied to clipboard
| Challenge: | Existing methods for updating large language models are inefficient in multi-client scenarios . Existing approaches assume a single-user setting and are ineffective in multiclient scenarios. |
| Approach: | They propose a new task that enables multiple clients to perform LEKE while preserving privacy and reducing computational overhead. |
| Outcome: | The proposed framework outperforms existing LEKE frameworks on two benchmark datasets and retains 96% of performance. |
Copied to clipboard
| Challenge: | Quantization is a viable solution for pre-trained language models, but most existing methods are task-specific and require customized training and quantization with a large number of trainable parameters. |
| Approach: | They propose a "quantize before fine-tuning" framework that allows for quantization with a large number of trainable parameters on each individual task. |
| Outcome: | The proposed framework is compatible with quantization-aware training and post-training quantization and corrects quantization errors. |
Copied to clipboard
| Challenge: | Existing approaches to multi-turn response generation for open-domain dialogues have a complexity problem . auxiliary tasks that relate to context understanding can guide the learning of the generation model . |
| Approach: | They propose a multi-turn response generation model that has a simple structure yet can effectively leverage conversation contexts for response generation. |
| Outcome: | The proposed model outperforms state-of-the-art models in response quality and human judgment . it also enjoys a faster decoding process . |
Copied to clipboard
| Challenge: | Existing methods for short TST are difficult to implement and can cause content degradation. |
| Approach: | They propose a method to vary the style polarity of text while preserving semantic content. |
| Outcome: | The proposed method improves over baselines and is highly efficient. |
Copied to clipboard
| Challenge: | Existing methods for parameter-efficient finetuning (PEFT) are limited and only finetune a small number of parameters using limited instruction data. |
| Approach: | They propose a method that inserts an adaptation layer into an LLM to integrate embeddings of entities appearing in the input text. |
| Outcome: | The proposed method can activate parameterized knowledge in an LLM without changing its parameters or input prompts. |
Copied to clipboard
| Challenge: | Multimodal embedding models encode multimedia inputs into latent vector representations. |
| Approach: | They propose to synthesize multimodal multilingual data using a multimodal large language model . they identify three criteria for high-quality synthetic multimodal data . |
| Outcome: | The proposed model outperforms existing models on the MMEB Benchmark and the XTD benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to incorporate bilingual dictionaries into Neural Machine Translation (NMT) models have been criticized for lack of integration of bilingual lexical information into the neural architecture. |
| Approach: | They propose a neural architecture to incorporate bilingual dictionaries into Neural Machine Translation models by introducing three new components: Pointer, Disambiguator, and Copier. |
| Outcome: | The proposed method achieves the following merits inherently compared with previous efforts: (1) Pointer leverages the semantic information from bilingual dictionaries, for the first time, to better locate source words whose translation in dictionary can potentially be used; (2) Disambiguator synthesizes contextual information from source view and target view, both of which contribute to distinguishing translation of a specific source word from multiple candidates in dicaries; (3) Copier systematically connects Pointer and Disambiguators based on a hierarchical |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) fail under RDEI, leading to disrupted structure and evidence-unsupported hallucinations. |
| Approach: | They propose a backbone-agnostic, evidence-driven pipeline that treats off-the-shelf MLLMs as interchangeable components to improve stem consistency and figure consistency. |
| Outcome: | The proposed pipeline improves stem consistency by 1.01-3.18%, figure consistency by 0.50-49.16%, and refusal F1 by 1.06-10.88% across question types. |
Copied to clipboard
| Challenge: | Sentence scoring and sentence selection are two main steps in extractive document summarization systems. |
| Approach: | They propose an end-to-end neural network framework for extractive document summarization by jointly learning to score and select sentences. |
| Outcome: | The proposed framework outperforms the state-of-the-art summarization models on the CNN/Daily Mail dataset. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks with limited references may not accurately reflect the quality of the model’s hypotheses. |
| Approach: | They propose a method to enrich evaluation benchmarks by diversifying the expression of a single reference into multiple high-quality ones to cover the semantic space of the reference sentence as much as possible. |
| Outcome: | The proposed method can enhance evaluation benchmarks by diversifying the expression of reference into multiple high-quality ones to cover the semantic space of the reference sentence as much as possible. |
Copied to clipboard
| Challenge: | a tokeniser that splits "2025-03-14" into "20", "25", "-0", "3", "-1", "4" obscures temporal cues and obscures structure . excessive fragmentation correlates with accuracy drops of up to 10 points on uncommon dates . |
| Approach: | They propose a date fragmentation ratio measure that measures how faithfully a tokeniser preserves multi-digit date components. |
| Outcome: | The proposed method shows that excessive fragmentation correlates with accuracy drops of up to 10 points on uncommon dates like historical and futuristic dates. |
Copied to clipboard
| Challenge: | Existing multilingual evaluation benchmarks neglect cultural nuances and lack language coverage in subjective tasks. |
| Approach: | They propose a framework that categorizes evaluation tasks into three cultural layers and nine cognitive sub-layers. |
| Outcome: | The proposed framework surpasses prior coverage by up to 111% on 20+ LLMs. |
Copied to clipboard
| Challenge: | DA-Code is a code generation benchmark designed to assess LLMs on agent-based data science tasks. |
| Approach: | They propose a code generation benchmark specifically designed for LLMs on agent-based data science tasks. |
| Outcome: | The benchmark performs better than existing frameworks, but lacks accuracy . it is based on real-world data, and includes examples that cover a wide range of tasks . |
Copied to clipboard
| Challenge: | Large language models (LLMs) lack large-scale, systematically constructed benchmarks for evaluating their alignment with real-world social attitudes. |
| Approach: | They propose a benchmark to assess LLMs' alignment with real-world social attitudes . they find LLM models achieve only 30–40% accuracy when simulating individuals . |
| Outcome: | The proposed benchmark shows that LLMs achieve only 30% accuracy when simulating individuals in complex survey scenarios. |
Copied to clipboard
| Challenge: | Identifying and modeling puns is challenging as they involve implicit semantic or phonological tricks. |
| Approach: | They propose a method to detect puns in a sentence and then locate them in it . they propose to capture phonetic associations between the context and phonetic symbols . |
| Outcome: | The proposed method outperforms state-of-the-art methods in pun detection and location tasks. |
Copied to clipboard
| Challenge: | a new generation of (M)LLMs is enabling the creation of superintelligent AI assistants . OS Agents can complete tasks autonomously and have the potential to significantly enhance the lives of billions of users worldwide. |
| Approach: | They propose to build OS Agents that operate within operating systems' GUIs and GUIs . they examine evaluation metrics and benchmarks to identify promising directions . |
| Outcome: | The proposed agents are based on operating systems (OS) and operating systems frameworks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have sparked interest in collaborative LLM agents. |
| Approach: | They propose to integrate various ordinal preferential voting mechanisms into LLMs to improve reasoning capabilities and robustness. |
| Outcome: | The proposed method improves reasoning capabilities and robustness of leading LLMs without complex system designs. |
Copied to clipboard
| Challenge: | PoDA pre-trains encoders and decoders by denoising noise-corrupted text . Unlike encoder-only or decode-only methods, it can be used for text generation tasks without using any task-specific techniques. |
| Approach: | They propose a sequence-to-sequence (seq2sequ) pre-training method PoDA which denoises autoencoders by denoising noise-corrupted text. |
| Outcome: | The proposed method improves model performance over strong baselines without using any task-specific techniques and significantly speed up convergence. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on well-structured tables and fail to reflect irregular structures and complex reasoning commonly encountered in real-world scenarios. |
| Approach: | They propose a benchmark to evaluate TableQA under complex reasoning and irregular table conditions. |
| Outcome: | The proposed framework improves generalization and realism of large language models under complex and irregular table conditions. |
Copied to clipboard
| Challenge: | Existing GUI agents for real desktop environments require large amounts of high-quality interaction data, but collecting human demonstrations is expensive. |
| Approach: | They propose a framework that bootstraps scalable desktop supervision from seed demonstrations. |
| Outcome: | Experiments on standard desktop benchmarks show that the framework improves on zero-shot agents and representative synthesis baselines. |
Copied to clipboard
| Challenge: | Large Language Model (LLM)-based optimization has shown promise for autonomous problem solving, but most approaches cast LLMs as passive constraint checkers rather than proactive strategy designers. |
| Approach: | They propose an end-to-end Automated Constraint Optimization method that tightly couples operations-research principles of constraint relaxation with LLM reasoning. |
| Outcome: | Extensive experiments on three challenging COP benchmarks validate AutoCO’s consistent effectiveness and superior performance, especially in hard regimes where current methods degrade. |
Copied to clipboard
| Challenge: | Experimental results prove that language models can learn from human feedback better, irrespective of sequence length . emergence of length bias often induces the model to favor longer outputs . |
| Approach: | They propose to separate reward modeling from the influence of sequence length by using the Product-of-Experts technique. |
| Outcome: | The proposed approach shows that language models perform better regardless of sequence length . the main expert is focused on understanding human intents, while the biased expert targets the identification and capture of length bias. |
Copied to clipboard
| Challenge: | Existing urban itinerary planning studies focus on traditional tourism, but they lack the precision and accuracy needed to create a personalized itinerary. |
| Approach: | They propose an open-domain urban itinerary planning system that integrates spatial optimization with large language models to provide customized urban itineraries based on user needs. |
| Outcome: | The proposed system can generate personalized urban itineraries based on user needs and scale with existing methods. |
Copied to clipboard
| Challenge: | Existing evaluation metrics are not capable of evaluating text quality. |
| Approach: | They propose a metric that compares system output against reference texts based on semantics rather than surface forms. |
| Outcome: | The proposed metric shows a high correlation with human judgment of text quality on a number of text generation tasks. |
Copied to clipboard
| Challenge: | a new method for generating chart annotations is proposed to improve visual reasoning in multimodal large language models. |
| Approach: | They propose a code-as-intermediary translation method for distilling visual reasoning abilities from LLMs to MLLMs. |
| Outcome: | The proposed method is cost-effective, efficient and scalable. |
Copied to clipboard
| Challenge: | Entity linking is a task of assigning entity mentions to referent entities in a knowledge base. |
| Approach: | They propose to use ultra-fine-grained type information to improve the generalization ability of EL models by utilizing a low-level task to extract ultra-finish entity type information. |
| Outcome: | The proposed model achieves state-of-the-art in the zero-shot entity linking task . |
Copied to clipboard
| Challenge: | Existing work exploits easily accessible co-occurrence information of events to learn event representations. |
| Approach: | They propose a weakly supervised contrastive learning method and a prototype-based clustering method for event representation learning. |
| Outcome: | The proposed framework outperforms baselines on Hard Similarity and Transitive Sentence Similarity tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of tasks, including mathematical problem-solving. |
| Approach: | They propose a framework that connects the subgoal breakdown process and the probability of solving problems by identifying better subgoals with theoretical guarantees. |
| Outcome: | The proposed framework outperforms existing methods on two benchmarks, GSM8K and MATH, highlighting the potential of SEGO in AI-driven mathematical problem-solving. |
Copied to clipboard
| Challenge: | Existing controllable Text-to-Speech methods limited to inter-utterance-level control . utterance expressiveness remains a challenge in building human-like TTS synthesis systems . |
| Approach: | They propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. |
| Outcome: | The proposed framework achieves state-of-the-art intra-utterance consistency while maintaining baseline-level speech quality. |
Copied to clipboard
| Challenge: | Existing line-based chunking heuristics often break semantic structures, splitting functions or merging unrelated code. |
| Approach: | They propose a structure-aware method that breaks large AST nodes into smaller chunks . this method generates self-contained, semantically coherent units across programming languages . |
| Outcome: | The proposed method boosts Recall@5 by 4.3 points on RepoEval retrieval and Pass@1 by 2.67 points on SWE-bench generation. |
Copied to clipboard
| Challenge: | Existing knowledge evolution benchmarks are static and fail to capture the evolving nature of LLMs and knowledge. |
| Approach: | They propose an evolving dataset that categorizes information into stable, evolved, and uncharted states. |
| Outcome: | The proposed dataset is auto-updatable and enables evaluation of continuously changing knowledge and newly released LLMs. |
Copied to clipboard
| Challenge: | Existing VSD work focuses on skewed spatial understanding of target objects . Existing work merely models the 2D geometrical vision features . |
| Approach: | They propose to incorporate 3D scene features into visual spatial description tasks by sampling topologically-diverse subgraphs from Go3D-S2G. |
| Outcome: | The proposed framework outperforms baselines on two VSD datasets and produces more spatially-diversified generation. |
Copied to clipboard
| Challenge: | Existing methods for enhancing large language models (LLMs) lack explicit mechanisms for guiding diverse exploration and instead prioritize efficiency and performance over diversity. |
| Approach: | They propose a reinforcement learning-based framework that decomposes the generation process into explicitly planned intermediate steps and introduces divergence at the planning phase based on diversity variation. |
| Outcome: | The proposed method significantly outperforms existing baselines on creative writing benchmarks on a semi-structured long chain-of-thought (CoT) it introduces divergence at the planning phase based on diversity variation, alongside a group-aware diversity reward to encourage distinct trajectories. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases. |
| Approach: | They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process. |
| Outcome: | The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations. |
Copied to clipboard
| Challenge: | Existing approaches to textual robustness evaluation focus on slightly modifying the input data, which maintains the original meaning and results in a different prediction. |
| Approach: | They propose a multilingual robustness evaluation toolkit for NLP that integrates universal text transformations, task-specific transformations and adversarial attack. |
| Outcome: | The toolkit includes universal text transformation, task-specific transformation, adversarial attack, subpopulation, and their combinations to provide comprehensive robustness analyses. |
Copied to clipboard
| Challenge: | Existing models for KG-to-text generation are based on pretrained language models. |
| Approach: | They propose to automatically generate a text that describes the facts in knowledge graph (KG) they leverage the excellent capacities of pretrained language models (PLMs) in language understanding and generation. |
| Outcome: | The proposed model outperforms all comparison methods on fully-supervised and fewshot settings. |
Copied to clipboard
| Challenge: | Low-bit floating-point formats like MXFP and NVFP4 offer new opportunities for precision and efficiency. |
| Approach: | They evaluate HiFloat (HiF8 and HiF4), a family of floating-point formats tailored for Ascend NPUs. |
| Outcome: | The proposed formats excel with high-variance data and are compatible with state-of-the-art quantization frameworks. |
Copied to clipboard
| Challenge: | Existing prompt tuning methods for RC are limited by label spaces and rigid prompt restrictions. |
| Approach: | They propose a generative prompt tuning method to reformulate relation classification as an infilling problem by adding cloze-style phrases to masked language modeling problems. |
| Outcome: | The proposed method exploits rich semantics of entity and relation types and can predict label verbalizations with varying lengths at multiple predicted positions. |
Copied to clipboard
| Challenge: | a new framework for training data processing for Chinese medical language models is proposed . experimental results show that the framework significantly improves model accuracy . |
| Approach: | They propose a data processing framework for Chinese medical language models training and deployment . the framework is based on a question-oriented model training strategy and privacy preservation . |
| Outcome: | The proposed framework significantly improves model accuracy and reduces privacy leakage by 27%. |
Copied to clipboard
| Challenge: | Existing safeguards relying on pre-filtering or fine-tuning are costly and diminish overall utility. |
| Approach: | They propose a lightweight method that leverages LVLMs’ inherent multimodal alignment for zero-shot toxic image detection. |
| Outcome: | The proposed method achieves a 66.9% defense success rate with only 3.2% false positive rate and 7.2% overhead. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly being explored for problem-solving tasks . their strategic planning capability is often viewed with skepticism due to their limited planning capabilities. |
| Approach: | They propose a framework that coordinates agent recruitment and communication through LLM specialized MCTS. |
| Outcome: | The proposed framework achieves 76% accuracy on HotpotQA and 80% on WebShop . it relies on extensive sampling simulations to approximate the true reward distribution . |
Copied to clipboard
| Challenge: | Existing methods for detecting malicious queries rely on residual stream activations, resulting in limited discriminative power and unreliable interventions. |
| Approach: | They propose to use feed-forward networks (FFNs) to generate more discriminative signals for intervention, since these activations more explicitly reflect the intent of a query. |
| Outcome: | Experiments show that the proposed approach achieves state-of-the-art defense performance against various jailbreak attacks while maintaining the model's original performance on benign tasks. |
Copied to clipboard
| Challenge: | Existing alignment paradigms for creative writing use static reward signals and supervised data. |
| Approach: | They propose a constraint-aware reward model that synthesizes query-specific criteria to provide fine-grained preference judgments. |
| Outcome: | The proposed framework aligns models with human preferences across content quality and structural paradigms without supervised fine-tuning and ground-truth references. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have highlighted the challenge of handling long-context tasks. |
| Approach: | They propose a chain-of-thought framework that teaches models to generate high-quality reasoning paths for enhanced long-context performance. |
| Outcome: | The proposed framework generalizes across most long-context scenarios and amplifys with increasing context length. |
Copied to clipboard
| Challenge: | Recent studies have shown that cloze-style reading comprehension is a popular task for measuring the progress of natural language understanding. |
| Approach: | They propose a multi-perspective framework which can be seen as joint training of heterogeneous experts and aggregate context information from different perspectives. |
| Outcome: | The proposed framework achieves new state-of-the-art over previous strong baselines on a recently released cloze-test dataset. |
Copied to clipboard
| Challenge: | Recent research has highlighted a significant inefficiency associated with the slow thinking paradigm . models often overthink simple tasks while underthinking complex challenges . |
| Approach: | They propose a framework for adaptive reasoning preference control that dynamically adjusts reflection preferences based on group-level distributional statistics of reasoning length and reflection intensity. |
| Outcome: | The proposed framework reduces average response length by 17.1%-21.4% while improving performance by 6.12-6.59 points under 32K token budgets. |
Copied to clipboard
| Challenge: | Existing methods that confuse tool utilization with knowledge reasoning harm readability and give rise to tool invocation hallucinations. |
| Approach: | They propose to decouple LLM from tool invocation tasks by establishing a memory module with explicit descriptions of query statements and a query memory module to facilitate the KGQA process. |
| Outcome: | The proposed method achieves state-of-the-art on WebQSP and CWQ benchmarks. |
Copied to clipboard
| Challenge: | Creating spreadsheet formulas remains a tedious and error-prone task for many end-users . a novel task is proposed to generate spreadsheet formulae from a user's NL query . |
| Approach: | They propose a task to generate formulas that are grounded on a spreadsheet table given a Natural Language query as input. |
| Outcome: | The proposed task generates formulas that are grounded on a spreadsheet table, given a natural language query as input. |
Copied to clipboard
| Challenge: | Existing approaches to training agents for visual-language models trap them in local optima, hindering exploration and error correction with the environment. |
| Approach: | They propose a hierarchical training recipe that bridges atomic action execution and strategic task completion. |
| Outcome: | The proposed training recipe bridges atomic action execution and strategic task completion. |