Papers by Liang Wang
Copied to clipboard
| Challenge: | Hallucination is a significant barrier to the effective application of Large Language Models (LLMs). |
| Approach: | They propose an Attention-Guided SElf-Reflection approach for hallucination detection in Large Language Models. |
| Outcome: | The proposed method significantly outperforms existing methods in zero-shot hallucination detection on four widely-used LLMs across three different halluciation benchmarks. |
Copied to clipboard
| Challenge: | Existing unlearning paradigms are mired in vague forgetting boundaries, erasing knowledge indiscriminately. |
| Approach: | They propose a benchmark to evaluate if unlearning erases essential knowledge . they propose 'knowUnDo' which uses copyrighted content and privacy domains . |
| Outcome: | The proposed method is superior to existing methods in both precise knowledge unlearning and general knowledge retaining of LLMs. |
Copied to clipboard
| Challenge: | None Large language models (LLMs) are emerging as a key tool for automated programming. |
| Approach: | They compare performance of None Large language models with language understanding models on functional programming and object-oriented programming benchmarks. |
| Outcome: | The models perform relatively well on functional programming (FP) and object-oriented programming (OOP) benchmarks, while exhibiting poor performance on OOP benchmarks. |
Copied to clipboard
| Challenge: | Recent reinforcement learning approaches have advanced radiology report generation (RRG) however, there are two limitations: report-level rewards offer limited evidence-grounded guidance for clinical faithfulness . |
| Approach: | They propose a method that uses group-wise evidence-aware alignment rewards and self-correcting preference learning to build a reliable, disease-agnostic preference dataset without human supervision. |
| Outcome: | ESC-RL promotes clinically faithful, disease-aligned reward and supports continual self-improvement during training. |
Copied to clipboard
| Challenge: | Existing benchmarks for role-playing agents only evaluate surface-level fidelity and provide limited insight into decision making under role–alignment value conflicts. |
| Approach: | They propose a benchmark to evaluate RPAs under role–alignment value conflicts . they use 8k diverse role profiles and 240k dilemma instances to evaluate role-aware decision making . |
| Outcome: | The proposed benchmark covers 8k diverse role profiles and scenarios and nearly 240k dilemma instances across three difficulty levels and eight role categories. |
Copied to clipboard
| Challenge: | Existing clickbait detection models rely on analyzing the objective semantics of posts or correlating posts with article content only, but fail to identify and exploit the manipulation intention of clickbaiting from a user’s subjective perspective. |
| Approach: | They propose a multiview clickbait detection model to model subjective and objective preferences simultaneously to capture clickbaiting from a user's subjective perspective. |
| Outcome: | The proposed model outperforms state-of-the-art models on two real-world datasets and shows that it integrates subjective and objective preferences simultaneously. |
Copied to clipboard
| Challenge: | Detecting fraudulent online text is essential as they exploit human greed and deceive individuals. |
| Approach: | They propose to build a long-term dataset of Chinese fraudulent texts collected over 12 months. |
| Outcome: | The proposed dataset includes 59,106 entries extracted from billions of web pages and includes large language model-based detectors and pre-trained language model approaches. |
Copied to clipboard
| Challenge: | Existing methods for style transfer are based on an inductive learning approach, which represents the style as embeddings, decoder parameters, or discriminator parameters and directly applies these general rules to the test cases. |
| Approach: | They propose a retrieval-based context-aware style representation that involves top-K relevant sentences in the target style in the transfer process. |
| Outcome: | The proposed method outperforms several strong baselines and is general and effective to the task of unsupervised style transfer. |
Copied to clipboard
| Challenge: | Existing defenses rely on impractical assumptions about trigger settings to mitigate backdoor attacks . a recent study found that small amounts of training data can systematically induce harmful behaviors in large language models. |
| Approach: | They propose a backdoor defense framework that requires no prior knowledge of trigger settings . they use a two-stage process to aggregate backdoor representations and fine-tune recovery . |
| Outcome: | The proposed defense reduces the average Attack Success Rate to 4.41% across multiple benchmarks . the proposed framework generalizes across different types of backdoors, confirming its robustness in practical deployment scenarios. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) focus on replicating human cognition in specific contexts, overlooking the inherently dynamic nature of cognition. |
| Approach: | They propose a task to assess cognitive dynamics of large language models (LLMs) they introduce a benchmark and two evaluation metrics to validate the benchmark and evaluate it through participant surveys. |
| Outcome: | The proposed task overcomes the limitations of existing methods and is available for download. |
Copied to clipboard
| Challenge: | Existing work only encodes entity types and textual context within individual instances, which limits the performance of sentence-level relation extraction (RE). |
| Approach: | They propose a module that aggregates the features from sentences to learn global representations of properties and augments local features within individual sentences. |
| Outcome: | The proposed module can learn global representations of properties from sentences and augment local features within individual sentences. |
Copied to clipboard
| Challenge: | Program-of-Thought (PoT) replaces natural language-based Chain-ofThough (CoT) but introduces more reasoning errors, such as incorrect formulas or flawed logic, compared to CoT. |
| Approach: | They propose a method that integrates CoT and Program-of-Thought to achieve more accurate reasoning and reinforcement learning. |
| Outcome: | The proposed method achieves an average improvement of 6.5% on the Llama-Base model and 4.3% on the Mistral-Bass model across 8 mathematical calculation datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have recently achieved remarkable progress on complex reasoning tasks by leveraging extended Chain-of-Thought (CoT) techniques. |
| Approach: | They propose a method that uses Extended Chain-of-Thought (EFT) to reduce the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
| Outcome: | The proposed method reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have advanced beyond traditional Large Language Models, yet they pose heightened safety risks. |
| Approach: | They propose a first jailbreak attack targeting Large Reasoning Models . they exploit a Chaos Machine component to transform attack prompts with diverse one-to-one mappings based on the reasoning chain . |
| Outcome: | The proposed attack exploits the unique vulnerabilities of LRMs by integrating a Chaos Machine. success rates of the mousetrap attack are as high as 96%, 86% and 98% respectively. |
Copied to clipboard
| Challenge: | Current multimodal large language models (MLLMs) show limited understanding of dental images. |
| Approach: | They propose a dental-specialized multimodal large language model trained via staged multimodal alignment and reinforcement learning. |
| Outcome: | The proposed model outperforms state-of-the-art models on disease classification and dental VQA tasks. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has improved text- and vision-language models, but its application in SDMs is hindered. |
| Approach: | They propose a dual-axis Generative Reward Model that provides semantic quality and interaction timing for SDMs. |
| Outcome: | The proposed model achieves state-of-the-art performance on interaction-quality assessment across a wide spectrum of datasets. |
Copied to clipboard
| Challenge: | Existing methods for unlearning large language models struggle to balance effective forgetting with maintaining model utility. |
| Approach: | They propose a human-inspired unlearning framework that simulates forgetting on fuzzy data and represents them in hyperbolic and Euclidean spaces. |
| Outcome: | The proposed framework is able to forget sensitive content while maintaining the model’s language understanding, fluency, and benchmark performance. |
Copied to clipboard
| Challenge: | Previous work shows that large language models generate hallucinations, yet the origins and mechanisms of these signals remain unclear. |
| Approach: | They propose to validate and disentangle two different pathways for truthfulness cues . they also propose to use the same mechanism to derive self-contained evidence from the generated answer . |
| Outcome: | The proposed applications improve hallucination detection performance by integrating two different inputs. |
Copied to clipboard
| Challenge: | Multi-domain learning is a good solution for solving domain tasks but it requires retraining when adding a new domain. |
| Approach: | They propose to exploit unlabeled data from the same distributions of the older domains to avoid catastrophic forgetting. |
| Outcome: | The proposed framework exploits unlabeled data from the same distributions of the older domains to avoid catastrophic forgetting. |
Copied to clipboard
| Challenge: | Existing methods for storing key-value caches during long-horizon rollouts cause performance collapses. |
| Approach: | They propose a new training paradigm that empowers stable RL training under sparse rollouts. |
| Outcome: | The proposed model reduces rollout overhead while maintaining the performance. |
Copied to clipboard
| Challenge: | Pairwise data constructed from weakly supervised signals is widely used for training deep learning models. |
| Approach: | They propose two methods to refine pairwise data that are aimed to obtain subsets that are more useful as learning examples. |
| Outcome: | The proposed methods achieve most machine translation gains in the first iteration, but following iterations further improve its intrinsic evaluation. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks that use large language models as referees are insufficient for accurately assessing their alignment with human intent. |
| Approach: | They propose a calibration framework to address positional bias in large language models as evaluators by manually annotating the “win/tie/lose” outcomes of responses from ChatGPT and Vicuna-13B in the Vicun A Benchmark’s question prompt. |
| Outcome: | The proposed framework alleviates evaluation bias, resulting in closer alignment with human judgments. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. |
| Approach: | They propose a multilingual chart question answering benchmark that enables efficient multilingual generation via data translation and code reuse. |
| Outcome: | The proposed benchmark systematically evaluates multilingual chart understanding on state-of-the-art LVLMs and shows a significant performance gap between English and other languages. |
Copied to clipboard
| Challenge: | Various watermarking methods have been proposed to identify LLM-generated texts . lack of unified evaluation platforms has left many critical questions unanswered . |
| Approach: | They systematize existing LLM watermarkers and watermark removal attacks and develop a unified platform that integrates them. |
| Outcome: | The proposed systematizes existing LLM watermarkers and watermark removal attacks, mapping out their design spaces. |
Copied to clipboard
| Challenge: | Large language models (LLMs) like ChatGPT are only accessible through restricted APIs, which creates barriers to new research and advancements in the field. |
| Approach: | They propose a framework to enhance and regulate the translation abilities during chat . they reformulate translation data into the instruction-following style and introduce a "Hint" field . |
| Outcome: | The proposed framework enhances and regulates the translation abilities during chat . it reformulates translation data into the instruction-following style and introduces a "Hint" field . |
Copied to clipboard
| Challenge: | Object hallucination has been an Achilles’ heel which hinders the broader applications of large vision-language models (LVLMs). |
| Approach: | They propose a logical closed loop-based framework for Object Hallucination Detection and Mitigation that uses logical consistency probing to raise questions with logical correlations to determine hallucinations. |
| Outcome: | The proposed method can be applied to all existing LVLMs and is effective and general. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis models are susceptible to learning spurious correlations between words . a recent study shows that feature engineering is time-consuming and costly . |
| Approach: | They propose to use a template to prompt LLMs to generate an appropriate explanation for the sentiment polarity of each aspect to reduce spurious correlations. |
| Outcome: | The proposed methods improve ABSA models and their generalization ability. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated good performance in many reasoning tasks, but struggle with some more complex reasoning tasks including logical reasoning. |
| Approach: | They propose five concrete tasks from three cognitive dimensions of WHAT, WHY, and HOW to evaluate LLMs’ capability of logical fallacy understanding. |
| Outcome: | The proposed dataset can be used to evaluate LLMs’ LFU capability and to fine-tune LLM models to obtain significantly enhanced performance on logical reasoning. |
Copied to clipboard
| Challenge: | Conventional speculative decoding methods use a predefined length policy for proposing drafts, but the reality deviates from this assumption. |
| Approach: | They propose a self-verification length policy that adaptively determines the lengths of draft sequences by referring to the draft entropy. |
| Outcome: | The proposed method achieves 17% speedup on MT-Bench and 22% speedup in long-form reasoning. |
Copied to clipboard
| Challenge: | Existing methods for text-to-SQL generation are prone to hallucinations and grounding . authors present a novel reasoning paradigm that transforms text- to-Sql from unverifiable textual rationales into step-wise executable semantics. |
| Approach: | They propose a reasoning paradigm that transforms text-to-SQL from unverifiable textual rationales into step-wise executable semantics. |
| Outcome: | The proposed reasoning paradigm transforms text-to-SQL from unverifiable textual rationales into step-wise executable semantics. |
Copied to clipboard
| Challenge: | Existing speech codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks. |
| Approach: | They propose a neural speech codec with semantic-acoustic dual-stream quantization that disentangles semantic and acousian modeling into two dedicated streams. |
| Outcome: | The proposed codec outperforms state-of-the-art speech tokenizers in auto-propagating text-to-speech models. |
Copied to clipboard
| Challenge: | MCTS methods retain only the single highest-reward trajectory, discarding comparative signals present in the many explored paths. |
| Approach: | They propose a framework that transforms supervision extraction into a synthesis procedure. |
| Outcome: | The proposed framework matches or exceeds baselines on 60K CRPS-synthesized examples on out-of-domain benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to multi-turn dialogues lack contextual consistency and dependencies, and models struggle to maintain factual faithfulness as interaction turns increase. |
| Approach: | They propose an adaptive context refactoring framework that monitors and reshapes the interaction history to mitigate contextual inertia and state drift. |
| Outcome: | The proposed model outperforms baselines while reducing token consumption. |
Copied to clipboard
| Challenge: | Multi-hop Question Answering (MHQA) is a challenging task that requires models to answer multiple questions with multiple passages. |
| Approach: | They propose a self-guided prompting finite state machine to improve multi-hop reasoning abilities by iterating over multiple questions and correcting itself to improve accuracy. |
| Outcome: | The proposed approach outperforms baselines on Musique and other datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) produce hallucinated text, compromising their practical utility in professional contexts. |
| Approach: | They have developed an unconstrained hallucination generation evaluation benchmark that contains hallucines generated by large language models with minimal restrictions. |
| Outcome: | The proposed benchmarks are based on a Chinese-language dataset that is lacking in the field. |
Copied to clipboard
| Challenge: | Recent prompt learning-based approaches have shown promising improvements on the ECI task . however, they are subject to the delicate design of multiple prompts and positive correlations between the main task and derivate tasks. |
| Approach: | They propose an event causality identification model that uses contrastive learning to enhance both positive and negative demonstrations. |
| Outcome: | The proposed model improves on the event-related causality identification task . it uses contrastive learning to enhance both positive and negative demonstrations . |
Copied to clipboard
| Challenge: | Existing methods to identify unseen multimodal entities struggle with limited knowledge and generalization. |
| Approach: | They propose a framework that leverages the strengths of small fine-tuned models and MLLMs to generate unambiguous predictions. |
| Outcome: | Extensive experiments show that the proposed framework retains the in-domain knowledge of small models while utilizing the capabilities of MLLMs to handle unseen entities. |
Copied to clipboard
| Challenge: | Recent work on LLMs has focused on fine-grained skill decomposition and consistency probing at the propositional level. |
| Approach: | They propose a benchmark evaluating immediate inference that evaluates elemental operations over categorical propositions and proposes a model that uses immediate inferential reasoning. |
| Outcome: | The proposed benchmark demonstrates that models lack robust operator grounding, oscillating between structural reasoning and surface pattern matching, inconsistent handling of quantifiers and negation. |
Copied to clipboard
| Challenge: | generative search engines rely on in-line citations as the key gateway to original webpages . a recent study shows that LLMs tend to cite left-leaning sources at higher rates compared to traditional retrieval systems . |
| Approach: | They construct a dataset of news articles labeled with left- or right-leaning stances . they find that LLMs tend to cite left-leansing sources at higher rates than traditional retrieval systems . |
| Outcome: | The proposed dataset shows that LLMs tend to cite left-leaning sources at higher rates than traditional retrieval systems. |
Copied to clipboard
| Challenge: | Named Entity Recognition models are feature-engineering and machine learning based. |
| Approach: | They propose a new NER learning framework that uses entity mentions to improve model performance. |
| Outcome: | The proposed model achieves better performance on OOV entities on various settings and datasets. |
Copied to clipboard
| Challenge: | Existing multi-agent systems lack agent coordination and rely on predefined procedures . existing systems lack adaptive task coordination when task is big and complex . |
| Approach: | They propose a large-scale autonomous LLM-based multi-agent system that generates agents based on task complexity and enables dynamic task decomposition, parallel execution, efficient communication and comprehensive system monitoring. |
| Outcome: | The proposed system outperforms existing systems in task completion efficiency and scalability. |
Copied to clipboard
| Challenge: | Natural Language Inference (NLI) is a foundational understanding task in language understanding. |
| Approach: | They propose a framework to construct counterfactual reasoning data and fine-tune LLMs to reduce attestation bias. |
| Outcome: | The proposed framework reduces hallucinations from attestation bias on original and bias-neutralized datasets while keeping hypotheses unchanged. |
Copied to clipboard
| Challenge: | Existing knowledge distillation techniques for neural machine translation lack special treatment on the top-1 information, which is limiting the potential of KD. |
| Approach: | They propose a method to distill knowledge from top-1 predictions of teachers and a technique to infuse more additional knowledge by distilling on the data without ground-truth targets. |
| Outcome: | The proposed method outperforms the vanilla word-level KD and outperfies the existing methods on three different students with different capacity gaps. |
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is a task arousing increasing interest in many fields. |
| Approach: | They propose a novel GNN-based ERC model that captures speaker and position information. |
| Outcome: | The proposed model captures speaker and position-aware conversation structure information. |
Copied to clipboard
| Challenge: | comparative method allows linguists to infer protoforms from their reflexes based on sound change . authors argue that this approach ignores one of the most important aspects of the comparative approach . |
| Approach: | They propose a comparative method that allows linguists to infer protoforms from their reflexes . they propose to use a system where candidate protoform from a reconstruction model are reranked by a reflex prediction model. |
| Outcome: | The comparative method surpasses state-of-the-art methods on Chinese and Romance datasets. |
Copied to clipboard
| Challenge: | Existing evaluation methods based on large language models (LLMs) are expensive and lack expertise due to limitations in human expertise. |
| Approach: | They propose an open-source automatic evaluation model with 13B parameters specifically engineered to measure the question-answering proficiency of medical LLMs. |
| Outcome: | The proposed model surpasses baselines in terms of correlation with human judgments. |
Copied to clipboard
| Challenge: | Non-autoregressive translation models are weak at learning high-mode knowledge, argues a new study . despite the improved learning difficulty, there are still complicated word orders and structures in the synthetic sentences, making the NAT performance sub-optimal. |
| Approach: | They propose to train non-autoregressive translation models to learn fine-grained lower-mode knowledge . they break down sentence-level examples into three types and increase granularities . |
| Outcome: | The proposed method improves phrase translation accuracy and model reordering ability against strong NAT baselines. |
Copied to clipboard
| Challenge: | Relation extraction (RE) is an essential topic in natural language processing and has attracted extensive attention. |
| Approach: | They propose a case-oriented construction framework to build a hard case relation extraction dataset with 65,225 relational facts annotated from 9,231 documents. |
| Outcome: | The proposed model achieves a high 96% F1 score on data quality and is far lower than humans. |
Copied to clipboard
| Challenge: | SDiaReward is an end-to-end spoken dialogue system that integrates paralinguistic nuances and spontaneous nature of human conversation. |
| Approach: | They propose an end-to-end multi-turn reward model trained on SDiaReward-Dataset . it is a collection of episode-level preference pairs targeting modality and colloquiality gaps . |
| Outcome: | The proposed model outperforms general-purpose audio LLMs in episode-level evaluation. |
Copied to clipboard
| Challenge: | a paper focuses on automatically generating the text of an ad to capture user interest for achieving higher click-through rate. |
| Approach: | They propose a CTR-driven advertising text generation approach to generate ad texts based on user reviews. |
| Outcome: | The proposed approach outperforms existing approaches on industrial datasets and on large-scale unpaired reviews. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for memory-augmented dialogue systems lack practical value . current evaluation methods only consider passive memory retrieval while ignoring diverse memory recall with rich triggering factors. |
| Approach: | They propose to use long-term memory to create human-like dialogues using chatbots. |
| Outcome: | The proposed benchmark covers memory retrieval and memory recognition tasks with both passive and proactive memory recall data. |
Copied to clipboard
| Challenge: | Existing approaches to reduce memory footprint of long-context LLMs rely on RoPE-induced oscillations. |
| Approach: | They propose a frequency-domain framework that converts RoPE-induced oscillations into sparse spectral representations. |
| Outcome: | The proposed framework achieves efficient compression with performance comparable to FP16 benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Emotion Recognition in Conversations models struggle due to lack of Common Sense Knowledge (CSK). |
| Approach: | They propose a multimodal approach to integrate multiple knowledge into the edge representations by integrating textual and visual CSK. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two popular datasets. |
Copied to clipboard
| Challenge: | Large Language Models exhibit a level of intelligence that is both impressive and everevolving, but their ability to refuse generating unsafe content is a double-edged sword. |
| Approach: | They propose a method to tackle a refusal position bias within safety tuning data that compromises the models’ ability to appropriately refuse generating unsafe content. |
| Outcome: | The proposed method significantly improves model safety without compromising performance and surpasses baseline methods in defending against attacks. |
Copied to clipboard
| Challenge: | Existing methods that render words or characters into images separately, but instead use text's visual features as input, we use 3-dimensional convolutions to learn distributed sentence representation. |
| Approach: | They propose to use text's visual features as input to learn distributed sentence representation using 3-dimensional sentence tensors and multiple 3-dimensional convolutions with different lengths are applied to the sentence . |
| Outcome: | The proposed model performs well on several downstream natural language processing tasks. |
Copied to clipboard
| Challenge: | Existing approaches to text generation often neglect event structures that shape real-world narratives. |
| Approach: | They propose a framework that integrates structured event semantics with iterative retrieval and inference to enhance text generation. |
| Outcome: | Experiments on UltraDomain and MultiHopRAG show that the proposed framework outperforms baseline RAG systems in generation effectiveness, logical consistency, and multi-hop reasoning accuracy. |
Copied to clipboard
| Challenge: | Medical record reviewers must produce consistent, traceable, guideline-compliant outcomes . longcontext inference is expensive and often degrades as inputs grow . |
| Approach: | a new method compiles textual guidelines into a fixed review tree . a cost-aware split-and-prune search is used to update the tree offline . the algorithm produces consistent, traceable, guideline-compliant outcomes . |
| Outcome: | The proposed system outperforms the strongest non-expert baselines by 84.5–92.8 Macro-F1 . it reduces average I/O volume to 74K input+output characters and average latency to 22s . |
Copied to clipboard
| Challenge: | Unsupervised contrastive learning is emerging as a powerful technique for extracting knowledge from unlabeled data. |
| Approach: | They propose a momentum contrastive learning model with negative sample queue for sentence embedding with a simulated model with EMA update mechanism. |
| Outcome: | The proposed model achieves a Spearman’s correlation of 77.27% on the semantic text similarity task and a maximum traceable distance metric. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are shifting the focus from single verifiable tasks toward complex, open-ended real-world scenarios. |
| Approach: | They propose a framework that automatically adjusts reward weights and data importance to synchronize learning intent with data utility for optimal performance. |
| Outcome: | The proposed framework improves model capabilities across all domains and scales. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in various domains but face challenges when applied to data science workflows due to their complex, multi-stage nature. |
| Approach: | They propose a hierarchical graph-based agent that represents complexity and a progressive strategy for step-by-step verification, refinement, and consistent context management. |
| Outcome: | The proposed agent surpasses state-of-the-art baselines on the MATH dataset and performs better on InfiAgent-DABench. |
Copied to clipboard
| Challenge: | Patronizing and condescending language is an essential branch of toxic language . pre-trained language models perform poorly in detecting PCL due to its implicit toxicity traits . |
| Approach: | They propose a novel LLM benchmark for patronizing and condescending language . they use a dataset to analyze the toxicity of patronizing condescending languages . |
| Outcome: | The proposed model can detect patronizing and condescending language (PCL) the model can be used to analyze the toxicity of the language and to improve the detection. |
Copied to clipboard
| Challenge: | Existing models struggle to balance predictive accuracy with human-understandable rationales. |
| Approach: | They propose to enhance LLMs by leveraging rationale distillation and domain knowledge injection for trustworthy multimodal rationale generation. |
| Outcome: | Experiments on real-world medical datasets show that ClinRaGen achieves state-of-the-art performance in disease diagnosis and rationale generation. |
Copied to clipboard
| Challenge: | Existing literature on mechanistic interpretation (MI) treats it as an observational science, leaving practical applications underexplored. |
| Approach: | They propose a survey structured around the pipeline to identify and improve MI models. |
| Outcome: | The proposed framework enables tangible improvements in Alignment, Capability, and Efficiency. |
Copied to clipboard
| Challenge: | Recent advances in NLP have been driven by the development of Large Language Models (LLMs). |
| Approach: | They propose a self-renewal approach to optimize LLM outputs to better align with human preferences without supervised fine-tuning. |
| Outcome: | The proposed approach improves outputs to better align with human preferences across LLMs and tasks without supervised fine-tuning. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) generate coherent reasoning paths before conclusions, but they introduce new vulnerabilities. |
| Approach: | They propose a framework that leverages a weaker but less-aligned model to simulate execution reasoning for initial hijacking attempts and iteratively refines attacks by exploiting reasoning patterns leaked through the target LRM’s refusals. |
| Outcome: | The proposed framework achieves 100% success rate within one or few turns, neutralizing reasoning-based defenses even when evaluated by robustly aligned external models. |
Copied to clipboard
| Challenge: | Existing studies focus on improving MMS models by filtering summary-unrelated visual features with implicit learning or explicitly complex training objectives. |
| Approach: | They propose a multimodal multimodal summarization task that aims to generate summaries in any language with document inputs in any languages and the corresponding image sequence. |
| Outcome: | The proposed task can generate summaries in any language with document inputs in any languages and the corresponding image sequence. |
Copied to clipboard
| Challenge: | Existing MAS initialization methods do not fully account for the collaborative needs of the generated agents in subsequent stages. |
| Approach: | They propose to use a Natural Language to Format mechanism to optimize the structure of agent teams and incorporate a natural language to format mechanism to ensure consistency and standardization. |
| Outcome: | The proposed method outperforms state-of-the-art initialization methods and pre-defined strategies across various frameworks and tasks while reducing token consumption. |
Copied to clipboard
| Challenge: | Existing data synthesis methods suffer from limited diversity and lack precise control over problem difficulty, making them insufficient for efficient training paradigms such as curriculum learning. |
| Approach: | They propose a data synthesis paradigm that generates high-quality, difficulty-controllable mathematical reasoning problems through hybrid and decomposed strategies. |
| Outcome: | The proposed paradigm outperforms existing methods and improves mathematical reasoning abilities. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have demonstrated impressive capabilities in code generation across various domains, but their ability to replicate complex, multi-panel visualizations remains largely unassessed. |
| Approach: | They propose a large-scale benchmark to evaluate chart generation from large- scale raw data and assess iterative code refinement in a multi-turn conversational setting. |
| Outcome: | The new benchmark evaluates 14 leading VLMs on real-world data and shows they struggle with complex plot structures and authentic data. |
Copied to clipboard
| Challenge: | Existing approaches to speed up parallel scaling have relied on similarity-based or confidence-based pruning, but these signals do not reliably indicate trace quality. |
| Approach: | They propose a pruning framework that evaluates reasoning steps using hidden states and dynamically prunes unpromising traces during generation. |
| Outcome: | The proposed framework reduces end-to-end inference latency by 45%–70% on average compared to self-consistency while improving reasoning accuracy. |
Copied to clipboard
| Challenge: | Multi-modal entity alignment (MMEA) is a critical task that aims to identify equivalent entity pairs across multi-modal knowledge graphs (MMKGs). |
| Approach: | They propose a novel MMEA transformer that hierarchically introduces neighbor features, multi-modal attributes, and entity types to enhance alignment task. |
| Outcome: | The proposed transformer hierarchically introduces neighbor features, multi-modal attributes, and entity types to enhance the alignment task. |
Copied to clipboard
| Challenge: | Generative audio modeling has been fragmented into specialized tasks such as text-to-speech (TTS), text- to-music (TTM), and text-ta (TTA) specialized models require reference audio for timbre cloning and strict phoneme alignment, whereas TTA models generate unstructured textures from open-ended captions. |
| Approach: | They propose a unified flow-matching framework capable of synthesizing speech, music, sound effects . they propose 'token injection mechanism' that projects unstructured environmental sounds into structured temporal latent space . |
| Outcome: | The proposed framework achieves state-of-the-art performance in instruction-based TTS and TTM while maintaining competitive fidelity in TTA. |
Copied to clipboard
| Challenge: | Texar is an open-source text generation toolkit that supports a broad set of text generation tasks. |
| Approach: | They introduce Texar, an open-source text generation toolkit that supports text generation tasks. |
| Outcome: | Texar supports machine translation, summarization, dialog, content manipulation, and more. |
Copied to clipboard
| Challenge: | Document images are characterized by higher resolutions, denser content, and more complex structural layouts. |
| Approach: | They propose a 1.2B-parameter document parsing vision-language model that decouples layout analysis from local content recognition. |
| Outcome: | The proposed model surpasses general-purpose and domain-specific models on multiple benchmarks while maintaining significantly lower computational overhead. |
Copied to clipboard
| Challenge: | Event Argument Extraction (EAE) aims to extract arguments for specified events from a text . previous work focused on long-distance dependencies of arguments, modeling co-occurrence relationships . |
| Approach: | They propose a model that takes inductive biases as targets to locate prototypes . they set multiple prototypes to represent each role to capture intra-class differences . |
| Outcome: | The proposed model achieves state-of-the-art on the RAMS and WikiEvents datasets. |
Copied to clipboard
| Challenge: | Existing GUI agents struggle to adapt to dynamic and interconnected nature of real-world digital environments, authors show . |
| Approach: | They propose a benchmark to evaluate the transferability of GUI agents across three key dimensions . transBench includes 15 app categories with diverse functionalities . |
| Outcome: | The proposed benchmark shows that existing GUI agents struggle to adapt to dynamic, interconnected environments. |
Copied to clipboard
| Challenge: | Existing methods for multimodal aspect-based sentiment analysis focus on fusing image regional information and textual words. |
| Approach: | They propose a multimodal aspect-based sentiment analysis method that integrates regional and global image information with global image data. |
| Outcome: | Experiments show that the proposed method outperforms state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for argument quality assessment do not consider multi-perspective evaluation due to subjective nature of arguments. |
| Approach: | They propose a multi-persona framework for argument quality assessment that simulates diverse evaluator perspectives through large language models. |
| Outcome: | The proposed framework outperforms baselines while providing comprehensive multi-perspective rationales on IBM-Rank-30k and IBM-ArgQ-5.3kArgs datasets. |
Copied to clipboard
| Challenge: | sarcasm is a form of irony conveying mockery and contempt . social media has become increasingly popular for identifying sarcasm . |
| Approach: | They develop a method to detect sarcasm from social media using augmented potentials. |
| Outcome: | The proposed method outperforms baselines on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing machine learning models may lead to poor performance in discriminative natural language understanding tasks. |
| Approach: | They propose to use ChatGPT to query large amounts of human-written text to find the answer to a question. |
| Outcome: | The proposed model has a high chance to select labels at earlier positions as the answer. |
Copied to clipboard
| Challenge: | Existing document question answering methods reduce inference costs and input tokens. |
| Approach: | They propose a retrieval-augmented generation method that automatically extracts useful entities and generates summaries from documents. |
| Outcome: | The proposed method surpasses baseline retrieval-augmented generation (RAG) and long-context question answering (LC) methods achieve higher accuracy by processing entire documents, but at the cost of increased computational Corresponding authors. |
Copied to clipboard
| Challenge: | Existing studies on LLM confidence estimations in languages other than English have been limited to English. |
| Approach: | They propose to use question-related language to prompt LLMs to assess their confidence in large language models. |
| Outcome: | The proposed model improves on question-related language prompts for LS tasks, while English exhibits notable linguistic dominance in confidence estimations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can call tools effectively, but they remain brittle in multi-turn execution. |
| Approach: | They propose a framework that converts execution errors into on-policy corrective supervision within the RL training loop. |
| Outcome: | The proposed framework improves the error recovery rate of Qwen3-8B by 5.7% absolute and overall accuracy by 4.0% on BFCL v4 Multi-Turn. |
Copied to clipboard
| Challenge: | Existing methods define important nodes as important and target them for attacks if the model treats nodes’ predictive influence more uniformly . Existing approaches target high predictive influence nodes but are vulnerable to malicious message injection attacks. |
| Approach: | They propose a defense mechanism that encourages the model to learn graph representations where nodes with varying importance have a more uniform influence on predictions. |
| Outcome: | Extensive experiments on the Twitter and Weibo datasets show that similarizing the predictive Influence of nodes with Contrastive Learning significantly enhances resistance against LLM-driven message injection attacks. |
Copied to clipboard
| Challenge: | Existing top-performing methods for Lexical Relation Mining rely on pre-trained language models yet fail to distinguish nuanced lexical relations. |
| Approach: | They propose a framework to leverage structured sememe knowledge to enhance LRC and LE. |
| Outcome: | The proposed method outperforms existing methods on benchmarks and outperformed the LLMs. |
Copied to clipboard
| Challenge: | Existing supervised fine-tuning methods struggle to generalize across document types, leading to poor performance. |
| Approach: | They propose layoutRL, a reinforcement learning framework that optimizes layout understanding through composite rewards integrating normalized edit distance, paragraph count accuracy, and reading order preservation. |
| Outcome: | The proposed model outperforms specialized document parsing systems and general-purpose vision-language models on a broad range of document types, languages, and structural complexities. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have led to an influx of AI-generated content on the internet, transforming corpus of Information Retrieval (IR) systems from human-written to a coexistence with LLM-generated contents. |
| Approach: | They propose a benchmark named Cocktail that compares IR models with LLMs to find relevant documents and passages from a corpus. |
| Outcome: | The proposed benchmark aims to evaluate IR models in the mixed-sourced data landscape of the LLM era. |
Copied to clipboard
| Challenge: | Existing approaches to personalized text generation rely on retrieval-augmented generation and parameter-efficient fine-tuning. |
| Approach: | They propose a training-free framework that disentangles and represents personalized writing style as a vector in LLM’s activation-space. |
| Outcome: | The proposed framework achieves 8% relative improvement in personalized generation while reducing storage requirements by 1700 over PEFT method. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) are evolving rapidly and require data with human supervision to achieve better alignment. |
| Approach: | They introduce VLFeedback, the first large-scale vision-language feedback dataset . they train Silkie, an LVLM fine-tuned via direct preference optimization . |
| Outcome: | The proposed model outperforms its base model in helpfulness, visual faithfulness, and safety metrics and exhibits enhanced resilience against red-teaming attacks. |
Copied to clipboard
| Challenge: | Existing methods for long chain-of-thought (LCoT) are coarse-grained, reward hacking, and poor generalization. |
| Approach: | They propose a Long Chain-of-Thought (LCoT) model that integrates reinforcement learning with verifiable rewards with a process-aware verification approach. |
| Outcome: | The proposed model improves reasoning and code generation tasks while reducing the cost of training and performance bottlenecks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable progress in dialogue and reasoning, but they struggle to solve strictly constrained dialogue tasks. |
| Approach: | They construct a dataset that contains 12,705 high-quality Chinese dialogue instructions from 440 flowcharts containing 5,055 process nodes. |
| Outcome: | The proposed model outperforms GPT-4o models on backward transitions and outperformed GPT-42 models on the same dataset. |
Copied to clipboard
| Challenge: | Existing methods for misinformation detection are limited by domain knowledge and expert experience. |
| Approach: | They propose a Multi-Agent Framework for cross-domain misinformation detection with Automated Decision Rule Optimization (MARO) they first employ multiple expert agents to analyze target-domain news, then introduce a question-reflection mechanism that guides expert agents for higher-quality analysis. |
| Outcome: | The proposed framework improves on a common dataset and shows that iteratively improves over existing methods. |
Copied to clipboard
| Challenge: | Chinese Search Query Spell Correction is a task designed to identify and correct typographical errors within queries. |
| Approach: | They propose a large-scale benchmark specifically developed for Chinese Query Spell Correction. |
| Outcome: | The proposed benchmark covers a broad range of topics, including formal entities, everyday colloquialisms and idiomatic expressions. |
Copied to clipboard
| Challenge: | Existing methods to embed knowledge graphs have ignored the fact that they contain two fundamentally different views: high-level ontology-view concepts and fine-grained instance-view entities. |
| Approach: | They propose a novel geometric representation that jointly embeds the two views of a KG using dual geometric representations. |
| Outcome: | Experiments on the public DBpedia KG and a newly-created industrial KG show the proposed method works well. |
Copied to clipboard
| Challenge: | Existing work on the Mutual Reinforcement Effect in information extraction has not been empirically validated . 76 percent of the 21 sub-datasets exhibit the Mutual Reforcement effect across languages . |
| Approach: | They propose a multilingual MRE mix dataset that integrates 21 sub-datasets covering English, Japanese, and Chinese. |
| Outcome: | The proposed framework reduces manual annotation effort while preserving structural requirements of MRE tasks. |
Copied to clipboard
| Challenge: | Existing knowledge editing techniques show limitations when applied to multi-hop reasoning . residual single-hop knowledge causes edited models to revert to original answers . |
| Approach: | They propose a knowledge editing method that incorporates a Knowledge Erasure mechanism for Large language model Editing (KELE) they propose an erasure function for residual knowledge and an injection function for new knowledge . |
| Outcome: | The proposed method significantly improves multi-hop reasoning capability of edited models. |
Copied to clipboard
| Challenge: | Unfairness is a well-known challenge in Recommender Systems (RSs) some approaches have started to improve fairness in offline or static contexts, but it often exacerbates over time, leading to significant problems like the Matthew effect, filter bubbles, and echo chambers. |
| Approach: | They propose a framework to promote multi-interest diversity fairness in RSs by establishing diverse hypergraphs through contrastive learning. |
| Outcome: | The proposed framework achieves state-of-the-art performance while effectively alleviating unfairness in two CRS-based datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for insurance claims adjudication are limited to information retrieval or simple multiple-choice setups. |
| Approach: | They propose a benchmark that provides complete reasoning traces linking factual inputs, relevant policy clauses, and final verdicts. |
| Outcome: | The proposed benchmark shows that models often produce correct decisions but fail to provide precise justifications, highlighting a critical discrepancy between decision accuracy and logical reasoning capabilities. |
Copied to clipboard
| Challenge: | Existing CRS datasets suffer from data inextensibility and semantic inconsistency . |
| Approach: | They introduce the LLM-REDIAL dataset to facilitate the research in CRS by leveraging large language models to generate high-quality dialogues. |
| Outcome: | The proposed dataset is the largest multi-domain CRS dataset which consists of 47.6k multi-turn dialogues with 482.6k utterances across 4 domains. |
Copied to clipboard
| Challenge: | Existing methods for MAS suffer from high token consumption and inefficiency due to frequent generation and communication among multiple agents. |
| Approach: | They propose a multi-agent system based on large language models that identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance. |
| Outcome: | The proposed method reduces prompt token consumption and completion token consumption by 18.4% and improves task performance by 1.14. |
Copied to clipboard
| Challenge: | Existing studies have focused on integrating large language models (LLMs) with information extraction (IE) however, the best approach to incorporate information with LLMs for IE remains an open question. |
| Approach: | They propose to use a Chinese IT dataset to perform RA-IT for IE . they use semantically similar examples from the training dataset as the context . |
| Outcome: | The proposed approach is evaluated in English and Chinese scenarios. |
Copied to clipboard
| Challenge: | Argumentation mining (AM) aims to detect arguments and their inherent relations from textual compositions. |
| Approach: | They propose a method to model the inter-relationships among three subtasks within a generative framework. |
| Outcome: | The proposed method achieves state-of-the-art performance on two AM benchmarks. |
Copied to clipboard
| Challenge: | Current approaches focus on isolated meme analysis, either for harmful content detection or standalone interpretation, overlooking a fundamental challenge: the same meme can express different intents depending on its conversational context. |
| Approach: | They propose a benchmark to evaluate how large vision language models understand memes in their original context. |
| Outcome: | The proposed benchmark evaluates how large vision language models understand meme intent in their original context. |
Copied to clipboard
| Challenge: | Existing methods for text-to-image synthesis lack systematic error analysis and refinement strategies, resulting in limited reliability and effectiveness. |
| Approach: | They propose a plug-and-play multi-agent system called GenPilot that integrates error analysis, clustering-based adaptive exploration, fine-grained verification and a memory module for iterative optimization. |
| Outcome: | The proposed method improves text consistency and structural coherence on images with a plug-and-play system. |
Copied to clipboard
| Challenge: | Incomplete learning is widespread and heterogeneous in large language models . authors identify five recurrent sources of incomplete learning: missing prerequisite knowledge, conflicts between SFT supervision and pre-training knowledge, internal inconsistencies within SFT data, left-side forgetting during sequential fine-tuning, and insufficient optimization for rare or complex patterns. |
| Approach: | They propose a diagnostic-first framework that maps incomplete learning to causes . they identify five recurrent sources of incomplete learning: missing prerequisite knowledge, conflicts between supervision and pre-training knowledge, internal inconsistencies, left-side forgetting during sequential fine-tuning, and insufficient optimization for rare or complex patterns. |
| Outcome: | The proposed framework maps incomplete learning to causes using observable training and inference signals. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) parsing aims to translate sentences to semantic representation with a hierarchical structure, but there is a gap between their flat training objective and the hierarchic structure, which limits the model generalization. |
| Approach: | They propose a Hierarchical Curriculum Learning framework with Structure-level (SC) and Instance-level curricula (IC) that aims to translate sentences to semantic representation with a hierarchical structure. |
| Outcome: | Experiments on AMR2.0, AMR3.0, structure-complex and out-of-distribution situations confirm the effectiveness of the proposed framework. |
Copied to clipboard
| Challenge: | Existing methods to enhance the zeroshot generalization of DST fail to effectively decouple semantics of samples, limiting the zero-shot performance of the system. |
| Approach: | They propose a new learning schema that explicitly disentangles the semantics of seen data and leverages the performance and robustness with the mixture-of-experts mechanism. |
| Outcome: | The proposed model achieves state-of-the-art on multiWOZ2.1 with 10M trainable parameters and is robust to the mixture-of experts mechanism. |
Copied to clipboard
| Challenge: | Existing research on sentiment analysis based on eye movement signals has been attributed importance. |
| Approach: | They propose a linguistic probing eye movement paradigm to extract eye movement features based on the relationship between linguistic features and human reading behavior. |
| Outcome: | The proposed graph architecture achieves state-of-the-art performance on two sentiment analysis datasets with eye movement signals and three sentiment analysis data without eye movement signal. |
Copied to clipboard
| Challenge: | Recent research in large vision-language models has shown promising results, but the issue of hallucination remains. |
| Approach: | They propose an instruction-based method to reduce hallucinations in large vision-language models . they use disturbance instructions to exacerbate hallucinosity in multimodal fusion modules . |
| Outcome: | The proposed method reduces hallucinations in multimodal fusion modules by reducing alignment uncertainty and subtracting hallucines from the original distribution. |
Copied to clipboard
| Challenge: | Existing studies focus on evaluating large language models' ability to handle disagreement cases. |
| Approach: | They evaluate the performance of large language models in detecting offensive language at varying levels of agreement. |
| Outcome: | The proposed model improves detection accuracy and model alignment with human judgment by using disagreement samples in training. |
Copied to clipboard
| Challenge: | Numerical reasoning is ubiquitous in scientific research and financial analysis, but few benchmarks evaluate them by integrating numerical processing and mathematical reasoning. |
| Approach: | They propose a numerically-integrated hierarchical benchmark with 27,215 questions derived from 7,404 math word problems that spans 4 key cognitive aspects, 14 subcategories, and 2 modalities. |
| Outcome: | The proposed model improves Qwen-2.5 score with SOLVE and IRPO training. |
Copied to clipboard
| Challenge: | Existing preference-based reward modeling methods face a recursive dependency where each verifier requires a meta-verifier, leading to continuous and costly dependence on human annotation. |
| Approach: | They propose a dual RM that couples discriminative and generative reward models under a non-parametric meta-reward. |
| Outcome: | The proposed model achieves strong performance across major preference benchmarks and even when trained exclusively on language modality, it exhibits robust cross-modal transfer on Omni-RewardBench. |
Copied to clipboard
| Challenge: | Existing studies on fact verification lack a high-quality dataset for explainability . existing systems lack evidence retrieval and veracity prediction, limiting the ability to verify a claim . |
| Approach: | They propose a dataset for multi-hop explainable fact verification that summarises and modifies Wikipedia documents. |
| Outcome: | The proposed dataset aims to improve the accuracy of multi-hop explainable fact verification systems. |
Copied to clipboard
| Challenge: | Existing approaches to deploy large language models (LLMs) into RecSys have limited prompt length, unstructured item information, and un-constrained generation of recommendations. |
| Approach: | They propose a taxonomy-guided recommendation framework that empowers LLMs with category information in a systematic approach. |
| Outcome: | The proposed framework significantly improves recommendation quality compared to zero-shot approaches. |
Copied to clipboard
| Challenge: | Entropy-Guided Stepwise Scaling (EGSS) is a novel TTS framework for software engineering tasks. |
| Approach: | They propose an entropy-guided stepwise scaling framework that balances efficiency and effectiveness through entropic-guide encoding and robust test-suite augmentation. |
| Outcome: | EGSS boosts performance by 5–10% across all evaluated models, and reduces inference-time token usage by over 28% . compared to existing methods, EGS reduces token usage and reduce inference time by over 20% . |
Copied to clipboard
| Challenge: | Existing methods for multi-document reading comprehension cannot make full of the advantages of both approaches. |
| Approach: | They propose a multi-view fusion and multi-decoding method that integrates multiple documents for answering questions. |
| Outcome: | The proposed method improves on two mainstream multi-document reading comprehension datasets. |
Copied to clipboard
| Challenge: | Chinese Spelling Correction (CSC) is a task of detecting and correcting misspelled charac- ters in Chinese texts. |
| Approach: | They propose a model to learn detection and correction parts together from a multi-task learning perspective. |
| Outcome: | The proposed model can learn detection and correction parts together from a multi-task learning perspective. |
Copied to clipboard
| Challenge: | rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. |
| Approach: | They organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. |
| Outcome: | The authors summarize the current state of research in three main areas: hypothesis formulation, hypothesis validation, and manuscript publication. |
Copied to clipboard
| Challenge: | Recent advances in language models have led to significant improvements in mathematical reasoning across benchmarks. |
| Approach: | They analyze the prevalence of false positives in language models by using heuristic evaluation methods . they find that false positive models produce correct final answers but with flawed deduction paths . |
| Outcome: | The proposed model performance improvements are based on the proposed model and its evaluation metrics. |
Copied to clipboard
| Challenge: | a large pre-trained language model can cause computational burdens in inference time due to multiple forward passes. |
| Approach: | They propose a method to learn fixed text representations with source tasks . they learn a task-specific prefix for each source task independently and combine them . |
| Outcome: | The proposed method improves generalizability of representations with source tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are prone to forced generation when confronting ambiguous evidence or complex recursive dependencies. |
| Approach: | They propose a framework that imposes semantic and structural constraints via a financial metric knowledge graph. |
| Outcome: | a neuro-symbolic framework outperforms existing models on financial metric knowledge graphs. |
Copied to clipboard
| Challenge: | Simultaneous translation is notoriously dif- ficult due to word-order differences. |
| Approach: | They propose a prefix-to-prefix framework that implicitly learns to anticipate in a single translation model. |
| Outcome: | The proposed framework achieves low latency and reasonable qual- ity on 4 directions. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to achieve ecological validity, signal clarity, and reliable fine-grained labeling in multimodal Emotion Recognition (MER) Existing datasets lack spontaneity of real-life interactions, resulting in poor quality and inconsistent data quality. |
| Approach: | They propose a bilingual benchmark to resolve limitations of ecological validity and noise in existing datasets by combining strictly filtered static slices with a dynamic Streaming Monologue subset. |
| Outcome: | EmoS provides trusted ground truth that captures continuous emotional evolution. |
Copied to clipboard
| Challenge: | Existing work on math word problem solvers replace real numbers with symbolic placeholders to focus on logic reasoning. |
| Approach: | They propose to inject numerical properties into symbolic placeholders with contextualized representation learning schema to solve number representation dilemma. |
| Outcome: | The proposed model can solve MWP problems on English and Chinese benchmarks. |
Copied to clipboard
| Challenge: | Despite of significant achievements in improving instruction-following capabilities of large language models, the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge. |
| Approach: | They construct multi-turn instruction with 1.1K high-quality multi-turned conversations using the human-in-the-loop approach and examine their capabilities. |
| Outcome: | The proposed model shows that it is difficult to integrate multiple turns and balance competing objectives when instructions intersect or conflict. |
Copied to clipboard
| Challenge: | Existing studies use synthetic speech to train and evaluate SpeechRE models, hindering their development . modality gap issue limits performance of existing models, limiting future researches . |
| Approach: | They propose to use speech data to train and evaluate SpeechRE models by using real speech . they propose to train a cross-modal alignment model to bridge the modality gap . |
| Outcome: | The proposed model can train to bridge the modality gap between speech encoder and text decoder . the proposed model is based on two real SpeechRE datasets . |
Copied to clipboard
| Challenge: | DoTAT is a domain-oriented text annotation tool that can reduce the time for event annotation by 19.7% . the tool supports multi-person collaborative process with automatically merging and review . |
| Approach: | They propose a domain-oriented text annotation tool called DoTAT . it provides multi-person collaborative process with automatic merging and review . |
| Outcome: | The proposed tool can reduce the time for event annotation by 19.7% compared with existing tools. |
Copied to clipboard
| Challenge: | Existing evaluations rely on synthetic Gaussian noise or simplistic single-source interference, failing to capture the intricate, multi-layered acoustic dynamics that characterize authentic physical environments. |
| Approach: | They propose a robustness benchmark to stress-test Audio Large Models (ALLMs) using high-fidelity auditory scene simulations. |
| Outcome: | The proposed model performs well on a wide range of tasks, including automatic speech recognition, speech translation, and audio-based reasoning. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) based agents suffer from brittle procedural memory that is manually engineered or entangled in static parameters. |
| Approach: | They propose a procedural-memory repository that distills past agent trajectories into fine-grained, step-by-step instructions and higher-level, script-like abstractions. |
| Outcome: | The proposed repository can be used to improve agents' performance on travelplanner and Alfworld. |
Copied to clipboard
| Challenge: | Large language models (LLMs) suffer from severe hallucination issues due to the knowledge misalignment between the pre-training stage and the supervised fine-tuning stage. |
| Approach: | They propose a training objective with an abstention mechanism that selectively rejects tokens that misalign with the desired knowledge distribution via a special [REJ] token. |
| Outcome: | The proposed model selectively rejects tokens that misalign with the desired knowledge distribution via a special [REJ] token. |
Copied to clipboard
| Challenge: | Existing approaches to training pre-trained language models (PLMs) focus on static image modality; inevitably encounter modality gaps and noise; and treat all modalities. |
| Approach: | They propose a multimodal-augmented framework that can infuse multimodal semantics into PLMs and facilitate a self-balancing calibration of information allocation. |
| Outcome: | The proposed framework outperforms baselines on multiple NLP tasks and outperformed existing frameworks. |
Copied to clipboard
| Challenge: | Existing alignment strategies that focus on diverse and high-quality data often overlook the intrinsic uncertainty of tasks, learning all data samples equally. |
| Approach: | They propose to introduce the sample uncertainty into the alignment of different task scenarios by a simple fashion by setting the label smoothing value of training according to the uncertainty of individual samples. |
| Outcome: | The proposed model outperforms standard supervised fine-tuning on high-entropy tasks and complex low-entropic tasks. |
Copied to clipboard
| Challenge: | Low-code LLM is a visual programming interface that allows users to incorporate their ideas into the process without writing trivial prompts. |
| Approach: | They propose a human-LLM interaction framework that incorporates low-code visual programming interactions to achieve more controllable and stable responses. |
| Outcome: | The proposed framework enables users to incorporate ideas into the process without writing trivial prompts. |
Copied to clipboard
| Challenge: | Existing RLVRs lack visual faithfulness due to text-dominated reasoning . a novel framework to reinforce visual focus during policy optimization is proposed . |
| Approach: | They propose a framework to reinforce visual focus during policy optimization using visual attention compensation mechanism. |
| Outcome: | The proposed framework exhibits better visual activation and superior performance in multimodal reasoning and visual-dependent tasks. |
Copied to clipboard
| Challenge: | Knowledge editing is a promising technique to adapt large language models to new knowledge without retraining from scratch. |
| Approach: | They propose to use a multilingual dataset to translate a large-scale cross-lingual synthetic dataset from English to Chinese and then to evaluate their performance in Chinese. |
| Outcome: | The proposed method can change model performance on several special cases without retraining from scratch. |
Copied to clipboard
| Challenge: | Existing approaches to improve machine reading comprehension models are vulnerable and not robust to adversarial examples. |
| Approach: | They propose to construct positive example pairs which have same answer by augmentation and then introduce stability and contrastive loss to improve invariance of representation. |
| Outcome: | The proposed approach boosts the robustness of QA models across different tasks and attack sets significantly and consistently. |
Copied to clipboard
| Challenge: | Existing LLMs generate responses based on the dialogue context, overlooking the underlying linguistic cues about the user status exhibited in the context. |
| Approach: | They propose a linguistic cue-based chain-of-thoughts method which enhances the LLMs inference with an intermediate reasoning step to find cues exhibited in the dialogue. |
| Outcome: | The proposed method outperforms standard prompting methods on in-depth dialogue questions and linguistic cues exhibited in the context. |
Copied to clipboard
| Challenge: | Asymmetrical text matching is a fundamental problem in information retrieval and natural language processing. |
| Approach: | They propose a method that regularizes features vectors projected from different domains . WD-Match can be used to improve different text matching methods . |
| Outcome: | The proposed method outperforms existing methods and benchmarks on four datasets. |
Copied to clipboard
| Challenge: | Large Language Model (LLM) agents are transforming education by automating complex tasks and enhancing both teaching and learning processes. |
| Approach: | This survey analyzes recent advances in applying Large Language Model agents to educational settings . it highlights ethical issues, hallucination and overreliance, and integration with existing ecosystems . |
| Outcome: | The authors analyze the technologies enabling LLM agents and highlight key challenges in deploying them in educational settings. |
Copied to clipboard
| Challenge: | Extensive experiments demonstrate that our approach significantly alleviates task interference and forgetting. |
| Approach: | They propose a framework for supervised fine-tuning for large language models . they first fine-tail the model on each task to identify its core parameter regions . |
| Outcome: | The proposed framework outperforms vanilla fine-tuning and baselines on multiple public benchmarks on reasoning, dialogue, instruction following, and more. |
Copied to clipboard
| Challenge: | Existing work on affected package identification is limited by large language models . a recent study shows that 84% third-party packages contain security vulnerabilities . |
| Approach: | They propose a method to use LLM to generate the affected package . they propose supervised fine-tuning, retrieval augmented generation and a local search algorithm . |
| Outcome: | The proposed method has an average precision of 0.806 for identifying vulnerable packages in four most popular ecosystems in GitHub Advisory. |
Copied to clipboard
| Challenge: | Entity alignment (EA) aims to match identical entities across knowledge graphs (KGs) Graph neural network-based entity alignment methods have achieved promising results in Euclidean space, but KGs often contain complex local and hierarchical structures, which are hard to represent in a single space. |
| Approach: | They propose a method which unifies dual-space embedding to preserve the intrinsic structure of KGs. |
| Outcome: | The proposed method achieves state-of-the-art in structure-based EA on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing studies rely on entity information for sentence-level relation extraction (RE) but this can leak superficial and spurious clues of relations. |
| Approach: | They propose to use entity mentions to extract relations from textual context . they use a causal graph to model dependencies between variables in RE models . |
| Outcome: | The proposed method yields significant gains on both effectiveness and generalization for RE. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning Large Language Models are slow and lack of performance. |
| Approach: | They propose a Zeroth-Order optimization framework that uses forward passes to fine-tune Large Language Models. |
| Outcome: | The proposed framework achieves 1.7 to 3.0 wall-clock acceleration on LLaMA and OPT models. |
Copied to clipboard
| Challenge: | Large visionlanguage models (LVLMs) are a powerful visual-language reasoning tool. |
| Approach: | They propose to integrate attention analysis with LLaVA-CAM to determine interactions between visual representations. |
| Outcome: | The proposed approach can be used to determine interactions between visual representations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have emerged as a promising avenue for time series forecasting . existing approaches face limitations such as marginalized role in model architectures and lack of interpretability. |
| Approach: | They propose a framework that exploits LLM causal reasoning to discover and use directed causal associations among covariates. |
| Outcome: | The proposed model improves predictive accuracy while yielding transparent, traceable reasoning about variable interactions. |
Copied to clipboard
| Challenge: | Existing language evaluation benchmarks for English are limited to English . lack of such benchmarks makes it difficult to replicate success in other languages . |
| Approach: | They introduce a large-scale Chinese language understanding evaluation benchmark . the benchmark uses a set of current state-of-the-art pre-trained Chinese models . |
| Outcome: | The first large-scale Chinese Language Understanding Evaluation (CLUE) benchmark is released . the benchmark evaluates models across a wide range of tasks on original Chinese text . existing language evaluation benchmarks are mostly limited to English . |
Copied to clipboard
| Challenge: | Short-form video hashtag recommendation (SVHR) is a classification or ranking problem that selects hashtags from a set of limited candidates. |
| Approach: | They propose a short-form video hashtag recommendation task that better represents how hashtags are created naturally by retrieving relevant hashtags from a large-scale hashtag pool as extra guidance signals. |
| Outcome: | The proposed model outperforms strong classification baselines on two short-form video datasets and the guidance signals boost the performance by 8.11 and 2.17 absolute ROUGE-1 scores on average. |
Copied to clipboard
| Challenge: | Automated theorem proving (ATP) benchmarks focus on symbolic inference but rarely involve understanding complex number combination reasoning. |
| Approach: | They propose a benchmark that requires a model to reduce a trigonometric expression with step-by-step proof and evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
| Outcome: | The proposed benchmark evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
Copied to clipboard
| Challenge: | Recent QA with logical reasoning questions requires passage-level relations among the sentences. |
| Approach: | They propose a discourse-aware graph network that aggregates passage-level clues for QA by using discourse-based information. |
| Outcome: | The proposed model achieves competitive results on two logical reasoning QA datasets. |
Copied to clipboard
| Challenge: | Faithful opinion summarization task involves generating a summary that covers the majority and minority opinions in documents. |
| Approach: | They propose a facets-guided opinion summarization method that induces facets and partitions documents into multiple facet-specific sets. |
| Outcome: | The proposed method outperforms state-of-the-art methods and multiple LLMs on two representative datasets and shows it can be used in specialty domains. |
Copied to clipboard
| Challenge: | Existing retrieval approaches often overlook patient-specific factual knowledge embedded in EHRs . existing retrieval frameworks often overlook this factual information, limiting its effectiveness in clinical decision-making. |
| Approach: | They propose a recurrence generation-augmented retrieval framework that synergizes factual and conceptual knowledge from dual sources. |
| Outcome: | The proposed framework improves on factual-aware medical QA benchmarks. |
Copied to clipboard
| Challenge: | Existing graph-based models excel at capturing structural information within TKGs but lack semantic comprehension abilities. |
| Approach: | They propose a plug-and-play module to enhance the performance of graph-based TKG models by exploring high-order histories step-by-step. |
| Outcome: | Experiments on three datasets and backbones show that CoH is effective in capturing high-order historical information for LLMs. |
Copied to clipboard
| Challenge: | Morphemes are a strong linguistic feature to capture lexical semantics, but lack of morpheme-informed resources and the expense of manual annotations hinder morphme-enhanced methods. |
| Approach: | They propose a task of Morpheme Sense Disambiguation with two subtasks in-text and in-word to generalize morpheme features on more tasks. |
| Outcome: | The proposed tasks are based on two morpheme-annotated datasets for Chinese . the best model yields a promising precision of 77.66% on in-text and 88.19% on in word . |
Copied to clipboard
| Challenge: | In the evolving landscape of large language models, the predominant focus has been on English and Chinese. |
| Approach: | They propose to utilize Arabic-specific vocabulary in the tokenizer to accelerate decoding. |
| Outcome: | The proposed model achieves decent performance comparable to the best Arabic LLMs across various Arabic benchmarks. |
Copied to clipboard
| Challenge: | Existing Large Language Models suffer from "Reasoning Collapse" on mathematical reasoning tasks where stochastic sampling produces lexical variations of the same erroneous logic rather than genuine semantic exploration. |
| Approach: | They propose a geometric inference framework that uses a spectral orthogonal probe to introduce semantically heterogeneous reasoning signals into the teacher's orthogonale complement of its dominant subspace. |
| Outcome: | The proposed framework improves accuracy and sampling efficiency over baseline methods on logic and code generation benchmarks. |
Copied to clipboard
| Challenge: | Existing work focuses on domain-specific enhancements during fine-tuning, the challenge of which lies in catastrophic forgetting of knowledge across other domains. |
| Approach: | They propose a data composition framework that allows LLMs to enhance their multi-domain capabilities during supervised fine-tuning. |
| Outcome: | The proposed framework improves multi-domain fostering performance by 29.77% compared to uniform weights. |
Copied to clipboard
| Challenge: | Recent Large Reasoning Models (LRMs) have demonstrated the ability to generate long chains of thought (LongCoT) LongCoT still faces challenges such as redundancy and logical incoherence. |
| Approach: | They propose a neural-symbolic reasoning framework that generates chains of thought . they propose Logic-Thinker, which transforms symbolic solvers into chains of thoughts . |
| Outcome: | The proposed framework outperforms models fine-tuned with ThinkerCoT on logic reasoning tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated superior ability to serve as ranking models, but they will exhibit discriminatory ranking behaviors based on users’ sensitive attributes (gender). |
| Approach: | They propose an evaluation method to investigate the severity of implicit ranking unfairness and a pair-wise regression method to conduct fair-aware data augmentation for LLM fine-tuning. |
| Outcome: | The proposed method outperforms existing methods in ranking fairness, achieving this with only a small reduction in accuracy. |
Copied to clipboard
| Challenge: | Existing models focus on a single therapy, but complex cases require flexible strategies among various therapies. |
| Approach: | They propose a multi-session, multi-therapy, and highly realistic benchmark . it is designed to address three key challenges: 1) can we train a highly realistic AI counselor? 2) How to systematically evaluate an AI counselor?" |
| Outcome: | The proposed benchmark is annotated with extensive professional skills and includes over 677 meta-skills and 4577 atomic skills. |
Copied to clipboard
| Challenge: | MLLMs are able to integrate multiple modalities into a single model to tackle complex tasks in real-world scenarios. |
| Approach: | They propose a comprehensive survey of Omni-MLLMs to address the challenges and opportunities of multimodal modeling. |
| Outcome: | The proposed model can integrate multiple modalities into a single model and provide novel perspectives. |
Copied to clipboard
| Challenge: | Existing methods such as GRPO often break down when task difficulty exceeds the model’s capacity, resulting in sparse rewards and inefficient training. |
| Approach: | They propose to measure the compatibility between external guidance and a model's intrinsic policy by introducing an adaptive framework to enhance reasoning performance while explicitly preserving high Affinity. |
| Outcome: | The proposed framework outperforms baseline models while maintaining high Affinity. |
Copied to clipboard
| Challenge: | Current models can not ensure the complexity of generated questions, so they may generate shallow questions that can be answered without multi-hop reasoning. |
| Approach: | They propose a controlled framework to generate multi-hop questions that contain key entities in multi- hop reasoning chains and a novel Transformer-based decoder to guarantee that key entities appear in the questions. |
| Outcome: | The proposed model outperforms the state-of-the-art model 25% on HotpotQA. |
Copied to clipboard
| Challenge: | Existing approaches to generating reward models rely on voting-based mechanisms to evaluate CoT outputs. |
| Approach: | They propose an efficient generative reward modeling framework grounded in model-internal uncertainty. |
| Outcome: | The proposed framework reduces inference cost while improving answer accuracy. |
Copied to clipboard
| Challenge: | Existing research on building ES conversation systems only considered single-turn interactions with users, which is over-simplified and has limited support for multi-turn systems. |
| Approach: | They propose a multi-turn ES conversation system that uses lookahead heuristics to estimate future user feedback after using particular strategies. |
| Outcome: | The proposed system significantly outperforms baselines in both dialogue generation and strategy planning. |
Copied to clipboard
| Challenge: | Recent efforts have turned to large language models (LLMs) as therapeutic agents for psychological therapy tasks, yet robustness across diverse patients remains underexplored. |
| Approach: | They propose a realistic role-play protocol for evaluating therapeutic dialogue agents and a de-identified, expert-annotated corpus of therapist–patient dialogues. |
| Outcome: | The proposed framework outperforms baselines on therapeutic outcomes and dialogue quality while improving conversational efficiency. |
Copied to clipboard
| Challenge: | Existing approaches to VideoQA often fail when complex reasoning or temporal relationships are involved. |
| Approach: | They propose a method that leverages reasoning processes generated by Multimodal Large Language Models to improve VideoQA models. |
| Outcome: | The proposed method improves VideoQA models on three benchmarks. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) is effective for aligning Large Language Models with human preferences, but its complex process limits its ability to continually learn human feedback. |
| Approach: | They propose a non-RL offline method to convert historical optimal policies into optimization constraints when continually learning new preferences. |
| Outcome: | The proposed method outperforms strong CL baselines in terms of reward-based evaluations and human assessment. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods can modify concept-level definitions, but they can distort instantial knowledge in LLMs, leading to poor performance. |
| Approach: | They construct a benchmark dataset ConceptEdit and establish new metrics for evaluation to investigate the editing capability of LLMs. |
| Outcome: | The proposed methods can modify concept definitions but can distort instantial knowledge in LLMs, leading to poor performance. |
Copied to clipboard
| Challenge: | a recent study shows that large language models have limited generalization in low-resource languages like Chinese. |
| Approach: | They propose to evaluate the zero-shot generalizability of large language models to the Chinese language . they release only half of the dataset publicly, with the remainder kept private . |
| Outcome: | The Chinese Instruction-Following Benchmark evaluates the generalizability of LLMs to the Chinese language. |
Copied to clipboard
| Challenge: | Existing TTRL methods rely on positive pseudo-labeling strategies to enhance reasoning capabilities. |
| Approach: | They propose a test-time reinforcement learning framework that mitigates label noise amplification by deriving pseudo-rewards from majority voting consensus. |
| Outcome: | The proposed framework mitigates label noise amplification by implementing selective positive pseudo-labeling and entropy-gated negative p-labeled pruning. |
Copied to clipboard
| Challenge: | Existing answer selection approaches for community question answering lack additional answer summaries due to redundancy and lengthiness issues of crowdsourced answers. |
| Approach: | They constructed a dataset which contains a corresponding reference summary for each original lengthy answer. |
| Outcome: | The proposed model improves the performance of a question and candidate answer on a WikiHowQA dataset. |
Copied to clipboard
| Challenge: | Recent advances in open-source Large Language Models (LLMs) have achieved notable successes in natural language processing. |
| Approach: | They propose a Parameter Efficient Fine-Tuning paradigm for improved fine-tuning and parameter efficiency in multi-task learning. |
| Outcome: | The proposed model outperforms existing methods on multi-task learning while reducing training costs by over 80% without losing general capability. |
Copied to clipboard
| Challenge: | Existing methods for evaluating reasoning paths are not efficient, but they are prone to errors. |
| Approach: | They propose a probabilistic self- and cross-consistency framework for mathematical reasoning that employs an accept-reject mechanism to encourage high-quality reasoning paths. |
| Outcome: | The proposed framework improves on 9 LLMs across 4 challenging benchmarks. |
Copied to clipboard
| Challenge: | Questionnaires are a professional research methodology used for qualitative and quantitative analysis of human opinions, preferences, and behaviors. |
| Approach: | They propose a questionnaire-based dataset that consists of 13,168 human-written questionnaires. |
| Outcome: | The proposed dataset contains 13,168 human-written questionnaires gathered from online platforms. |
Copied to clipboard
| Challenge: | Existing work on multilingual summarization and cross-lingual summmarization has been limited due to their different definitions. |
| Approach: | They propose to unify MLS and CLS into a more general setting, i.e. many-to-many summarization. |
| Outcome: | The proposed model outperforms the state-of-the-art models in the zero-shot directions. |
Copied to clipboard
| Challenge: | Existing GEC models fail to understand error patterns in varying contexts . a framework that generates copies of training instances with error-irrelevant contexts altered is proposed . |
| Approach: | They propose a framework that generates copies of training instances with error-irrelevant contexts altered. |
| Outcome: | The proposed framework outperforms baselines on the simulated tasks and outperformed existing models. |
Copied to clipboard
| Challenge: | Traditional recommender systems focus on the user-platform paradigm, where users are directly exposed under the control of the platform's recommendation algorithms. |
| Approach: | They propose a user-agent-platform paradigm where agent serves as the protective shield between user and recommender system that enables indirect exposure. |
| Outcome: | The proposed model improves 16.6% over baselines on four datasets and mitigates echo chamber effects and reduces model bias in disadvantaged users. |
Copied to clipboard
| Challenge: | Existing models ignore the inherent causality during related work generation, leading to spurious correlations which downgrade the models’ generation quality and generalizability. |
| Approach: | They propose a Causal Intervention Module for Related Work Generation (CaM) that captures causal relationships in related work generation and implements causal interventions to mitigate the negative impact of spurious correlations. |
| Outcome: | The proposed framework improves the quality and coherence of generated related work by capturing causalities in the generation process. |
Copied to clipboard
| Challenge: | Controlled Text Generation (CTG) aims to produce texts that exhibit specific desired attributes. |
| Approach: | They propose a pluggable CTG framework for Large Language Models to control text . they use attribute scorers to evaluate attributes of sentences and construct dynamic attribute graphs . |
| Outcome: | The proposed framework achieves a peak improvement of 19.29% over baseline methods in two tasks. |
Copied to clipboard
| Challenge: | Existing work on temporal relation extraction focuses on extracting temporal relations between events . previous work on relation extraction focused on focusing on event-centered tasks . |
| Approach: | They propose a temporal relation extraction model that unifies events, timexes and DCT . they propose combining event mentions, time expressions and document creation time into a sentence-style model . |
| Outcome: | The proposed model outperforms baselines on E-E, E-T and E-D significantly. |
Copied to clipboard
| Challenge: | Existing methods for model extraction attacks on large language models are inadequate . existing methods neglect the inconsistency between training tasks and LLM alignment . |
| Approach: | They propose a model extraction algorithm that uses a policy-gradient-style training task to guide the crafting of preference for the local model. |
| Outcome: | The proposed algorithm reduces query complexity while mitigating watermark protection . it can extract various state-of-the-art commercial LLMs while minimizing query complexity . |
Copied to clipboard
| Challenge: | Existing efforts to learn meaningful representations at the instance level are limited. |
| Approach: | They propose a deep embedded clustering model with cluster-level representation learning to jointly learn cluster and instance level representations. |
| Outcome: | The proposed model produces meaningful clusters on real-world short text datasets. |
Copied to clipboard
| Challenge: | Current outcome-centric verification paradigms neglect potential errors in the derivation process. |
| Approach: | They propose a process-aware RLVR training paradigm utilizing verifiers selected via **PRIME**. |
| Outcome: | The proposed approach outperforms the baseline verification paradigm on AIME24, AIME25, and Beyond-AIME models. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | a small model can be used to select effective prompt templates for a larger model. |
| Approach: | They propose a method to use a smaller model to select effective prompt templates for a larger model. |
| Outcome: | The proposed method significantly reduces the cost of prompt engineering while matching performance with optimal prompts among candidates. |
Copied to clipboard
| Challenge: | Existing knowledge embedding methods have limited performance on knowledge graph reasoning tasks . eureka is empowered to learn seen relations with sufficient training triples . |
| Approach: | They propose a neural insight learning framework called Eureka to bridge the “seen” to “unsea” gap . Eureca is empowered to learn seen relations with sufficient training triples while providing flexibility to learn unseen relations given only one trigger . |
| Outcome: | The proposed framework outperforms state-of-the-art models on seen and unseen relations . it can learn seen and unseen relationships with sufficient training triples . |
Copied to clipboard
| Challenge: | Existing LLMs are delicate and elusive in prompt words and styles. |
| Approach: | They propose an LLM-acquainted prompting technique that includes proficient "native-speaking" they propose to use in-context learning to prompt LLMs to perform high-performance reasoning . |
| Outcome: | The proposed technique achieves step-wise prompts in zero-shot scenarios while maintaining the prompt quality. |
Copied to clipboard
| Challenge: | Existing approaches to large language models are limited to historical backtesting and static data. |
| Approach: | a new large-language model is developed to simulate real-time trading in a virtual stock market . the agent trading arena simulates real-world bid-ask interactions and provides real-life trading scenarios . |
| Outcome: | The Agent Trading Arena simulates real-world market conditions and directly impacts price dynamics. |
Copied to clipboard
| Challenge: | Existing approaches to subjective assessment are inconsistent and inconsistent due to inconsistent scales and inherent preference biases. |
| Approach: | They propose a framework that operationalizes subjective assessment as comparative analysis and internalizes it via Language Buttons. |
| Outcome: | The proposed framework achieves state-of-the-art performance across multiple benchmarks and is scale-steerable. |
Copied to clipboard
| Challenge: | Existing efficiency-oriented methods attempt to shorten or mix reasoning strategies, yet often degrade reasoning capability. |
| Approach: | They propose a token-level dual-process framework that explicitly decouples efficiency and correctness signals during training. |
| Outcome: | The proposed framework reduces inference cost while maintaining strong reasoning ability across multiple benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to study the Matthew effect in Recommender Systems (RSs) however, it is amplified when the user interacts with the system over time. |
| Approach: | They propose a paradigm to alleviate the Matthew effect in conversational recommendation by learning multi-aspect preferences. |
| Outcome: | The proposed paradigm achieves state-of-the-art performance and superior of alleviating Matthew effect in conversational recommendation tasks. |
Copied to clipboard
| Challenge: | a new dataset is being developed to improve the capabilities of mobile GUI-control agents. |
| Approach: | They propose a dataset designed for generalist mobile GUI-control agents . they use screenshots from popular mobile applications to create a detailed GUI-annotated dataset . |
| Outcome: | The Android Multi-annotation EXpo (AMEX) is a large-scale dataset for generalist mobile GUI-control agents . it includes screenshots from popular mobile applications, which are annotated at multiple levels . |
Copied to clipboard
| Challenge: | Existing methods to assess similarity between sentences encounter over-estimation problem . compared to fuzzy representations, similarity is comparatively lower in terms of "The person's age". |
| Approach: | They propose a conditional contrastive learning framework that constructs positive and negative samples from two perspectives. |
| Outcome: | The proposed method achieves state-of-the-art performance with five models based on bi-encoder and tri-encoding architectures. |
Copied to clipboard
| Challenge: | Existing approaches to solve non-deterministic reasoning problems in large language models are limited by their complexity and lack of a clear understanding of the problem. |
| Approach: | They propose a method to diagnose and correct non-deterministic reasoning behaviors in large language models. |
| Outcome: | The proposed method outperforms baselines and WebQSP benchmarks on the widely used WebQ SP and CWQ benchmarks. |
Copied to clipboard
| Challenge: | Mobile GUI agents have attracted tremendous research participation recently. traditional approaches to mobile agent training rely on centralized data collection. |
| Approach: | They propose a benchmark for federated training and evaluation of mobile GUI agents . they find that federation algorithms consistently outperform local training . |
| Outcome: | The first benchmark for federated training and evaluation of mobile GUI agents is released . it features 6 datasets with 30+ subsets, 8 federation algorithms, 10+ base models, and over 800 apps across 5 categories . |
Copied to clipboard
| Challenge: | Existing studies ignore data imbalance in multilingual settings and do not utilize monolingual data. |
| Approach: | They propose a cross-lingual summarization model that aligns cross-linguistic data with high-resource monolingual data via contrastive and consistency loss. |
| Outcome: | The proposed model outperforms baseline models and consistently dominates on 45 language pairs. |
Copied to clipboard
| Challenge: | Recent advances in large language models have significantly improved automated code generation . however, the translation of complex mobile UI designs into high-fidelity front-end code remains a challenge . |
| Approach: | They propose a collaborative multi-agent system to reconstruct static single-page apps from mockups. |
| Outcome: | The proposed system outperforms existing methods in reconstructing complex app pages . the code and data will be released upon paper acceptance . |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models overlook intrinsic semantic associations between soft prompt tokens, leading to high discreteness and limited interactions. |
| Approach: | They propose a low-parameters Prompt Tuning method which leverages prompt decomposition and compressed outer product to facilitate multiple interactions among prompt tokens. |
| Outcome: | Experiments on six architectures and eight datasets show that the proposed method outperforms state-of-the-art methods in performance and efficiency. |
Copied to clipboard
| Challenge: | Recent studies reveal large language models lack logical reasoning abilities . logical relationship understanding is a core capability underlying genuine logical thinking . |
| Approach: | They propose a lightweight training framework targeting logical relationship understanding . they propose logical relation understanding framework that provides explicit supervision . |
| Outcome: | The proposed framework outperforms CoT-SFT training frameworks in logical relationship understanding tasks. |
Copied to clipboard
| Challenge: | Existing linear GCNs perform neural network operations in Euclidean space, which do not capture tree-like hierarchical structure of graphs. |
| Approach: | They propose a Lorentzian linear GCN framework that maps features into hyperbolic space and performs a feature transformation to capture the underlying tree-like structure of data. |
| Outcome: | The proposed framework achieves state-of-the-art accuracy on standard citation networks datasets and 81.3% on PubMed datasets. |
Copied to clipboard
| Challenge: | Existing hybrid question answering systems use a "prompt-and-pray" paradigm . context size limitations limit ability of many transformer-based LLMs to fit into a given prompt . |
| Approach: | They propose a superset of SQLite to act as a unified dialect for orchestrating reasoning across unstructured and structured data. |
| Outcome: | The proposed framework scales to massive datasets and improves performance while using 35% fewer tokens. |
Copied to clipboard
| Challenge: | Existing benchmarks designed to evaluate the reasoning capabilities of large models are limited in scope and lack flexibility to adapt difficulty according to evolving reasoning capacities of models. |
| Approach: | They propose a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models. |
| Outcome: | The proposed benchmark incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models. |
Copied to clipboard
| Challenge: | Current instruction tuning relies on teacher models or human intervention to generate and refine the instructions and responses for training, which are costly, non-sustainable, and may lack diversity. |
| Approach: | They propose a human/model-free compositional data synthesis method that can create rich and diverse augmentations from existing instruction tuning data to enhance large language models. |
| Outcome: | The proposed method improves performance over benchmarks and reduces training costs by 80% compared with original instruction tuning. |
Copied to clipboard
| Challenge: | Existing frameworks rely on static or rule-based topologies that fail to adapt to task requirements. |
| Approach: | They propose a generative framework that generates highly task-adaptive topologies . they validated the framework on multiple benchmarks and validated it on multiple platforms . |
| Outcome: | The proposed framework outperforms existing frameworks in task-adaptive communication topologies. |
Copied to clipboard
| Challenge: | Existing work relies on training with multi-lingual ability-related data, which may not be available for low-resource languages. |
| Approach: | They propose a multi-lingual ability-enhanced LLM that extracts language-agnostic ability-related weights from LLMs and combine them across different languages by simple addition and subtraction operations without training. |
| Outcome: | The proposed approach extracts language-agnostic ability-related weights from LLMs and combine them across different languages without training. |
Copied to clipboard
| Challenge: | Mobile GUI agents show promise in automating tasks but face significant generalization challenges in long-tail scenarios. |
| Approach: | They propose a benchmark framework for mobile GUI agents that measures the performance of GUI agents by analyzing their performance. |
| Outcome: | The LearnGUI benchmark outperforms existing methods in offline and online evaluations and demonstrates consistent gains across model architectures. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis is sensitive to multi-aspect challenges, resulting in multiple aspects in a sentence. |
| Approach: | They propose a framework that leverages an in-domain generator to construct more multi-aspect samples . they then boost the robustness of ABSA models via contrastive learning on these generated samples ." |
| Outcome: | The proposed framework outperforms baselines without any augmentations on accuracy and Macro- F1 . the proposed framework can generate more multi-aspect samples and boost the robustness of ABSA models . |
Copied to clipboard
| Challenge: | Existing work on dependency prior structure integration into pre-trained models is still unclear. |
| Approach: | They propose a dependency-based fusion attention paradigm which explicitly introduces dependency prior structure into pre-trained models and adaptively fuses it with semantic information. |
| Outcome: | The proposed model achieves state-of-the-art or competitive performance on 10 public datasets, demonstrating the benefits of adaptively fusing dependency structure in semantic matching task. |
Copied to clipboard
| Challenge: | Despite LLMs' impressive capabilities in musical knowledge, music reasoning remains an unsolved task. |
| Approach: | They propose an open-source large language model (LLM) that integrates intrinsic musical abilities into LLaMA2 and GPT-3.5. |
| Outcome: | The proposed model can understand and generate music with a pure text tokenizer without external multi-modal neural structures or tokenizers. |
Copied to clipboard
| Challenge: | Existing role-playing models focus on character knowledge and tones, but lack personality-indicative data to capture characters' minds. |
| Approach: | They propose to enhance role-playing agents (RPAs) via personality-indicative data by asking psychological scales to capture broad aspects of personality traits in individuals. |
| Outcome: | The proposed model exhibits advanced role-playing capabilities for both general and personality-related evaluations. |
Copied to clipboard
| Challenge: | Existing evaluation methods for mobile GUI agents rely on static frame assessments or offline static apps. |
| Approach: | They propose an evaluation system that leverages large language models as reward models to verify task completion and process achievement. |
| Outcome: | The proposed system addresses the limitations of traditional function based evaluation methods on online dynamic apps. |
Copied to clipboard
| Challenge: | Existing approaches for optimizing human annotation efforts are limited . et al., 2015) suggest that densely annotated image captions improve vision-language alignment . |
| Approach: | They propose an AI-in-the-loop methodology to maximize the number of annotated samples and improve their comprehensiveness under fixed budget constraints. |
| Outcome: | The proposed method improves annotation speed and retrieval performance over the parallel method. |
Copied to clipboard
| Challenge: | Existing methods operate by learning to fuse modalities, leading to frequent misjudgments. |
| Approach: | They propose a paradigm shift from *learning to fuse* to *learning the reason's process' inspired by the dual-process theory of human cognition, MIND operationalizes a self-improving loop. |
| Outcome: | The proposed model significantly outperforms baseline models and exhibits strong generalization. |
Copied to clipboard
| Challenge: | Recent studies on video large language models focus on model architectures and training datasets . interaction format between user and model is unsatisfactory for time-sensitive tasks . |
| Approach: | They propose a video-text duet interaction format that allows for continuous playback of the video . when a text message ends, the video continues to play, similar to the alternative of two performers in a duet. |
| Outcome: | The proposed format improves performance on time-sensitive tasks with minimal training efforts. |
Copied to clipboard
| Challenge: | Multi-modal Large Language Models have shown remarkable progress in visual contexts, yet their ability to convert visual figures into executable code remains underexplored. |
| Approach: | They propose to use a set of visual coding metrics to assess MLLMs' visual . pass rate, text-match ratio, and GPT-4V rating judgement to assess the quality of generated code and rendered images. |
| Outcome: | The proposed benchmark includes 132 high-quality matplotlib plots across six plot types, as well as 150 and 86 plots from Python’s and R’s plotly libraries respectively, totaling 368 plots. |
Copied to clipboard
| Challenge: | Recent methods to reduce the KV cache size fail to identify crucial KVs for generation while excluding others accurately, resulting in severe information loss. |
| Approach: | They propose an intention-aware KV cache eviction method that identifies and retains crucial KVs according to the attention distribution of intention, which semantically reflects the user’s goal and determines which part of the context is relevant. |
| Outcome: | The proposed method can maintain the model performance while reducing the KV cache size from 128K to 2K, leading to a 6.3x increase in decoding speed and 7.8x enhancement in memory efficiency compared to the default setting. |
Copied to clipboard
| Challenge: | Neural machine translation systems have become state-of-the-art approaches for Grammatical Error Correction (GEC) task. |
| Approach: | They propose a copy-augmented architecture for the Grammatical Error Correction task by copying unchanged words from the source sentence to the target sentence. |
| Outcome: | The proposed architecture outperforms all recently published state-of-the-art results by a large margin. |
Copied to clipboard
| Challenge: | Existing vision-Language-Action models are notoriously brittle to linguistic perturbations. |
| Approach: | They propose a probabilistic framework that disentangles physical affordance from semantic execution. |
| Outcome: | The proposed framework disentangles physical affordance from semantic execution. |
Copied to clipboard
| Challenge: | Input too many potential entity types would distract the model inevitably. |
| Approach: | They propose to use a generative model to exploit entity types' merit on promoting NER task by appending a type matching model to identify the entity types most likely to appear in the sentence. |
| Outcome: | The proposed framework exploits entity types' merit on promoting NER task by adding auxiliary task to the model to discover the entity types. |
Copied to clipboard
| Challenge: | Existing studies show that initializing NMT models with pre-trained language models (LM) can speed up the model training and boost the model performance. |
| Approach: | They propose a method to control copying behaviors in NMT models by initializing them with pre-trained language models (LM) they propose to use a metric called copy ratio to control the copying behavior in decoding. |
| Outcome: | The proposed method improves translation performance by controlling copying behaviors for pre-training based models. |
Copied to clipboard
| Challenge: | Existing methods for generating SQL queries using natural language questions produce inconsistent NLQ-SQL pairs. |
| Approach: | They propose a text-to-SQL data synthesis framework that generates domain-relevant questions . they synthesize NLQ-SqL pairs that are domain-specific and intent-consistent . |
| Outcome: | The proposed method outperforms closed-source LLMs on the Text-to-SQL task. |
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is the preliminary step for training non-autoregressive translation models, but it can lose important information for translating low-frequency words. |
| Approach: | They propose a knowledge distillation method which trains NAT student on external monolingual data with AT teacher trained on the original bilingual data. |
| Outcome: | Extensive experiments on eight WMT benchmarks show that monolingual KD outperforms the standard KD by improving low-frequency word translation without introducing any computational cost. |
Copied to clipboard
| Challenge: | Existing molecule-language models obscure the hierarchical organization of chemical semantics . Existing models rely on linear or uniform encodings, causing structural distortion . |
| Approach: | They propose a framework that integrates intrinsic molecular topology into large language models. |
| Outcome: | The proposed framework improves on cross-modal retrieval, captioning, and property prediction benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to large vision-language models fail to capture interleaved nature of human visual-verbal reasoning processes. |
| Approach: | They propose a framework that integrates visuospatial and linguistic domains to facilitate multimodal slow thinking by enabling progressive visual-textual reasoning. |
| Outcome: | Experiments show that VisuoThink significantly improves reasoning capabilities even without fine-tuning. |
Copied to clipboard
| Challenge: | Various machine learning methods for tabular data lack accurate confidence estimation, which is needed for high-risk sensitive applications such as credit modeling and financial fraud detection. |
| Approach: | They propose a general post-training confidence calibration framework to calibrate the confidence of current machine learning models by employing graph neural networks to model the relationships between different samples. |
| Outcome: | The proposed framework improves the confidence estimation on tabular datasets by using graph neural networks to model the relationships between different samples. |
Copied to clipboard
| Challenge: | Existing dialogue systems process conversational turns in isolation, overlooking event structures that guide natural interactions. |
| Approach: | They propose a framework that explicitly models relationships between conversational events to generate more contextually appropriate dialogue responses. |
| Outcome: | Experiments on three dialogue datasets show that the proposed approach produces more natural responses while requiring less computational overhead. |
Copied to clipboard
| Challenge: | Text classification tasks often encounter few-shot scenarios with limited labeled data, and addressing data scarcity is crucial. |
| Approach: | They propose a self-evolution learning (SE) based mixup approach for data augmentation in text classification which generates more adaptive and model-friendly pseudo samples for the model training. |
| Outcome: | The proposed approach can generate more adaptive and model-friendly pseudo samples for the model training. |
Copied to clipboard
| Challenge: | Experimental results show that PT and BT are nicely complementary to each other. |
| Approach: | They introduce two probing tasks for PT and BT respectively and investigate their complementarity. |
| Outcome: | The proposed methods establish state-of-the-art on the WMT16 English-Romanian and English-Russian benchmarks. |
Copied to clipboard
| Challenge: | Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work. |
| Approach: | They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations. |
| Outcome: | The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can be effective at rewriting toxic content, but they often default to overly polite rewrites, distorting the emotional tone and communicative intent. |
| Approach: | They evaluate 17 large language models with variant architectures to evaluate their ability to rewrite toxic content while preserving the speaker's original intent. |
| Outcome: | The first Chinese detoxification dataset explicitly designed to preserve sentiment polarity is evaluated across five real-world scenarios. |
Copied to clipboard
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Copied to clipboard
| Challenge: | Recent advances in neural theorem-proving resort to large language models and tree searches. |
| Approach: | They propose a Dynamic-Tree Driven Theorem Solver to accommodate general theoremes by guiding the search procedure with state confidence and proof-level values. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two popular theorem-proving datasets with a 6.65% improvement on average in terms of success rate. |
Copied to clipboard
| Challenge: | Empirical evaluations demonstrate that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models. |
| Approach: | They propose a framework that reformulates supervision as a constrained optimization problem. |
| Outcome: | Empirical evaluations show that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models. |
Copied to clipboard
| Challenge: | Recent work shows that large-scale annotated datasets are essential for training state-of-the-art Question Answering (QA) models. |
| Approach: | They use large-scale annotated datasets to train question answering models . they use feedback data collected from deployed QA systems to provide cheaper supervision . |
| Outcome: | The proposed model improves on the large scale annotated datasets from QA systems . the proposed model can be easily supervised on large-scale unlabeled web data . |
Copied to clipboard
| Challenge: | Existing methods to analyze filter bubbles in the static recommendation environment are unable to burst them during user interactions. |
| Approach: | They propose a paradigm to learn multi-grained user preferences during dynamic user-system interactions via natural language conversations to burst filter bubbles. |
| Outcome: | The proposed paradigm achieves state-of-the-art performance and the superior of bursting filter bubbles in the conversational recommendation system. |
Copied to clipboard
| Challenge: | Extensive experiments demonstrate the effectiveness of SGTC across various tasks. |
| Approach: | They propose a generative tool invocation framework that introduces structure-aware semantic tokenization to encode tools as discrete code sequences. |
| Outcome: | The proposed framework reduces the size of the representation space and underutilizes collaborative signals among tools in downstream tasks. |
Copied to clipboard
| Challenge: | Recent studies have shown that pre-trained language models contain smaller matching subnetworks that are not robust to adversarial examples. |
| Approach: | They propose a method to find robust tickets hidden in pre-trained language models by learning binary weight masks and an adversarial loss objective to guide the search. |
| Outcome: | The proposed method improves on previous work on adversarial robustness evaluation. |
Copied to clipboard
| Challenge: | Existing reward models produce scalar scores and struggle to incorporate critiques in a natural language format. |
| Approach: | They propose a framework that predicts critiques and rewards using self-generated critiques without extra supervision. |
| Outcome: | The proposed framework improves reward modeling accuracy by 3.7%-7.3% compared to standard reward models and LLM judges. |
Copied to clipboard
| Challenge: | Effective content moderation is essential for video platforms to safeguard user experience and uphold community standards. |
| Approach: | They propose a method to transform a generative MLLM into a multimodal classifier using minimal discriminative training data. |
| Outcome: | The proposed method improves F1 score by 66.50% over traditional classifiers while requiring only 2% of the fine-tuning data. |
Copied to clipboard
| Challenge: | Spoken language understanding (SLU) suffers from error propagation from automatic speech recognition (ASR) in actual scenarios. |
| Approach: | They propose a framework which calibrates bias and errors and achieves adaptive-balanced decoupling training by a prototype-based loss model. |
| Outcome: | The proposed framework outperforms existing approaches and achieves state-of-the-art performance on three datasets. |
Copied to clipboard
| Challenge: | a recent study shows that vision-language models have modality gaps that persist even in well-aligned models. |
| Approach: | They propose a modality-dominance score to measure and leverage modality gaps . they propose automatic interpretability metrics to evaluate these features in a scalable manner . |
| Outcome: | The proposed framework allows for training-free probing and editing methods for understanding model perception across genders and generating adversarial examples. |
Copied to clipboard
| Challenge: | Existing efficient test-time scaling methods introduce budget constraints or early stop mechanisms to avoid overthinking for straightforward questions but add human bias to the reasoning process. |
| Approach: | They propose a framework that dynamically adapts reasoning depth based on question complexity. |
| Outcome: | Experimental results show that the proposed framework achieves higher accuracy than baseline methods and reduces computational overhead by up to 25.2%. |
Copied to clipboard
| Challenge: | Existing benchmarks measure common sense knowledge indirectly or without reasoning. |
| Approach: | They propose a benchmark to test whether a system can differentiate natural language statements that make sense from those that do not make sense. |
| Outcome: | The proposed benchmarks show that models trained on large corpora perform better than humans on some benchmarks. |
Copied to clipboard
| Challenge: | Despite significant advances in video-language modeling, hallucinations remain a persistent challenge in video large language models. |
| Approach: | They present a systematic taxonomy that categorizes hallucinations into two core types: dynamic distortion and content fabrication. |
| Outcome: | The proposed taxonomy categorizes hallucinations into two core types: dynamic distortion and content fabrication. |
Copied to clipboard
| Challenge: | Existing knowledge selection methods are costly to learn and difficult to interpret when errors arise in the generated responses. |
| Approach: | They propose a generator-agnostic knowledge selection method to select context-related knowledge among different knowledge structures and variable knowledge requirements. |
| Outcome: | The proposed method can select knowledge accurately in advance and reduce learning, adjustment, and interpretation burden of later models. |
Copied to clipboard
| Challenge: | Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese sentences. |
| Approach: | They propose to integrate phonetic and character representations to allow interaction between textual and phonetic information. |
| Outcome: | The proposed method is superior to other methods on three benchmarks. |
Copied to clipboard
| Challenge: | Existing neural solvers only generate binary expression trees that contain basic arithmetic operators and do not explicitly use the math formulas. |
| Approach: | They propose a sequence-to-general tree that generates interpretable and executable operation trees where nodes can be formulas with an arbitrary number of arguments. |
| Outcome: | The proposed tree generates interpretable and executable operation trees with formulas with an arbitrary number of arguments. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models (LLMs) have strong multilingual abilities, giving them the potential to perform M2MS in real applications. |
| Approach: | They propose to use many-to-many summarization (M2MS) to generate a brief summary in any language given a document also in any other language. |
| Outcome: | The proposed model outperforms zero-shot LLMs in terms of automatic evaluations. |
Copied to clipboard
| Challenge: | Experimental results show that SmartTrim accelerates the original model by 2-3 times with minimal performance degradation. |
| Approach: | They propose an adaptive acceleration framework which prunes redundant token representations and attention heads within each layer of the original model. |
| Outcome: | The proposed framework accelerates the original model by 2-3 times with minimal performance degradation across vision-language tasks. |
Copied to clipboard
| Challenge: | Existing methods for generating reward models focus on outcome-level supervision, neglecting analytical process quality, which constrains their potential. |
| Approach: | They propose a novel reward model that leverages self-reflection to assess analytical quality and enhance preference modeling. |
| Outcome: | The proposed model improves performance on four benchmarks and significantly mitigates positional bias. |
Copied to clipboard
| Challenge: | Recent work on domain adaptation for text summarization fails to account for the huge gap between dialogue and general articles. |
| Approach: | They propose a hypernetwork-assisted encoder-decoder architecture with parameter-efficient fine-tuning to disentangle domain-invariant knowledge from source domains while learning specific knowledge of the target domain. |
| Outcome: | The proposed model can disentangle domain-invariant knowledge from source domains while learning specific knowledge of the target domain. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained models are beneficial to Chinese Word Segmentation (CWS). However, these models lack task-specific prior segmentation knowledge. |
| Approach: | They propose a pre-trained Chinese word segmentation model MetaSeg which incorporates meta learning into a multi-criteria pre-training task. |
| Outcome: | Empirical results show that MetaSeg can achieve new state-of-the-art performance on twelve widely-used CWS datasets and significantly improve model performance in low-resource settings. |
Copied to clipboard
| Challenge: | Existing methods for debiasing factchecking models learn such biases instead of understanding the semantic relationship between the claim and evidence. |
| Approach: | They propose a counterfactual framework CLEVER which is augmentation-free and mitigates biases on the inference stage. |
| Outcome: | The proposed method is augmentation-free and mitigates biases on the inference stage. |
Copied to clipboard
| Challenge: | Recent models have extended Corresponding Author. context lengths to millions of tokens while maintaining reasoning and comprehension capabilities. |
| Approach: | They propose a benchmark to evaluate the ability of large language models to extract sequential information items from long contexts. |
| Outcome: | The proposed model achieves maximum accuracy of 63.50% on six well-known LLMs. |
Copied to clipboard
| Challenge: | Existing studies on learning social media content focus on single modal or bi-modal learning, but this approach is non-trivial and challenging because content is multi-modal and involves several types of data, including text, audio, and image. |
| Approach: | They propose to combine textual, acoustic, and visual information to learn social media content by fusing them jointly. |
| Outcome: | The proposed model outperforms the state-of-the-art approaches on real-world datasets by a large margin. |
Copied to clipboard
| Challenge: | Existing studies on prompt engineering have focused on optimizing models for performance under stylistic perturbations. |
| Approach: | They conduct the first analysis of n-gram token-level mechanisms . they find that higher average performance is inherently associated with lower variance and greater stability. |
| Outcome: | The proposed model reduces the variance of the generated code by 40% . the proposed model is based on a large-scale dataset of 132,000 prompt variants . |
Copied to clipboard
| Challenge: | Pre-trained language models have been widely applied to cross-domain NLP tasks like sentiment analysis, but fine-tuning them on the source domain tends to overfit, leading to inferior results on the target domain. |
| Approach: | They propose to pre-train a sentiment-aware language model (SentiX) via domain-invariant sentiment knowledge from large-scale review datasets and utilize it for cross-domain sentiment analysis tasks without fine-tuning. |
| Outcome: | The proposed model achieves state-of-the-art in all the cross-domain sentiment analysis tasks and can be trained with only 1% samples and better than BERT with 90% samples. |
Copied to clipboard
| Challenge: | Recent studies have shown that ChatGPT has limitations such as failing to ask clarifying questions to ambiguous queries or refusing problematic user requests. |
| Approach: | They propose a Proactive Chain-of-Thought prompting scheme which augments LLMs with the goal planning capability over descriptive reasoning chains to trigger proactivity. |
| Outcome: | The proposed scheme augments LLMs with the goal planning capability over descriptive reasoning chains to trigger the proactivity of LLM-based proactive dialogue systems. |
Copied to clipboard
| Challenge: | Existing methods to summarize dialogues are difficult due to insufficient training data and low information density. |
| Approach: | They propose a curriculum-based prompt learning method with self-training that gradually increases the degree of prompt perturbation, improving dialogue understanding and modeling capabilities. |
| Outcome: | The proposed model outperforms baseline models on the AMI and ICSI datasets and human evaluations show it is superior in the quality of the summary generation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning. |
| Approach: | They propose a virtual API server and stable evaluation system to assess the stability of large-scale real-time APIs. |
| Outcome: | The proposed benchmarks demonstrate the stability of the proposed system and its caching system. |
Copied to clipboard
| Challenge: | Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored. |
| Approach: | They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain. |
| Outcome: | MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows. |
Copied to clipboard
| Challenge: | Existing methods for Temporal Knowledge Graph reasoning capture intra- and inter-time latent relations between entities that appear at different times. |
| Approach: | They propose a Latent relations Learning method for TKG reasoning that captures latent relations between entities at different times. |
| Outcome: | The proposed method exploits the intra- and inter-time latent relations of entities at different times. |
Copied to clipboard
| Challenge: | Recent studies have investigated methods to improve the safety of large language models (LLMs) safety training involves fine-tuning the LLM with adversarial samples, which activate the LRM’s capabilities against jailbreak. |
| Approach: | They propose a safety training approach that integrates safety training and safeguards to train the LLM to perform harmfulness detection on its own outputs. |
| Outcome: | The proposed method reduces harmful output and adds a [harmful] or [harmless] tag to the end of the LLM's response. |
Copied to clipboard
| Challenge: | Existing methods to train large language models on private data are not effective because they rely on a local model for generation, resulting in a performance decline, or expose private data to API servers. |
| Approach: | They propose a client-server framework which enhances synthetic data quality and improves model performance while ensuring privacy. |
| Outcome: | The proposed framework improves model performance and privacy while learning local knowledge from the private data with differential privacy (DP) and distilling professional knowledge from server. |
Copied to clipboard
| Challenge: | Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. |
| Approach: | They propose a task that requires models to reason over multiple pieces of evidence . they construct a large-scale dataset comprising 15k multi-hop claims paired with multimodal evidence - generated and refined using large language models with additional input from human feedback. |
| Outcome: | The proposed method is based on human performance benchmarks and human reasoning hops. |
Copied to clipboard
| Challenge: | Existing methods to narrate movies with no actors are difficult to implement in real situations . a new metric is proposed to provide the best correlation with human evaluation . |
| Approach: | They propose a large-scale Chinese movie benchmark to help visually impaired enjoy movies . they propose metric called Movie Narration Score (MNScore) which achieves best correlation with human evaluation. |
| Outcome: | The proposed method outperforms baselines and the existing methods. |
Copied to clipboard
| Challenge: | Existing methods for code summarization are limited in resources and require atomic commands and category constraints to enhance code representations. |
| Approach: | They propose a framework that leverages limited atomic commands and category constraints to enhance code representations. |
| Outcome: | The proposed framework outperforms baseline methods in a number of domains and demonstrates superiority over competing frameworks. |
Copied to clipboard
| Challenge: | Long-form question answering requires two procedures: information retrieval and information synthesis. |
| Approach: | They propose a Chinese long-form question answering dataset called WebCPM . the dataset is based on a web search interface that engages with a search engine in real time . |
| Outcome: | The proposed dataset generates answers that are no worse than human-written ones . the dataset is the first Chinese LFQA dataset . |
Copied to clipboard
| Challenge: | Existing methods for question generation over knowledge bases rely on annotated data for fine-tuning . emergence of Large Language Models (LLMs) has shown impressive generalization ability in few-shot tasks. |
| Approach: | They propose to use a logical form to generate a question in a reasoning problem . they propose to extend the prompting method into a method that can generate questions in logical forms . |
| Outcome: | The proposed method outperforms baselines on three public KBQG datasets. |
Copied to clipboard
| Challenge: | despite advances in multimodal large language models, the challenge of interpreting long-form videos remains a challenge . despite advancements in video-language benchmarks, the inefficiency in temporal grounding and limited pre-trained context window size remains . |
| Approach: | They propose a framework that bootstraps MLLMs with advanced temporal grounding capabilities and broadens their contextual scope. |
| Outcome: | The proposed framework significantly enhances the temporal capabilities of existing MLLMs. |
Copied to clipboard
| Challenge: | Existing work uses sentences within the same batch as negatives, which suffers from easy negatives. |
| Approach: | They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment . |
| Outcome: | The proposed model achieves state-of-the-art on several tasks. |
Copied to clipboard
| Challenge: | Emotion Support Conversation (ESC) is a crucial application for reducing stress and providing emotional guidance. |
| Approach: | They re-organize 2,801 role-playing cards to define roles of role-players . they train a specific role- playing model called ESC-Role which behaves more like a confused person than GPT-4 . |
| Outcome: | The proposed model behaves more like a confused person than GPT-4, and the model performs better than GPLs. |
Copied to clipboard
| Challenge: | ConsistRM is a self-training framework that enables effective and stable GRM training without human annotations. |
| Approach: | They propose a self-training framework that enables effective and stable GRM training without human annotations. |
| Outcome: | The proposed framework outperforms vanilla Reinforcement Fine-Tuning (RFT) by 1.5% on five benchmark datasets. |
Copied to clipboard
| Challenge: | Form understanding is a complex task because of the textual contents and organizational structure of forms. |
| Approach: | They propose to use multimodal methods to extract key-value pairs from forms . they validate their method on two benchmarks and demonstrate their effectiveness . |
| Outcome: | The proposed method is validated on two benchmarks, MedForm and FUNSD. |
Copied to clipboard
| Challenge: | characterization imaging data is fundamental to acquiring materials information. |
| Approach: | a team of researchers develop a benchmark for materials characterization imaging data . the goal is to bridge this gap by addressing 1,500 questions that require expert-level expertise. |
| Outcome: | a new benchmark for materials characterization imaging data is presented . the benchmark reveals that MLLMs perform poorly when addressing higher-level questions . |
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is commonly used to construct synthetic data for training non-autoregressive translation models. |
| Approach: | They propose to use knowledge distillation to generate training data for non-autoregressive translation models by leveraging pretraining. |
| Outcome: | The proposed approach achieves 28.2 and 33.9 BLEU points on the WMT14 English-German and WMT16 Romanian-English datasets. |
Copied to clipboard
| Challenge: | Recent advances in audio large language models have led to their potential privacy implications unexplored. |
| Approach: | They propose a benchmark to examine whether ALLMs leak user privacy through acoustic voiceprints. |
| Outcome: | The proposed benchmark is constructed from over 22,000 real-world audio clips. |
Copied to clipboard
| Challenge: | High-quality data is the cornerstone of advancing large language models, but the supply of premium data is nearing depletion, while vast stale corpora remain underutilized. |
| Approach: | They propose a framework to restore stale data affinity by quantifying the latent value of samples and employing a dynamic renovation strategy selection mechanism to determine the optimal component-level strategy. |
| Outcome: | The proposed framework achieves performance improvements using less than 10% of the data volume, underscoring that the latent potential of stale corpora remains largely untapped. |
Copied to clipboard
| Challenge: | Existing multimodal neural machine translation models focus on bilingual translation, but experimental results show that they outperform the text-only baselines and multilingual multimodal methods by a large margin. |
| Approach: | They propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P) this framework aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. |
| Outcome: | The proposed framework outperforms previous text-only baselines and multilingual multimodal methods by a large margin. |
Copied to clipboard
| Challenge: | Existing methods for generating a entailment tree exhibit the reasoning chains from knowledge facts to predicted answers, but they have large fact search spaces and error accumulation problems resulting in the generation of invalid steps. |
| Approach: | They propose a Fact-Retrieval and Verification Augmented bidirectional entailment tree generation method that contains two systems. |
| Outcome: | The proposed method outperforms existing models and achieves state-of-the-art performance in fact selection and structural correctness. |
Copied to clipboard
| Challenge: | Existing studies on multimodal abstractive summarization focus on how to use extracted visual features to produce a concise summary given the multimodal data. |
| Approach: | They propose to improve the visual quality of the multimodal abstractive summarization model by capturing summary-oriented visual features. |
| Outcome: | The proposed approach achieves state-of-the-art under 44 languages and is highly effective on high-resource English datasets. |
Copied to clipboard
| Challenge: | Existing Chinese preference datasets suffer from limited scale, restricted domain coverage, and insufficiently rigorous data validation. |
| Approach: | They propose an LLM-based data annotation pipeline with no human intervention to annotate Chinese preference datasets. |
| Outcome: | The proposed pipeline outperforms existing Chinese preference datasets on AlignBench and Chinese Reward Benchmark. |
Copied to clipboard
| Challenge: | Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have substantially improved the reasoning abilities of Large Language Models (LLMs). |
| Approach: | They propose a method that balances exploration and exploitation in the hidden-state space of response trajectories. |
| Outcome: | The proposed model yields consistent improvements across models, algorithms and reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for learning complex sentences with multiple aspects are ill-equipped to learn complex sentences . |
| Approach: | They propose a mutual enhanced transformation network for the ABSA task . it improves representation learning of the aspect with contextual semantic features . |
| Outcome: | The proposed model improves representation learning of the aspect with contextual semantic features, giving the aspect more abundant information. |
Copied to clipboard
| Challenge: | Text-based methods lag behind graph embedding-based approaches for knowledge graph completion (KGC) |
| Approach: | They propose three types of negatives to improve contrastive learning to improve learning efficiency. |
| Outcome: | The proposed model outperforms embedding-based methods on several benchmark datasets. |
Copied to clipboard
| Challenge: | Large language models suffer from severe hallucinations, compromising performance in knowledge-oriented QA, dialogue, and writing. |
| Approach: | They propose to enhance the information searching and reflection ability of large language models by training them in position-agnostic multi-step QA tasks to improve their model's accuracy. |
| Outcome: | The proposed model improves in multi-doc QA and other benchmarks by 13.7% absolute gain in shuffled settings, by 21.5% in passage retrieval task. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) generate responses that are plausible but incorrect or unsupported—commonly referred to as hallucinations. |
| Approach: | They propose a representation-level intervention framework that modulates hallucination-related features during inference by probing their encoded features. |
| Outcome: | The proposed framework reduces hallucinations while maintaining the performance and generalization capabilities of Large Vision-Language Models (LVLMs). |
Copied to clipboard
| Challenge: | Existing methods to enhance credibility and verifiability of large language models (LLMs) mainly focus on passage-level or paragraph-level references or citations, which fall short in verifikatability. |
| Approach: | They propose a method that provides sentence-level citations in LLM-generated responses. |
| Outcome: | The proposed method achieves 90% accuracy in long-form question-answering tasks. |
Copied to clipboard
| Challenge: | Long-context modeling is crucial for next-generation language models, but high computational cost of standard attention mechanisms poses significant computational challenges. |
| Approach: | They propose a natively trained Sparse Attention mechanism that integrates algorithms with hardware-aligned optimizations to achieve efficient long-context modeling. |
| Outcome: | The proposed model maintains or exceeds Full Attention models across general benchmarks, long-context tasks, and instruction-based reasoning. |
Copied to clipboard
| Challenge: | Existing methods for generating explanations for recommender systems produce generic explanations that fail to incorporate user and item specific details. |
| Approach: | They propose a multi-scale distribution deepvariational autoencoder with a prior network that eliminates noise while retaining meaningful signals in the input. |
| Outcome: | The proposed models can generate explanations with concrete input-specific contents. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable reasoning capabilities, but they still face challenges in knowledge-intensive multi-hop reasoning. |
| Approach: | They propose a method that uses self-critique feedback to guide iterative reasoning by enabling iteration and self-evaluation of its intermediate reasoning steps. |
| Outcome: | The proposed method surpasses the previous SOTA by 8.6% on three multi-hop reasoning datasets. |
Copied to clipboard
| Challenge: | Existing methods for document comprehension rely on uniform supervision, resulting in a performance degradation in the intermediate sections. |
| Approach: | They propose a framework driven by Focal Preference Optimization to detect reading order in document layouts. |
| Outcome: | The proposed framework outperforms competing baselines and surpasses large-scale general VLMs. |
Copied to clipboard
| Challenge: | Synthetic data generation is an increasingly popular way of training models without the need for large, manually labeled datasets. |
| Approach: | They propose a framework that aligns open-source small models to efficiently generate large-scale embedding data. |
| Outcome: | The proposed framework outperforms state-of-the-art embedding models by using only 1/10 of the GPT API calls. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) implicitly involves compound semantic annotations. |
| Approach: | They propose to use auxiliary tasks which are semantically or formally related to enhance AMR parsing. |
| Outcome: | The proposed method achieves state-of-the-art performance on benchmarks especially in topology-related scores. |
Copied to clipboard
| Challenge: | Existing summarization methods compress content for gist browsing, but they break prerequisite logic in instructional videos. |
| Approach: | They propose a framework that decouples epistemic planning from content generation. |
| Outcome: | The proposed framework outperforms strong end-to-end baselines on Knowledge Progression Consistency and Learning Objective Coverage. |
Copied to clipboard
| Challenge: | a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models. |
| Approach: | They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models . |
| Outcome: | The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions . |
Copied to clipboard
| Challenge: | Existing methods for obtaining text embeddings require complex training pipelines . authors leverage proprietary LLMs to generate diverse synthetic data for text embeds based on 93 languages . |
| Approach: | They propose a method for obtaining high-quality text embeddings using only synthetic data and less than 1k training steps. |
| Outcome: | The proposed method achieves strong performance on competitive text embedding benchmarks without using any labeled data. |
Copied to clipboard
| Challenge: | CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview. |
| Approach: | They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
| Outcome: | The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
Copied to clipboard
| Challenge: | Existing reasoning-enhanced large language models fail to provide reliable attribution of reasoning behavior once it is transferred through knowledge distillation. |
| Approach: | They propose to embed a reasoning-length gap in a model by querying a target domain and training a local student to imitate its outputs. |
| Outcome: | et al. show that ReasMark outperforms baselines while preserving task utility. |
Copied to clipboard
| Challenge: | Existing knowledge graph embedding models measure the plausibility of triplets either through semantic matching or distance scoring functions. |
| Approach: | They propose to combine semantic matching with entity’s geometric distance to better measure the plausibility of triplets. |
| Outcome: | The proposed model outperforms existing models on well-known knowledge graph completion benchmark datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process. |
| Approach: | They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
| Outcome: | The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
Copied to clipboard
| Challenge: | Existing metrics for multimodal large language models only focus on token overlap and may not align with human judgment. |
| Approach: | They propose an open-source model that assesses the question answering abilities of multimodal large language models. |
| Outcome: | Experiments show that the ACE-M3 model performs better than existing models and is more reliable than existing metrics. |
Copied to clipboard
| Challenge: | Existing methods for QA use knowledge graphs, but they ignore subgraph optimization and subgraph deepening. |
| Approach: | They propose a dynamic heterogeneous-graph reasoning method with LMs and knowledge representation learning that optimizes the structure and knowledge representing of the HKG using a two-stage pruning strategy and knowledge-representation learning. |
| Outcome: | The proposed method improves on existing methods at CommonsenseQA and OpenBookQA. |
Copied to clipboard
| Challenge: | Existing vision-language-action models are unsuitable for simulated or physical-world deployments . current methods fail when confronted with inherent real-world dynamic variability. |
| Approach: | They propose a test-time reinforcement learning framework that enables on-the-fly policy adaptation during inference. |
| Outcome: | Empirical results show that the proposed framework improves adaptability, stability and task success in dynamic, previously unseen scenarios. |
Copied to clipboard
| Challenge: | Teaching large language models to use tools for solving complex problems can grant them human-like reasoning abilities. |
| Approach: | They propose a multi-agent system that enhances the Deep First Search Decision Tree (DFSDT) to address issues like error propagation and limited exploration in ReAct . |
| Outcome: | The proposed system reduces token usage by 60.9% compared to existing methods and performs on par with GPT-4-DFSDT. |
Copied to clipboard
| Challenge: | Existing methods face the trilemma of performance, data privacy, and computational costs, which hinders their application to LLMs. |
| Approach: | They propose a model-exclusive task arithmetic method for merging GPT-scale models which is data-agnostic and bypasses the heavy search process. |
| Outcome: | The proposed method achieves state-of-the-art performance on multiple tasks while minimizing the average loss difference between the merged model and each individual task model. |
Copied to clipboard
| Challenge: | Recent advances in long-context modeling have enhanced language models for complex tasks, but they struggle with multi-hop reasoning and noisy contexts. |
| Approach: | They propose an approach that prompts LMs to supply attributions for each assertion during reasoning. |
| Outcome: | The proposed model achieves competitive performance on multi-hop reasoning benchmarks, closely paralleling proprietary LMs such as ChatGPT and Claude-instant. |
Copied to clipboard
| Challenge: | Visual Language Action models have shown promise in decision-making tasks, but have been neglected in previous work . |
| Approach: | They propose a new paradigm for visual language action models that enhances the foundation model prior to action-specific tuning by first post-training it on a curated set of visual and linguistic tasks using self-supervised learning. |
| Outcome: | The proposed model outperforms the best agent baseline on a diverse set of atomic tasks and surpasses imitation learning-based policies in Minecraft. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) parsing aims to extract an abstract semantic graph from a sentence. |
| Approach: | They propose a new framework that allows for reversed linearization of AMR graphs . they propose to combine sequence-to-sequence approaches with a linearized graph . |
| Outcome: | The proposed framework outperforms the best AMR parser by 0.8 and 0.5 Smatch scores on the AMR 2.0 and AMR 3.0 datasets. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis is challenging because a sentence may contain multiple aspects or complicated relationships. |
| Approach: | They propose a bi-syntax aware Graph Attention Network to model the context of every aspect and sentiment relations across aspects for learning. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on four benchmark datasets. |
Copied to clipboard
| Challenge: | Event extraction is a task in natural language processing that involves identifying and extracting event information from unstructured text. |
| Approach: | They propose a paradigm that combines schema paraphrasing with schema retrieval-augmented generation. |
| Outcome: | The proposed paradigm retrieves paraphrased schemas and accurately generates targeted structures. |
Copied to clipboard
| Challenge: | specialized models fail to detect implicit hate speech due to its indirectly expressed hateful intent . advanced LLMs often misinterpret metaphorical implicit hate content, resulting in its propagation . |
| Approach: | They propose a Jailbreaking strategy and Energy-based Constrained Decoding techniques to detect implicit hate speech in large language models. |
| Outcome: | The proposed model can generate metaphorical implicit hate speech, but it fails to detect it effectively. |
Copied to clipboard
| Challenge: | Multi-tenant Model-as-a-Service (MaaS) workloads exhibit non-stationarity across multiple time scales . existing request schedulers often rely on a fixed policy that remains unchanged at runtime . |
| Approach: | They propose a hierarchical multi-agent scheduler that operates in a layered closed loop . they propose to maintain 1.2–3.0 higher Goodput than SGLang and vLLM . |
| Outcome: | Experiments show that H-MAS achieves 1.2–3.0 higher Goodput than SGLang and vLLM . it maintains more stable QoS under diverse request lengths and heterogeneous SLO targets . |
Copied to clipboard
| Challenge: | Recent approaches to data-to-text generation focus on improving content fidelity, but lack explicit control over writing styles. |
| Approach: | They propose a way to control writing styles by using existing sentences as "soft" templates . they conduct experiments in restaurants and sports domains to test their approach . |
| Outcome: | The proposed approach achieves stronger performance than a range of comparison methods. |
Copied to clipboard
| Challenge: | Existing persona-based dialogue models generate personalized responses using predefined persona information, but they lack personality. |
| Approach: | They propose a persona-based dual Alternating Learning Network that generates personalized responses using predefined persona information. |
| Outcome: | The proposed method produces more personalized responses than baseline methods. |
Copied to clipboard
| Challenge: | a semantic parser converts explanations into programmatic labeling functions . a standard protocol for obtaining a labeled dataset provides only one bit of information per example . |
| Approach: | They propose a framework where an annotator provides an explanation for each labeling decision . they use a semantic parser to convert these explanations into programmatic labeling functions . |
| Outcome: | The proposed framework trains classifiers faster by providing explanations instead of labels . the proposed framework is based on a rule-based semantic parser . |
Copied to clipboard
| Challenge: | Currently, personal AI assistants on the phone and AR glasses can assist our daily life in addressing our questions like "how to adjust the date for this watch?" |
| Approach: | They propose a task that asks a question about affordance of items in our daily life . they construct a dataset that contains 3.2k multimodal questions on 1.6k video segments . |
| Outcome: | The proposed task outperforms baseline methods while still having room for improvement in the future. |
Copied to clipboard
| Challenge: | Large language models extract useful information from conversation history to enhance the response in long-term conversations. |
| Approach: | They propose a Fragment-then-Compose framework to optimize memory utilization for long-term open-domain conversation. |
| Outcome: | The proposed framework can be used to extract useful information from conversation history . it can be adapted to different situations and improve response generation . |
Copied to clipboard
| Challenge: | SimLM uses a simple bottleneck architecture that learns to compress the passage information into a dense vector through self-supervised pre-training. |
| Approach: | They propose a simple yet effective pre-training method for dense passage retrieval that learns to compress the passage information into a dense vector through self-supervised pre-tuning. |
| Outcome: | The proposed method outperforms multi-vector approaches on large-scale passage retrieval datasets and shows significant improvements over baselines. |
Copied to clipboard
| Challenge: | Explicit reasoning trajectories increase performance but often trigger overthinking . despite its importance, this study examines how each step of reasoning affects the final outcome . |
| Approach: | They propose a Reasoning Completion Point Detector that detects the RCP by monitoring rank dynamics of termination tokens. |
| Outcome: | The proposed method reduces token usage by up to 44% while preserving accuracy. |
Copied to clipboard
| Challenge: | Existing methods to predict logical forms ignore the utilization of symbolic operations and lack reasoning ability and interpretability. |
| Approach: | They propose an operation-pivoted discrete reasoning framework that uses symbolic operations as neural modules to facilitate reasoning ability and interpretability. |
| Outcome: | Extensive experiments on DROP and RACENum datasets show the reasoning ability of OPERA. |
Copied to clipboard
| Challenge: | Existing methods for geometry theorem proving in Euclidean geometry are challenging and require a neural network to perform. |
| Approach: | They propose a method for adding auxiliary points in geometry that runs on CPUs without relying on neural network-based inference. |
| Outcome: | The proposed method achieves silver-medal-level human performance on IMO-30 benchmark. |
Copied to clipboard
| Challenge: | Random masking is a widely adopted classic baseline in large language models (LLMs). |
| Approach: | They propose a play-it-by-ear masking performance plug-in which enables LLMs to adaptively select masking target combinations for each task. |
| Outcome: | The proposed performance plug-in retains the advantages and mitigates the drawbacks of random masking in large language models. |
Copied to clipboard
| Challenge: | Existing methods to mitigate Matthew effect in offline recommendation systems are not effective . a number of studies have identified two root causes for the Matthew effect . |
| Approach: | They propose a framework to address the Matthew effect in conversational recommendation systems . they build hypergraphs to learn multi-level user interests to alleviate the Matthew effec . |
| Outcome: | The proposed framework achieves state-of-the-art performance on four CRS-based datasets . it improves on item-, entity-, word-oriented multiple-channel hypergraphs compared with existing methods . |
Copied to clipboard
| Challenge: | Recent advances in recommender systems have been overlooked due to their emphasis on textual content. |
| Approach: | They propose to introduce large language models into recommendation models to exploit the semantic understanding and strong transferability of LLMs. |
| Outcome: | The proposed approach significantly boosts an item’s exposure by altering its textual content during the testing phase, without requiring direct interference with the model’s training process. |
Copied to clipboard
| Challenge: | Existing studies have focused on synthetic supervision but have encountered data quality issues. |
| Approach: | They propose a fully synthetic supervision framework that aims at improving data quality via dual refinement of both tasks and trajectories. |
| Outcome: | The proposed framework outperforms existing methods on standardized benchmarks and shows promising results on a standardized test. |
Copied to clipboard
| Challenge: | In inference-time scaling, Chain-of-Thought (CoT) data is scarce or even unavailable. |
| Approach: | They propose a method which establishes an inference cycle to synthesize user queries and CoT data. |
| Outcome: | The proposed method achieves a 75.4% pass rate and a 79.6% win rate using small models in StableToolBench. |
Copied to clipboard
| Challenge: | Conventional training strategies only consider predefined senses for target words and learn each of them from relatively limited instances, neglecting the influence of similar ones. |
| Approach: | They propose a method to rank senses to improve the task of word Sense Disambiguation (WSD) by ranking an expanded list of sense definitions. |
| Outcome: | The proposed method achieves a SOTA F1 score of 79.6% in Chinese WSD and shows faster convergence than previous methods. |
Copied to clipboard
| Challenge: | Prior work has not explored the mechanisms underlying this sensitivity. |
| Approach: | They propose a synthetic benchmark to evaluate Large Language Models’ reasoning robustness against systematically controlled irrelevant context (IC). |
| Outcome: | The proposed model improves in-distribution and out-of-disttribution scenarios while training with strong distractors. |
Copied to clipboard
| Challenge: | Dialogue policy learning (DPL) aims to determine an abstract representation (also known as action) to guide what the response should be. |
| Approach: | They propose a joint Transformer-based model that generates a token-grained policy that allows more dynamic dialogue action generation without the need for predefined action candidates. |
| Outcome: | The proposed model outperforms existing models showing improvements of 9% and 13% in success rate and 34% and 37% in diversity of dialogue actions across two benchmark dialogue modeling tasks. |
Copied to clipboard
| Challenge: | Modern QA systems entail retrieval-augmented generation (RAG) for accurate and trustworthy responses, but the inherent gap between user queries and relevant documents hinders precise matching. |
| Approach: | They propose a retrieval-augmented generation (RAG)-based approach to bridge this gap by attaching document fingerprints to the embedding to estimate the expectation of potential queries. |
| Outcome: | Experiments across diverse datasets, languages, and embedding models confirm the proposed solution is simple-yet-effective with zero additional index storage, retrieval latency, training costs, or catastrophic forgetting and hallucination issues. |
Copied to clipboard
| Challenge: | Traditional benchmarks for evaluating foundation models often fail to accurately represent their general abilities for human-centric tasks. |
| Approach: | They propose a bilingual benchmark to assess foundation models in the context of human-centric standardized exams such as college entrance exams, law school admission tests, and math competitions. |
| Outcome: | The proposed benchmark exceeds the average human performance on SAT, LSAT, and math competitions with 95% accuracy and 92.5% on the Chinese college entrance English exam. |
Copied to clipboard
| Challenge: | Inductive reasoning is an important task for large language models (LLMs). |
| Approach: | They propose a survey of inductive reasoning for large language models . they categorize methods into three main areas: post-training enhancement, test-time exploration, and data augmentation. |
| Outcome: | The proposed method improves inductive reasoning in large language models. |
Copied to clipboard
| Challenge: | Existing approaches to solve few-shot aspect-based sentiment analysis (ABSA) are suboptimal for this task because of in-context examples . |
| Approach: | They propose to retrieve in-context examples for few-shot aspect-based sentiment analysis . they construct positive and negative pairs from three perspectives and train the retriever . |
| Outcome: | The proposed retrieval framework outperforms baselines on four ABSA datasets. |
Copied to clipboard
| Challenge: | Existing models for semantic sentence matching lack the ability to capture subtle differences. |
| Approach: | They propose to use a Transformer-based pre-trained language model to capture fine-grained differences in sentence pairs by introducing a dual attention module and a fusion module to learn the aggregation of difference and affinity features. |
| Outcome: | The proposed method is able to capture fine-grained differences in sentence pairs. |
Copied to clipboard
| Challenge: | Existing methods for named entity recognition are time-consuming and laborintensive. |
| Approach: | They propose a few-shot multimodal named entity recognition task that uses few examples to locate and identify named entities for a text-image pair. |
| Outcome: | The proposed framework outperforms baselines under several few-shot settings. |
Copied to clipboard
| Challenge: | Existing methods for recursive reasoning are limited due to lack of expert-curated data. |
| Approach: | They propose a method that unlocks the potential of Large Language Models for recursive reasoning through long-form Chain of Thought. |
| Outcome: | The proposed method outperforms preference optimization methods on the openAI o1-series models by 20% on 3k synthetic samples. |
Copied to clipboard
| Challenge: | Large language model editing methods suffer from overfitting, where factual updates can propagate beyond their intended scope, overemphasizing the edited target even when it’s contextually inappropriate. |
| Approach: | They propose a framework for precise and controllable knowledge editing that utilizes two-phase representations and a linear transformation to compute a directional "belief shift" vector. |
| Outcome: | The proposed framework significantly reduces overfitting across nearly all evaluation metrics and on COUNTERFACT and MQuAKE. |
Copied to clipboard
| Challenge: | Cross-lingual summarization is a task of generating a summary in one language for a given document in a different language. |
| Approach: | They present a systematic review of the literature on cross-lingual summarization . they summarize previous efforts and compare them with each other . |
| Outcome: | The proposed approach is compared with previous approaches and summarizes them to provide a deeper analysis. |
Copied to clipboard
| Challenge: | Current approaches use a numeric ID or text piece as the identifier, but these identifieres cannot cover a passage’s content well. |
| Approach: | They propose a new type of identifier that is generated based on the content of a passage and could integrate contextualized information that text pieces lack. |
| Outcome: | The proposed approach performs the best in generative retrieval on three public datasets. |
Copied to clipboard
| Challenge: | Publicly available datasets can be used to evaluate performance of large language models . however, contamination of test data can artificially inflate model performance . |
| Approach: | They propose a Contamination-resilient Evaluation strategy that updates data with real-world knowledge. |
| Outcome: | The proposed evaluation strategy can be used to update datasets with real-world knowledge. |
Copied to clipboard
| Challenge: | a framework to mitigate spurious optimization signals is proposed for test-time reinforcement learning (TTRL) Reinforcement learning with verifiable rewards (RLVR) is an effective paradigm for improving large language models on structured challenging reasoning tasks. |
| Approach: | They propose a framework to mitigate spurious optimization signals from label noise . they propose to use a frequency-based sampling strategy to exclude ambiguous samples . |
| Outcome: | The proposed framework outperforms existing TTRL baselines on three large language models across multiple mathematical reasoning benchmarks. |
Copied to clipboard
| Challenge: | Multi-agent LLMs generate multiple candidate responses that are aggregated by an LLM judge. |
| Approach: | They propose to advocate KV cache reuse across partially shared contexts and report substantial speedups for generation agents. |
| Outcome: | The proposed reuse strategies weaken cross-candidate attention, especially for later candidate blocks, and highlight judge-centric inference as a distinct regime that requires dedicated, risk-aware system design. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) frameworks struggle with identifying whether retrieved documents meaningfully contribute to answer generation. |
| Approach: | They propose a document-related metric to quantify the contribution of retrieved documents to correct answer generation. |
| Outcome: | The proposed framework outperforms existing approaches on both single and multiple retrieval paradigms. |
Copied to clipboard
| Challenge: | Large language models excel at processing unstructured data, but integrating time series data with text remains a challenge. |
| Approach: | They propose a self-supervised multimodal framework that uses prompt-guided learning to unify heterogeneous data types. |
| Outcome: | The proposed framework outperforms state-of-the-art approaches on disease diagnosis tasks using real-world datasets. |
Copied to clipboard
| Challenge: | Multimodal embedding models encode multimedia inputs into latent vector representations. |
| Approach: | They propose to synthesize multimodal multilingual data using a multimodal large language model . they identify three criteria for high-quality synthetic multimodal data . |
| Outcome: | The proposed model outperforms existing models on the MMEB Benchmark and the XTD benchmark. |
Copied to clipboard
| Challenge: | Multi-Modal Knowledge Graphs (MMKGs) are knowledge graphs that integrate and align information from diverse modalities (e.g., text and images). |
| Approach: | They propose a framework that integrates image-text pairs of long-tailed entities and a concept guidance module that offers explainability and enables human verification. |
| Outcome: | The proposed framework improves the accuracy of recognizing long-tailed image-text pairs compared to baselines and also offers flexibility and explainability. |
Copied to clipboard
| Challenge: | Contrastive Language-Image Pre-training (CLIP) is a standard for cross-modal image-text representation learning. |
| Approach: | They propose a framework that enhances pre-trained CLIP models by exploiting challenging text-image pairs within existing datasets. |
| Outcome: | The proposed framework improves CLIP models by exploiting text-image pairs in training. |
Copied to clipboard
| Challenge: | Existing datasets involve translation, but translationese is distinguished from original text . previous studies have shown that translationeses in CLS are not a problem in training sets . |
| Approach: | They propose to use cross-lingual summarization to generate a concise summary in a target language from a document in . existing datasets typically involve translation in their creation, but the translated text is distinguished from the original written in that language. |
| Outcome: | The proposed method systematically investigates how translationese affects CLS model evaluation and performance when it appears in source documents or target summaries. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have enabled social simulation through multi-agent systems. |
| Approach: | They propose a system for constructing and simulating book-based multi-agent societies that simulates established fictional worlds and characters. |
| Outcome: | The proposed system generates high-quality stories while maintaining fidelity to the source books, surpassing previous methods with a win rate of 75.36%. |
Copied to clipboard
| Challenge: | Existing approaches to cross-document relation extraction (RE) focus on identifying relations between head and tail entities from single sentence or document. |
| Approach: | They propose a hierarchical relation tree-based LLM-based hierarchic classification model for cross-document relation extraction (HCRE) based on predefined relations, the model can perform hierarchically classification level by level. |
| Outcome: | The proposed model outperforms existing baselines and validates its effectiveness. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) scales capacity via conditional computation, but lacks knowledge lookup primitive. |
| Approach: | They propose a conditional memory instantiated via Deep Sparse Embedding (DSE) they propose 'u-shaped scaling law' that identifies optimal balance between MoE experts and DSE memory . |
| Outcome: | The proposed model outperforms an iso-parameter and isoFLOPs MoE baseline across knowledge and reasoning benchmarks and is infrastructure-efficient. |
Copied to clipboard
| Challenge: | Currently, the top-performing models achieve a 48.8% task completion rate on realizing machine learning algorithms . |
| Approach: | They propose a benchmark to test machine learning's ability to generate ML code for humans . they propose an automatic evaluation framework with metrics such as task pass rate and time overhead . |
| Outcome: | The proposed benchmark is unique in its focus on interpreting complex human instructions and producing multi-step, high-complexity code. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown remarkable results on various NLP tasks. |
| Approach: | They propose to improve the feed-forward network (FFN) in BERT with a higher computational cost than improving the multi-head attention (MHA). |
| Outcome: | The proposed model is 6.9 smaller and 4.4 faster than BERTBASE and has competitive performances on GLUE and SQuAD Benchmarks. |
Copied to clipboard
| Challenge: | Existing studies indicate that System-2 thinking alone does not transfer to the general alignment domain. |
| Approach: | They propose to use priority-aware metacognition to help LRMs understand human preferences and monitor and regulate their thinking process. |
| Outcome: | The proposed model improves general alignment performance by 10 points on helpfulness and harmless benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown promising results in various domains, but their practical application in industry-relevant operations research presents significant challenges and opportunities. |
| Approach: | They propose a cognitive-inspired framework that enhances optimization through counterfactual reasoning . they use a workflow that transforms requirements into mathematical models and executable solver code . |
| Outcome: | Experiments show that ORMind outperforms existing methods in the NL4Opt dataset and ComplexOR dataset. |
Copied to clipboard
| Challenge: | a new generation of (M)LLMs is enabling the creation of superintelligent AI assistants . OS Agents can complete tasks autonomously and have the potential to significantly enhance the lives of billions of users worldwide. |
| Approach: | They propose to build OS Agents that operate within operating systems' GUIs and GUIs . they examine evaluation metrics and benchmarks to identify promising directions . |
| Outcome: | The proposed agents are based on operating systems (OS) and operating systems frameworks. |
Copied to clipboard
| Challenge: | Existing work shows that LLMs rely on single-paradigm reasoning that limits their effectiveness across diverse tasks. |
| Approach: | They propose a new framework that integrates multiple reasoning paradigms to enable synergistic collaboration. |
| Outcome: | The proposed model outperforms current SOTA models in theorem proving tasks and the MATH benchmark in arithmetic tasks. |
Copied to clipboard
| Challenge: | PoDA pre-trains encoders and decoders by denoising noise-corrupted text . Unlike encoder-only or decode-only methods, it can be used for text generation tasks without using any task-specific techniques. |
| Approach: | They propose a sequence-to-sequence (seq2sequ) pre-training method PoDA which denoises autoencoders by denoising noise-corrupted text. |
| Outcome: | The proposed method improves model performance over strong baselines without using any task-specific techniques and significantly speed up convergence. |
Copied to clipboard
| Challenge: | Existing prompt learning models for IDRR use multiple-prompt decisions from three different yet much similar connective prediction templates. |
| Approach: | They propose to fuse three related tasks to fuse the learned features of auxiliary tasks to create a prompt learning model that can be used to boost the main task. |
| Outcome: | The proposed model outperforms the ConnPrompt in the training phase and in the testing phase. |
Copied to clipboard
| Challenge: | Existing embedding models support only 512 input tokens, hindering their application in scenarios requiring long inputs. |
| Approach: | They evaluate the performance of existing embedding models by using a new benchmark and a training-free context window extension strategy. |
| Outcome: | The proposed model extends the input window of existing models by several folds. |
Copied to clipboard
| Challenge: | Existing methods for open relation extraction (OpenRE) focus on labeled and pre-defined instances, which are costly to acquire in reality. |
| Approach: | They propose a framework that can extract relations without pre-defined types from open-domain corpus with efficient knowledge transfer from a few pre-determined relational instances. |
| Outcome: | The proposed framework achieves the new SOTA results for OpenRE on different datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for Continual Language Learning (CLL) are limited due to the complexity of the task and the lack of unified benchmarks. |
| Approach: | They propose a Continual Language Learning Evaluation benchmark CLLE in multilingual translation. |
| Outcome: | The proposed method is effective when compared with other strong benchmarks. |
Copied to clipboard
| Challenge: | Modern large language models (LLMs) have shown remarkable performance on general language tasks but struggle on complex reasoning tasks. |
| Approach: | They propose a multi-agent debate framework that encourages divergent thinking in LLMs . they propose to break debate and use a judge to obtain a final solution . |
| Outcome: | The proposed framework encourages divergent thinking in large language models . it is able to generate novel thoughts even if initial stance is incorrect . |
Copied to clipboard
| Challenge: | Recent studies have focused on developing persona consistent dialogue models . order sensitivity affects the quality and consistency of generated response . |
| Approach: | They propose a model-agnostic framework to improve persona consistent dialogue response generation by concatenating persona texts and dialogue history as a single input sequence. |
| Outcome: | The proposed framework outperforms existing models on the Persona-Chat dataset and shows that it is more robust under different persona orders and more consistent with the persona profile. |
Copied to clipboard
| Challenge: | End-to-end spoken dialogue models have higher potential ceiling in expressiveness and perceptual ability than cascaded systems. |
| Approach: | They propose a modality-aware adaptive post-training recipe that constrains preference updates to the semantic channel and improves acoustic behavior via explicit anchoring. |
| Outcome: | The proposed model improves speech quality and expressiveness across spoken dialogue benchmarks and architectures. |
Copied to clipboard
| Challenge: | Fill-in-the-Middle (FIM) models suffer from performance degradation and prohibitive latency. |
| Approach: | They propose a search-and-replace infilling framework that integrates agentic verification and editing into a single-pass inference process. |
| Outcome: | The proposed framework harmonizes completion tasks with the instruction-following priors of Chat LLMs, extending the paradigm from static infilling to dynamic context-aware editing. |
Copied to clipboard
| Challenge: | Existing KV cache compression methods enforce a fixed pattern, neglecting task-specific characteristics, which hampers the effective retention of essential information while discarding less important tokens. |
| Approach: | They propose a Task-Aware KV cache mechanism that dynamically adjusts the KV caching size across different layers based on the characteristics of the tasks. |
| Outcome: | The proposed method surpasses state-of-the-art methods by 11% on the LongBench dataset even under extreme compression (0.9%) |
Copied to clipboard
| Challenge: | APEX optimizes for text-to-image generation by combining learning potential, conflict penalty, and progress need. |
| Approach: | They propose an algorithm that stabilizes heterogeneous rewards and dynamically schedules objectives . they propose a method that achieves better Pareto trade-offs across four heterogenous objectives based on P3 Adaptive Priorities . |
| Outcome: | The proposed algorithm achieves better pareto trade-offs across four heterogeneous objectives while maintaining competitive OCR accuracy. |
Copied to clipboard
| Challenge: | Existing work has investigated the title optimization for Product Listing Ads (PLAs) however, little work has examined the effectiveness of this method. |
| Approach: | They propose a method to rewrite product listing ads titles without considering the fluency and information priority. |
| Outcome: | The proposed solution reduces the cost and improves CTR in the offline test and real-world online test by a large amount. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) methods are efficient for a large language model with reduced computational costs. |
| Approach: | They propose a layer-wise expert numbers and ranks allocation strategy with GuidedSelection Vectors. |
| Outcome: | The proposed method achieves superior or comparable performance to all baselines on three backbone models. |
Copied to clipboard
| Challenge: | Existing methods for sparse and dense retrieval have limited success on popular datasets. |
| Approach: | They propose a query expansion approach that generates pseudo-documents by few-shot prompting large language models and then expands the query with generated pseudo-docs. |
| Outcome: | The proposed method boosts the performance of BM25 on ad-hoc IR datasets by 3% to 15% without any model fine-tuning. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a powerful tool for test-time scaling, but they are often used under time constraints. |
| Approach: | They propose to use LLMs to make models think before answering questions . they also use self-correction and best-of-N decoding to encourage deeper thinking . |
| Outcome: | The proposed models are able to achieve higher inference accuracy with extra inference computation under time constraints. |
Copied to clipboard
| Challenge: | Existing Text-to-SQL research focuses on specific database systems, limiting adaptability to different dialects. |
| Approach: | They propose a framework that employs Object Relational Mapping (ORM) code as an intermediate language to bridge this gap. |
| Outcome: | The proposed framework outperforms existing methods that generate SQL queries directly. |
Copied to clipboard
| Challenge: | Existing methods for automatic text dating overlook the evolution of word meanings in texts spanning long periods. |
| Approach: | They propose a temporal-aware text representation that dynamically captures both semantic variance and invariance. |
| Outcome: | The proposed approach outperforms existing methods on two diachronic datasets. |
Copied to clipboard
| Challenge: | Existing script event prediction task forcasts the subsequent event based on an event script chain, but the evolution of historical events is more complicated in real world scenarios. |
| Approach: | They propose a Causality Graph Event Prediction task that forecasts consequential event based on an Event Causity Graph (ECG). |
| Outcome: | The proposed model outperforms the advanced competitors for the CGEP task. |
Copied to clipboard
| Challenge: | Existing unlearning methods suffer from a geometric mismatch, causing catastrophic forgetting or unsafe substitution. |
| Approach: | They propose a framework for surgical semantic pruning within the Lorentz manifold. |
| Outcome: | Experiments on MLLMU-Bench show that LOTUS significantly outperforms baselines while maintaining general utility. |
Copied to clipboard
| Challenge: | Existing methods for generating responses following a desired style are lacking of parallel data for training. |
| Approach: | They propose a KL loss and a style classifier to fine-tune response generation . they show that their model can significantly outperform state-of-the-art methods . |
| Outcome: | The proposed model outperforms state-of-the-art models in style consistency and contextual coherence with two public datasets. |
Copied to clipboard
| Challenge: | Existing datasets for Chinese instruction tuning are not well-aligned with Chinese users’ interaction patterns. |
| Approach: | They propose to use Chinese instruction tuning datasets to improve instruction fine-tuning for Chinese users. |
| Outcome: | The proposed dataset shows that Chinese models achieve competitive performance in diverse benchmarks. |
Copied to clipboard
| Challenge: | Existing rationalization approaches are susceptible to degeneration due to lack of effective control over the learning direction of the model during training. |
| Approach: | They propose an agent-guided rationalization approach that guides the next step of the model based on its current training state. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on BeerAdvocate and HotelReview datasets. |
Copied to clipboard
| Challenge: | Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored. |
| Approach: | They propose to use a benchmark to evaluate large language models' financial domain knowledge and practical abilities. |
| Outcome: | The proposed benchmark evaluates large language models' financial domain knowledge and practical abilities. |
Copied to clipboard
| Challenge: | Existing approaches to improve in-context learning performance are highly sensitive to the quality of the incontext examples provided. |
| Approach: | They propose a framework to iteratively train dense retrievers that can identify high-quality in-context examples for large language models. |
| Outcome: | The proposed model improves performance by retrieving examples with similar patterns, and the gains are consistent across LLMs of varying sizes. |
Copied to clipboard
| Challenge: | Existing multi-domain neural machine translation models lack adaptation to individual domains. |
| Approach: | They propose a multi-domain neural machine translation model with individual modules for each domain . they use word-level, adaptive and layer-wise domain mixing to achieve this . |
| Outcome: | The proposed model outperforms existing models in several NMT tasks. |
Copied to clipboard
| Challenge: | Existing approaches to embed multimodal models face limitations such as suboptimal causal attention in VLMs and limited diversity in training objectives and data. |
| Approach: | They propose a framework for transforming pre-trained VLMs into bidirectional multimodal embedding models. |
| Outcome: | The proposed model improves performance across MMEB and ViDoRe-v2 benchmarks and exhibits strong scalability with both model size and training data on MMEF. |
Copied to clipboard
| Challenge: | Using Large Language Models (LLMs)-based agents can enhance their understanding of environments and tasks. |
| Approach: | They propose a framework that allows agents to synthesize possible scenarios with multi-step action invocation within the action space and perform Monte Carlo Tree Search exploration to refine their action knowledge in the current environment. |
| Outcome: | The proposed framework synthesizes possible scenarios with multi-step action invocation within the action space and performs Monte Carlo Tree Search exploration to refine action knowledge in the current environment. |
Copied to clipboard
| Challenge: | Several perspectives of robustness for pre-trained language models have been studied independently, but lacking a unified consideration in multiple perspectives. |
| Approach: | They propose a technique to enhance the multi-perspective robustness of LMs by introducing adversarial perturbation while the model parameters are selectively updated upon their relative importance. |
| Outcome: | The proposed technique improves the robustness of LMs by incorporating four perspectives on model robustness. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) have demonstrated impressive performance on general video comprehension benchmarks, but their robustness needs to be thoroughly investigated for broader applications. |
| Approach: | They propose a temporal robustness benchmark which introduces temporal inconsistency perturbations separately at the visual and textual modalities to assess the robustness of models. |
| Outcome: | The proposed method improves the model’s robustness and reliability in temporal analysis. |
Copied to clipboard
| Challenge: | Existing methods to learn behavioral sequences fail to capture complex behavioral patterns due to a lack of deep reasoning capabilities and world knowledge. |
| Approach: | They propose a framework that integrates the reasoning power of Large Language Models with the sequential modeling strengths of traditional KT methods via multi-level plug-and-play alignment. |
| Outcome: | Extensive experiments on four standard datasets show that the proposed framework outperforms existing methods on state-of-the-art questions. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have expanded to more complex repository-level tasks. |
| Approach: | They propose a first approach to leveraging visual data to enhance the issue-resolving capabilities of Large Language Models (LLMs) they demonstrate the effectiveness of CodeV and provide valuable insights into leveraging visualization to resolve GitHub issues. |
| Outcome: | The proposed approach improves the issue-resolving capabilities of Large Language Models (LLMs) by using visual data. |
Copied to clipboard
| Challenge: | Existing methods to predict missing triples in Knowledge Graphs are limited by semantic information. |
| Approach: | They propose a method to leverage sememe knowledge to enhance LP . LP is a technique that integrates structural and textual information into a Knowledge Graph . |
| Outcome: | The proposed method improves LP performance in English and Chinese . it improves on WN18RR, HN7 and CWN5, respectively . |
Copied to clipboard
| Challenge: | Existing approaches to encoding compositional generalization are lacking . et al., 2017) argue that neural networks lack compositional ability . |
| Approach: | They propose a method to encode compositionality in neural networks using two representations . they reduce the entropy in each representation to improve generalization . |
| Outcome: | The proposed approach improves performance on five NLP tasks including instruction learning and machine translation. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue generation models struggle with dull and repetitive outputs, a problem commonly termed as text degeneration. |
| Approach: | They propose a framework that allows the model to "cheat" the objective by duplicating knowledge segments in a superficial pattern matching based on overlap. |
| Outcome: | The proposed framework can be applied to a WoW dataset and shows that it works across models and decoding strategies. |
Copied to clipboard
| Challenge: | Existing studies have shown that non-autoregressive translation models can predict all tokens independently and simultaneously. |
| Approach: | They propose to enhance signals of neighbour source tokens into conventional cross-attention to address a locality perception problem in NAT cross- attention. |
| Outcome: | The proposed approach improves translation quality over strong NAT baselines on representative datasets. |
Copied to clipboard
| Challenge: | Existing controllable Text-to-Speech methods limited to inter-utterance-level control . utterance expressiveness remains a challenge in building human-like TTS synthesis systems . |
| Approach: | They propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. |
| Outcome: | The proposed framework achieves state-of-the-art intra-utterance consistency while maintaining baseline-level speech quality. |
Copied to clipboard
| Challenge: | despite significant strides in multimodal tasks, MLLMs are plagued by the critical issue of hallucination. |
| Approach: | They propose a meta-evaluation benchmark to facilitate evaluation of advancements in hallucination detection methods. |
| Outcome: | The proposed framework validates hallucinations robustly and provides strategic insights . MHaluBench is a meta-evaluation benchmark designed to facilitate evaluation . |
Copied to clipboard
| Challenge: | Neural machine translation systems often produce inadequate translations for named entities. |
| Approach: | They propose a data augmentation strategy to enhance the accuracy of named entity translation by retraining the target named entity pair. |
| Outcome: | The proposed method improves translation accuracy across test sets and terminology tests. |
Copied to clipboard
| Challenge: | Large language models (LLMs) augmented with retrieval systems have significantly advanced natural language processing tasks by integrating external knowledge sources. |
| Approach: | They propose a method that conditions large language models to generate answers even in the absence of reliable knowledge. |
| Outcome: | The proposed approach balances accuracy with appropriate abstention, enhancing the reliability and trustworthiness of retrieval-augmented systems. |
Copied to clipboard
| Challenge: | Sparse autoencoders (SAEs) extract interpretable and monosemantic features in large language models . prior work focused on feature extraction from a single layer, failing to capture activations that span multiple layers. |
| Approach: | They propose a framework that integrates a routing mechanism with a shared SAE to efficiently extract features from multiple layers. |
| Outcome: | The proposed framework extracts features from multiple layers while incurring minimal parameter overhead while achieving high interpretability and flexibility. |
Copied to clipboard
| Challenge: | Existing graph-based methods for text classification cannot capture contextual word relationships within each document nor can they produce inductive learning of new words. |
| Approach: | They propose to use Graph Neural Networks to learn the local word representations and then aggregate the word nodes as the document embeddings. |
| Outcome: | The proposed method outperforms state-of-the-art methods on four benchmark datasets. |
Copied to clipboard
| Challenge: | Recent studies have shown the effectiveness of long chain-of-thought (CoT) in reasoning tasks such as math and coding tasks. |
| Approach: | They propose to use Qwen2.5 and LLama-3.1 as backbones to train long thought models to bring the success of long chain-of-thought (CoT) to neural machine translation. |
| Outcome: | The proposed model outperforms vanilla LLMs and LLM models which are fine-tuning on paired sentences without long thought and outperformed vanilla LRMs. |
Copied to clipboard
| Challenge: | Extensive experimental results on several popular logical benchmarks (ProofWriter, PrOntoQA, PrONtoQA-OOD, and FOLIO) and mathematical benchmark (DI-GSM) show that COP significantly outperforms previous state-of-the-art methods. |
| Approach: | They propose a reasoning approach called Concise and Organized Perception (COP) that carefully analyzes the given statements to identify the most pertinent information while eliminating redundancy efficiently. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on several popular logical benchmarks and mathematical benchmarks. |
Copied to clipboard
| Challenge: | Pre-trained protein language models have been used in protein engineering, but their adoption is limited due to data collection, task benchmarking, and application challenges. |
| Approach: | They propose a versatile engine that integrates biological data retrieval, standardized task benchmarking, and modular fine-tuning of PLMs. |
| Outcome: | The proposed engine integrates biological data retrieval, task benchmarking, and modular fine-tuning of PLMs. |
Copied to clipboard
| Challenge: | Existing approaches enhance reasoning through Chain-of-Thought, Program-ofThough, and Tool-Integration. |
| Approach: | They propose a tool-awareness training method that leverages both forward and backward data generation strategies to strengthen the model’s conscious and selective tool utilization in multi-step reasoning tasks. |
| Outcome: | The proposed method improves the model's tool utilization capabilities, including proactivity and execution success rates. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) achieve excellent performance through pretraining on extensive data. |
| Approach: | They propose an efficient selective layer intervention based on parameter-efficient fine-tuning methods to select the optimal steering layer to modulate LLM semantics. |
| Outcome: | The proposed approach is based on a model-agnostic framework and is safe to deploy. |
Copied to clipboard
| Challenge: | Existing approaches to integrating external knowledge into large language models (LLMs) however, the incorporation of external knowledge increases the vulnerability of LLMs . |
| Approach: | They propose a benchmark to evaluate the RAG security using a dataset . they classify attack tasks into silver noise, inter-context conflict, soft ad, and white Denial-of-Service . |
| Outcome: | The proposed benchmark evaluates the security of RAG against 14 representative RAG components. |
Copied to clipboard
| Challenge: | Distantly supervised relation extraction (DSRE) methods are not capable of extracting relation labels for individual sentences. |
| Approach: | They propose a semi-supervised learning relation extraction framework for sentence-level DSRE . they discard only the labels of the noisy samples and utilize them as unlabeled samples . |
| Outcome: | The proposed framework achieves significant performance enhancements on two real-world datasets. |
Copied to clipboard
| Challenge: | Recent LLM-based Text-to-SQL methods suffer from performance degradation on “huge” databases and complex user questions that require multi-step reasoning. |
| Approach: | They propose a framework that integrates a decomposer agent and auxiliary agents to generate SQL queries from natural language text. |
| Outcome: | The proposed framework achieves comparable execution accuracy on SQL-Llama tasks compared to the baseline model. |
Copied to clipboard
| Challenge: | Recent studies have shifted paradigms and leveraged Large Language Models (LLMs) to tackle the challenging task of Text-to-SQL. |
| Approach: | They propose a framework that leverages large language models to generate SQL queries . they exploit prior knowledge from the LLM to enhance embedding-based retriever . |
| Outcome: | The proposed method improves embedding-based retriever and reduces cost. |
Copied to clipboard
| Challenge: | Existing work models time implicitly, making it difficult to handle complex relationships . a novel temporal reasoning framework explicitly models the temporal relationships among facts by multi-view temporal graphs . |
| Approach: | They propose a multi-view temporal graph-based temporal reasoning framework that explicitly models the temporal relationships among facts by multi-visit temporal charts. |
| Outcome: | The proposed framework gives more consistent answers under question perturbations. |
Copied to clipboard
| Challenge: | supervised fine-tuning (SFT) is crucial for multimodal large language models, yet a comprehensive scaling law is lacking . et al.: scaling laws focus on model size, pre-training tokens, and MLLM SFT data volumes . |
| Approach: | They propose two scaling laws to guide optimal model-data configuration . they propose one applicable when training data volumes are well defined by researchers . |
| Outcome: | The proposed scaling laws provide valuable recommendations for optimal resource allocation . they show that the proposed laws are more accurate than existing models . |
Copied to clipboard
| Challenge: | Position encoding (PE) is used to preserve word order information for natural language processing tasks, generating fixed position indices for input sequences. |
| Approach: | They propose to augment SANs with cross-lingual position representations to model bilingually aware latent structure for the input sentence. |
| Outcome: | The proposed model significantly improves translation quality over baselines on EnglishGerman, JapaneseEnglish, and ChineseEnglish translation tasks. |
Copied to clipboard
| Challenge: | Recent research has demonstrated the potential of using LLMs to develop role-playing language agents (RPLAs) however, imitative decision-making necessitates a more nuanced understanding of personas. |
| Approach: | They propose a method that uses persona-based memory retrieval to improve RPLAs. |
| Outcome: | The proposed method significantly advances RPLAs on this task. |
Copied to clipboard
| Challenge: | Recent studies have shown that cloze-style reading comprehension is a popular task for measuring the progress of natural language understanding. |
| Approach: | They propose a multi-perspective framework which can be seen as joint training of heterogeneous experts and aggregate context information from different perspectives. |
| Outcome: | The proposed framework achieves new state-of-the-art over previous strong baselines on a recently released cloze-test dataset. |
Copied to clipboard
| Challenge: | Chinese medical large language models (LLMs) are underperforming on this benchmark, especially where medical reasoning and factual consistency are vital. |
| Approach: | They propose a benchmark with 14 expert-guided clinical scenarios to assess the medical ability of large language models across 7 pivot dimensions. |
| Outcome: | The proposed benchmark has been validated in several ways. |
Copied to clipboard
| Challenge: | Existing methods that confuse tool utilization with knowledge reasoning harm readability and give rise to tool invocation hallucinations. |
| Approach: | They propose to decouple LLM from tool invocation tasks by establishing a memory module with explicit descriptions of query statements and a query memory module to facilitate the KGQA process. |
| Outcome: | The proposed method achieves state-of-the-art on WebQSP and CWQ benchmarks. |
Copied to clipboard
| Challenge: | Large language models generate biased stances due to spurious correlations and preference towards certain individuals and topics. |
| Approach: | They propose a counterfactual Augmented Calibration Network to calibrate potential bias in stance detection of large language models. |
| Outcome: | The proposed calibration network can mitigate biases of large language models, achieving state-of-the-art results. |
Copied to clipboard
| Challenge: | Existing parallel code localization agents suffer from a 34.9% redundant tool invocation rate . specialized localization agent that operate as dedicated search components is needed to achieve high localization accuracy. |
| Approach: | They propose a parallel code localization system that reframes parallel code execution as a quality–efficiency co-optimization problem. |
| Outcome: | The proposed method matches SOTA performance while being 93.6% faster. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is a key task in vehicular systems. |
| Approach: | They propose a benchmark that encompasses diverse automotive scenarios . they use images from front, side, and rear cameras, various road types, weather conditions, and interior views . |
| Outcome: | The proposed benchmark includes images from front, side, and rear cameras, various road types, weather conditions, and interior views. |
Copied to clipboard
| Challenge: | Existing model editing methods focus on single-round editing and often face significant challenges in sequential model editing. |
| Approach: | They propose a model editing method that optimizes the target layer’s hidden states using the model’s original weights to prevent model failure. |
| Outcome: | The proposed method outperforms existing model editing methods and is available on the open-source platform 4open.science. |
Copied to clipboard
| Challenge: | Homographic puns have a long history in human writing, widely used in written and spoken literature, which intended as jokes. |
| Approach: | They propose a WordNet-encoded model to settle polysemy of homographic puns and a word weighted model for recognizing them. |
| Outcome: | The proposed model can distinguish between homographic pun and non-homographic pun texts. |