Papers by Yang Li
Copied to clipboard
| Challenge: | Existing approaches to multi-agent problem solving rely on hand-crafted protocols or automatically designed topologies. |
| Approach: | They propose a state-driven framework that formulates multi-agent problem solving as a finite-state execution process. |
| Outcome: | The proposed framework outperforms baselines on diverse benchmarks by 6.74%–19.39% while reducing token consumption. |
Copied to clipboard
| Challenge: | Existing large language model (LLM) agents are unable to adapt to changing domain knowledge and rules. |
| Approach: | They propose an LLM agent framework that continuously learns updated domain knowledge at test time. |
| Outcome: | The proposed agent improves on a customer due diligence name screening task on . the agent learns updated domain knowledge at test time. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can solve reasoning and mathematical problems using the Chain-of-Thought technique, but require costly and long CoT data and fine-tuning. |
| Approach: | They propose a method that uses Sparse Autoencoders to extract interpretable features from vanilla CoT and use them to steer the LLM's internal states. |
| Outcome: | The proposed method uses Sparse Autoencoders (SAEs) to extract interpretable features from vanilla CoT and steer the LLM's internal states during generation. |
Copied to clipboard
| Challenge: | Existing work on robustness in text or code tasks has focused on classification, while robustness for code generation tasks is an uncharted area. |
| Approach: | They propose a robustness evaluation benchmark for code generation models that customizes over 30 transformations specifically for code on docstrings, function and variable names, code syntax, and code format. |
| Outcome: | The proposed model performs better on human annotators and on SOTA models with human annnotators. |
Copied to clipboard
| Challenge: | Large Language Models are increasingly utilized as role-playing agents to simulate personas in interactive settings. |
| Approach: | They propose a role-playing agent trained to explicitly ground responses in individual identity. |
| Outcome: | The proposed agent can generate persona-consistent responses in long-context dialogues while maintaining general instruction-following capabilities. |
Copied to clipboard
| Challenge: | Existing benchmarks focus more on end-to-end performance, but neglect the underlying principles of knowledge acquisition and generalization. |
| Approach: | They propose a benchmark specifically designed to explore the problem-solving principles by decomposing 6.5K visual math problems into 10.9K step-level questions for evaluation. |
| Outcome: | The proposed benchmark covers 6.5K visual math problems and 10.9K step-level questions spanning 5 layers of knowledge granularity and 67 hierarchical knowledge concepts. |
Copied to clipboard
| Challenge: | Existing systems cannot automatically determine when to switch between modes based on text content. |
| Approach: | They propose a unified framework that implicitly infers vocal modes from text context to pioneer SCS Synthesis. |
| Outcome: | The proposed framework infers vocal modes solely from text context to pioneer SCS Synthesis. |
Copied to clipboard
| Challenge: | Existing methods focus on replicating dialogues in textual form, neglecting the role’s voice traits as a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. |
| Approach: | They propose a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. |
| Outcome: | The proposed model exhibits role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. |
Copied to clipboard
| Challenge: | Recent advances in retrieval-augmented generation (RAG) have substantially improved question-answering systems, particularly for factoid ‘5Ws’ questions. |
| Approach: | They propose a data organization paradigm where large language models transform documents into more structured and loosely interconnected LUs. |
| Outcome: | Experiments in open-domain and industrial settings show that the proposed paradigm outperforms existing paradigms and shows high adaptability across diverse document formats. |
Copied to clipboard
| Challenge: | Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, but academic research remains non-reproducible due to the lack of publicly available training data. |
| Approach: | They propose a system for long-form song generation with fine-grained style conditioning that includes a licensed synthetic dataset and a song generation model, Muse. |
| Outcome: | The proposed system achieves competitive performance on phoneme error rate, text–music style similarity, and audio aesthetic quality while enabling controllable segment-level generation across different musical structures. |
Copied to clipboard
| Challenge: | Large Language Models struggle with complex, multi-step operational tasks because they remain static during inference and cannot learn from past experience. |
| Approach: | They propose a framework that organizes cross-domain insights to facilitate orchestration of long-horizon workflows. |
| Outcome: | The proposed framework outperforms existing methods on the TAC productivity benchmark and shows strong cross-task transferability. |
Copied to clipboard
| Challenge: | Lexicalized parsing has traditionally been neglected in favor of unlexicalized, span-based methods. |
| Approach: | They propose a latent lexicalization framework that dynamically infers lexicals from data without relying on predefined head-finding rules. |
| Outcome: | The proposed model learns lexical dependencies directly from data, offering greater adaptability across languages and datasets. |
Copied to clipboard
| Challenge: | Topic modeling (TM) is a classic unsupervised learning task in the field of natural language processing. |
| Approach: | They propose a new taxonomy that emphasizes the role of LLMs and the design of end-to-end workflows. |
| Outcome: | The proposed taxonomy emphasizes the role of LLMs and the design of end-to-end workflows. |
Copied to clipboard
| Challenge: | Existing research has demonstrated that the ability of large language models (LLMs) to generate humorous sentences is limited to producing 25 unique jokes. |
| Approach: | They propose a multi-stage curriculum preference learning framework to optimize both pun structure preferences and humor preferences by a Chinese Pun dataset. |
| Outcome: | The proposed method significantly outperforms baseline models on Chinese and English benchmark datasets. |
Copied to clipboard
| Challenge: | Large language models are trained on static corpora but deployed in a dynamic world . a foundational tension remains between time and the ability to understand it . |
| Approach: | They formalize temporal queries in an information-theoretic framework based on parametric reachability of temporal premises and answers. |
| Outcome: | The proposed framework formalizes temporal queries in an information-theoretic framework based on parametric reachability of temporal premises and answers . the framework induces four temporal information regimes corresponding to internal reasoning, answer recency, premise anchoring, and genuine world indeterminacy . |
Copied to clipboard
| Challenge: | Existing vision-language models lack spatial reasoning capability, despite their ability to comprehend spatial arrangements and model structural relations. |
| Approach: | They propose a benchmark to evaluate vision-language models' spatial perception, structural understanding, and reasoning capabilities by minimizing reliance on domain-specific knowledge. |
| Outcome: | The proposed benchmark is based on 1,100 carefully curated real-world images with high spatial complexity. |
Copied to clipboard
| Challenge: | Existing methods to detect large language models (LLMs) use binary or ternary classifications, which can only distinguish pure human/LLM text or collaborative text at best. |
| Approach: | They propose a fine-grained method that characterizes distinct signatures of creator and editor by using Rhetorical Structure Theory to construct a logic graph for creator's foundation and extracting Elementary Discourse Unit (EDU)-level features for the editor's style. |
| Outcome: | The proposed method outperforms 12 baselines in identifying fine-grained types with low false alarms, offering a policy-aligned solution for LLM regulation. |
Copied to clipboard
| Challenge: | Inductive reasoning is a core component of human intelligence. |
| Approach: | They propose a task to induce natural language rules from natural language facts using natural language as representation for knowledge instead of formal language. |
| Outcome: | The proposed task surpasses baselines in both automatic and human evaluations. |
Copied to clipboard
| Challenge: | Existing methods for product attribute value extraction are noisy and incomplete with missing values for most retailers. |
| Approach: | They propose a Structure Mltimodal trAnsformeR for producT Attribute Value Extraction which jointly encodes the structured product information from multiple modalities. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two multimodal product datasets. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have advanced beyond traditional Large Language Models, yet they pose heightened safety risks. |
| Approach: | They propose a first jailbreak attack targeting Large Reasoning Models . they exploit a Chaos Machine component to transform attack prompts with diverse one-to-one mappings based on the reasoning chain . |
| Outcome: | The proposed attack exploits the unique vulnerabilities of LRMs by integrating a Chaos Machine. success rates of the mousetrap attack are as high as 96%, 86% and 98% respectively. |
Copied to clipboard
| Challenge: | Recent research has focused on pushing weight-only quantization to extremely low-bit due to numerical representation limitations. |
| Approach: | They propose a vector-based quantization approach that pushes LLMs to extremely low-bit . they propose scalar-based weight quantization that reduces memory requirements and optimizes storage costs . |
| Outcome: | The proposed method reduces model quantization perplexity by 0.01-0.34 on LLaMA-2, 0.38-0.68 on mistral-7B, 4.41-7.34, on llaMA-3 on QA tasks on average. |
Copied to clipboard
| Challenge: | Recent studies show that Large Language Models (LLMs) have shown remarkable intelligence in question answering. |
| Approach: | They propose to reframe the Question Answering task as Programming to overcome this limitation by leveraging LLMs' superior ability in understanding both natural language and programming language. |
| Outcome: | The proposed approach improves on time-sensitive question answering datasets by 14.5% over baselines. |
Copied to clipboard
| Challenge: | Current multimodal large language models (MLLMs) show limited understanding of dental images. |
| Approach: | They propose a dental-specialized multimodal large language model trained via staged multimodal alignment and reinforcement learning. |
| Outcome: | The proposed model outperforms state-of-the-art models on disease classification and dental VQA tasks. |
Copied to clipboard
| Challenge: | Existing approaches to composable text operations often require plug-and-play . a single LM can perform arbitrary text operation composition in the latent space . |
| Approach: | They propose an efficient approach for composable text operations in the latent space of text . they connect pretrained LMs to the laten space and adapt them to the space . |
| Outcome: | The proposed approach improves on existing methods in the latent space of text. |
Copied to clipboard
| Challenge: | Existing studies focus on language-agnostic settings, neglecting the inherently multilingual nature of modern software development. |
| Approach: | They propose a proportion-dependent scaling law that prioritizes high-utility languages . they propose PLs to have varying effects during pre-training that affect model performance . |
| Outcome: | The proposed scaling law is based on 1000+ experiments across multiple languages and models. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have catalyzed the development of autonomous agents capable of executing complex, multi-turn tasks. |
| Approach: | They propose a framework for agentic reinforcement learning that integrates turn-level tree search with tree search to address key challenges. |
| Outcome: | The proposed framework addresses key challenges: limited exploration diversity, sparse credit assignment, and misaligned policy optimization. |
Copied to clipboard
| Challenge: | Existing methods for creating versatile MLLMs rely on joint training with paired instruction data, which is resource-intensive and challenging to extend to new modalities. |
| Approach: | They propose a new paradigm for multimodal large language models by reusing modality encoders and merging LLM parameters. |
| Outcome: | The proposed model retains the modal understanding capabilities of each original model. |
Copied to clipboard
| Challenge: | Existing methods for unlearning large language models struggle to balance effective forgetting with maintaining model utility. |
| Approach: | They propose a human-inspired unlearning framework that simulates forgetting on fuzzy data and represents them in hyperbolic and Euclidean spaces. |
| Outcome: | The proposed framework is able to forget sensitive content while maintaining the model’s language understanding, fluency, and benchmark performance. |
Copied to clipboard
| Challenge: | Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs. |
| Approach: | They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information. |
| Outcome: | The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests. |
Copied to clipboard
| Challenge: | Existing inference frameworks for natural language processing are not the best choice for online service of sequence processing problems. |
| Approach: | They propose a highly efficient inference library for Transformer models that includes GPU optimization techniques to streamline computation and reduce memory footprint. |
| Outcome: | The proposed library achieves 14x speedup compared with TensorFlow and 1.4x speed up compared to a concurrent CUDA implementation. |
Copied to clipboard
| Challenge: | Existing top-k attention methods struggle to strike a balance between efficiency and accuracy. |
| Approach: | They propose a top-k attention approach that integrates low-overhead techniques into the Top-k Attention process to achieve 7.2 speedup compared to vanilla full attention. |
| Outcome: | The proposed approach achieves 7.2 speedup compared to current top-k attention methods while maintaining model accuracy. |
Copied to clipboard
| Challenge: | Previous work shows that large language models generate hallucinations, yet the origins and mechanisms of these signals remain unclear. |
| Approach: | They propose to validate and disentangle two different pathways for truthfulness cues . they also propose to use the same mechanism to derive self-contained evidence from the generated answer . |
| Outcome: | The proposed applications improve hallucination detection performance by integrating two different inputs. |
Copied to clipboard
| Challenge: | Existing methods achieve promising performance in in-target stance detection when trained and tested on the same datasets. |
| Approach: | They propose a joint contrastive learning framework to generalize stance features for unseen targets. |
| Outcome: | The proposed framework achieves state-of-the-art on three benchmark datasets. |
Copied to clipboard
| Challenge: | zero-shot cross-lingual SLU is a challenging task in low-resource languages . a lack of labeled training data makes it difficult to align representations of similar sentences . |
| Approach: | They propose a framework that uses cyclical contrastive learning to achieve consistency between languages . they propose to use geodesic to measure the similarity to construct positive and negative pairs . |
| Outcome: | The proposed framework achieves state-of-the-art performance on multiATIS++ and MTOP datasets. |
Copied to clipboard
| Challenge: | Neural architecture search (NAS) uses weight-sharing supernets to generate diverse subnetworks without retraining. |
| Approach: | They propose a weight-sharing supernet that leverages mixture-of-experts to enhance supernet model expressiveness with minimal training overhead. |
| Outcome: | The proposed method achieves state-of-the-art (SoTA) performance in NAS for fast machine translation models, surpassing NAS-BERT and AutoDistil across various model sizes. |
Copied to clipboard
| Challenge: | Existing models for general intelligence fail to model how mental states interact and crystallize into group-level outcomes. |
| Approach: | They propose a multimodal benchmark for group-level Theory of Mind (ToM) to probe nonlinear collective behavior. |
| Outcome: | The proposed model performs significantly below human levels, exposing blind spots in modeling social structures and nonlinear collective behavior. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can generalize domain datasets unseen during training but are not able to predict domain adaptation performance. |
| Approach: | They propose to quantify dataset learning difficulty as the learning difficulty of generative summarization, which is determined by word-based compression rate and abstraction level. |
| Outcome: | The proposed model can predict performance on unknown domain datasets without training, and it is based on the findings. |
Copied to clipboard
| Challenge: | Existing methods to gauge model’s uncertainty through self-consistency in responses to the target query are misleading: an LLM may confidently provide an incorrect answer to a target query, yet give a confident and accurate answer to that same query when answering a knowledge-preserving perturbation of the query. |
| Approach: | They propose a method that uses multi-agent interaction to estimate black-box LLMs' uncertainty. |
| Outcome: | The proposed method outperforms existing self-consistency based methods and improves hallucination detection. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of tasks. |
| Approach: | They explore how LLMs can be extended to interact with and reason about the physical world through IoT sensors and actuators, a concept that they call "Penetrative AI". |
| Outcome: | The proposed approach extends LLMs' capabilities to interact with and reason about the physical world through IoT sensors and actuators. |
Copied to clipboard
| Challenge: | Existing studies on cognitive distortion have limited generalizability and performance of models in large-scale and cross-linguistic contexts. |
| Approach: | They propose a multi-task learning model based on teacher student architecture solution which improves generalization performance. |
| Outcome: | The proposed model improves generalizability and interpretability of the proposed model. |
Copied to clipboard
| Challenge: | Explicit /think> tags are used to expose intermediate reasoning and enable hybrid thinking behaviors. |
| Approach: | They propose a training-free prompting format that combines these triggers to achieve intermediate-budget reasoning, outperforming fixed-token and prompt-based baselines in terms of the accuracy–length trade-off. |
| Outcome: | The proposed method outperforms fixed-token and prompt-based prompts in accuracy–length trade-offs while improving Qwen3-8B on AIME from 69.8% to 72.4% and on GPQA from 58.5% to 61.1%. |
Copied to clipboard
| Challenge: | Experiments show that ChunkAttention can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation. |
| Approach: | They propose a prefix-aware self-attention module that can detect matching prompt prefixes across multiple requests and share their key/value tensors in memory at runtime. |
| Outcome: | The proposed module can speed up the self-attention kernel by 3.2-4.8 compared to the start-of-the-art implementation, with the length of the system prompt ranging from 1024 to 4096. |
Copied to clipboard
| Challenge: | escalating complexity of modern codebases has intensified the need for code retrieval systems capable of interpreting cross-component change intents. |
| Approach: | RepoAlignBench is a benchmark designed to evaluate repository-level code retrieval . the benchmark proposes an adversarial reflection-augmented dual-tower architecture . |
| Outcome: | The proposed framework achieves 12.2% Top-5 Accuracy and 7.1% Recall improvements over state-of-the-art benchmarks. |
Copied to clipboard
| Challenge: | Existing adapter-based transfer methods treat instruction-tuned models as passive targets . direct fine-tuning can disrupt this delicate balance and lead to instability or performance degradation. |
| Approach: | They propose a framework that incorporates instruction-level guidance into task adaptation. |
| Outcome: | The proposed framework outperforms direct fine-tuning and representative transfer-based baselines while maintaining robust generalization and favorable test-time scaling behavior. |
Copied to clipboard
| Challenge: | Existing models that retrain are time- and resource-consuming, but they lack the memory to support sequential and batch editing. |
| Approach: | They propose a model editing method that supports sequential and batch editing . they use a small amount of memory to store several hook layers that remain unchanged over time . |
| Outcome: | The proposed method is memory-friendly and can store hook layers that remain unchanged over time. |
Copied to clipboard
| Challenge: | Existing tools to generate structured content for research tasks are limited in their ability to generate high-quality roadmaps. |
| Approach: | They propose a benchmark to evaluate the ability of large language models (LLMs) to generate high-quality roadmaps for solving complex research problems. |
| Outcome: | The proposed system can improve LLMs’ ability for roadmap generation while saving 84% of the time required by human experts. |
Copied to clipboard
| Challenge: | anthropomorphic LLMs are being developed to serve diversified roles, but content safety concerns remain regarding their toxicity and toxicity. |
| Approach: | They propose to assign personality traits to large language models (LLMs) to reduce toxic language and social biases in their outputs by using the widely accepted HEXACO personality framework developed in social psychology. |
| Outcome: | The proposed model is able to perform on three toxic and bias benchmarks and shows that assigning personality traits reduces bias and toxicity similar to humans’ correlations between personality traits and toxic behaviors. |
Copied to clipboard
| Challenge: | Existing research focuses on benchmarking LLMs in single-turn dialogues, neglecting the nuanced nature of human feedback within real-world usage scenarios. |
| Approach: | They propose a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese. |
| Outcome: | The proposed benchmarks show that human feedback can significantly impact LLMs’ responsiveness in real-world usage scenarios. |
Copied to clipboard
| Challenge: | et al., 2022: ripple effect challenges knowledge editing for large language models. |
| Approach: | They propose a method to improve the accuracy of large language models by integrating Chain-of-Thought reasoning into the ICL editing approach. |
| Outcome: | RIPPLE-COT outperforms the state-of-the-art on the ripple effect, with gains ranging from 7.8% to 87.1%. |
Copied to clipboard
| Challenge: | Existing privacy protection methods for large language models suffer from performance degradation or large inference time overhead. |
| Approach: | They propose a plug-and-play method to protect the privacy of user inputs during LLM inference . they use offline restoration vectors to train restoration vector for each privacy span type . |
| Outcome: | The proposed method can prevent the linear growth of the privacy budget. |
Copied to clipboard
| Challenge: | Existing evaluation methods for transfer learning are limited in speech research . authors show that pre-trained models transfer well across multiple tasks . |
| Approach: | They propose a benchmark to evaluate pre-trained models by increasing task diversity and difficulty over SUPERB. |
| Outcome: | The proposed benchmark increases task diversity and difficulty over SUPERB-SG. |
Copied to clipboard
| Challenge: | Existing methods to predict scientific claims’ replicability use only hand-extracted statistics features without utilizing research papers’ text information. |
| Approach: | They propose two weakly supervised learning approaches that use automatically extracted text information of research papers to improve the prediction accuracy of research replication using both labeled and unlabeled datasets. |
| Outcome: | The proposed methods achieve an accuracy of 75.76% over real-world datasets. |
Copied to clipboard
| Challenge: | Existing automated review systems struggle with factual accuracy, rating consistency, and analytical depth. |
| Approach: | They propose a framework for generating comprehensive and factually grounded scientific paper reviews using supervised fine-tuning and reinforcement learning. |
| Outcome: | The proposed framework outperforms existing methods on ICLR 2025 papers. |
Copied to clipboard
| Challenge: | Existing benchmarks primarily evaluate planning and execution success, overlooking the self-reflective dimension of tool use. |
| Approach: | They propose a benchmark to assess LLMs’ self-reflective reasoning in tool-augmented multi-turn dialogues. |
| Outcome: | The proposed benchmark covers 10 domains with 88 distinct APIs and 968 annotated dialogues, systematically injecting diverse error types arising from both user and assistant behavior. |
Copied to clipboard
| Challenge: | Existing structural analysis methods for hieroglyphic scripts are script-specific and labor-intensive. |
| Approach: | They propose a hieroglyphic Stroke Analyzer framework that captures character-internal structures and semantics without handcrafted data. |
| Outcome: | The proposed framework captures character-internal structures and semantics without priors . it can be used to generalize hieroglyphic scripts across languages . |
Copied to clipboard
| Challenge: | Streaming automatic speech recognition models use high power consumption to improve usability and accuracy. |
| Approach: | They propose to optimize on-device speech recognition models by adjusting component energy sensitivities based on their specific energy sensitities to reduce power consumption. |
| Outcome: | The proposed approach achieves up to 47% lower energy usage while preserving comparable model accuracy and improving real-time performance compared to leading methods. |
Copied to clipboard
| Challenge: | Existing methods for stance detection for pure texts have limited results to multi-modal content. |
| Approach: | They propose a multi-modal stance detection framework that leverages target information to learn multi-modal stance features from textual and visual modalities. |
| Outcome: | The proposed framework achieves state-of-the-art in multi-modal stance detection on five datasets based on Twitter . |
Copied to clipboard
| Challenge: | Existing approaches to mathematical reasoning rely on static heuristics or pre-determined reasoning strategies. |
| Approach: | They propose an adaptive framework that integrates fuzzy theory into LLM-based mathematical reasoning. |
| Outcome: | The proposed framework outperforms state-of-the-art models while offering effective and interpretable diagnostics of intermediate problem-solving states. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have attracted considerable attention from academic and industrial communities due to their outstanding performance in various natural language processing tasks. |
| Approach: | They propose a Contrastive Learning Framework for Human Alignment to evaluate the noise within the data and dynamically adjust the training process. |
| Outcome: | The proposed framework surpasses other algorithms in terms of reward model scores, automatic evaluations, and human assessments on the widely used dataset "Helpful and Harmless" |
Copied to clipboard
| Challenge: | Named entity recognition (NER) suffers from the scarcity of annotated training data, especially for low-resource languages without labeled data. |
| Approach: | They propose a cross-lingual entity projection framework to enable zero-shot cross-linguistic NER with the help of a multilingual labeled sequence translation model. |
| Outcome: | The proposed method outperforms the baseline method on two benchmarks by a large margin of +3 7 F1 scores and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Pre-trained contextual representations like BERT have been widely used for NLP tasks. |
| Approach: | They propose to transform anisotropic sentence embedding distribution to smooth and isotropic Gaussian distribution by normalizing flows that are learned with an unsupervised objective. |
| Outcome: | The proposed method achieves significant performance gains over state-of-the-art embeddings on a variety of semantic textual similarity tasks. |
Copied to clipboard
| Challenge: | Current training recipes often rely on datasets dominated by short annotations with limited rationales, hindering the models' ability to generalize to tasks requiring comprehensive reasoning. |
| Approach: | They propose a two-stage post-training strategy that augments short answers with CoT reasoning generated by GPT-4o, enhancing the VLM's CoT capabilities through fine-tuning. |
| Outcome: | The proposed strategy enhances the model's CoT capabilities through fine-tuning and reinforcement learning. |
Copied to clipboard
| Challenge: | Existing prototypical networks for named entity recognition suffer from label dependency and tightly distributed prototypes, thus causing misclassifications. |
| Approach: | They propose an Entity-level Prototypical Network enhanced by dispersedly distributed prototypes to build entity-level prototypes and distribute them dispersionally. |
| Outcome: | The proposed system outperforms the previous models on two evaluation tasks and the Few-NERD settings in terms of overall performance. |
Copied to clipboard
| Challenge: | Text-to-Image Diffusion models generate high-quality images from textual descriptions, but they often produce images that do not fully align with the input prompts, resulting in semantic inconsistencies. |
| Approach: | They propose an automated repair approach to address catastrophic-neglect in T2I DMs. |
| Outcome: | The proposed model achieves 10.1%-16.3% higher Correct Rate in image generation compared to baselines. |
Copied to clipboard
| Challenge: | Existing benchmarks evaluate agents in simplified, idealized settings, relying on pre-packaged tool interfaces, overlooking critical steps, and assume inputs are clean and fully specified. |
| Approach: | They propose a framework that evaluates language agents in simplified, idealized settings . they show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 . |
| Outcome: | Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 . |
Copied to clipboard
| Challenge: | Existing user simulators lack authenticity and user-level diversity in interactions with large language models. |
| Approach: | They propose a user simulator with implicit user profiles that infers user profiles from human-machine interactions to simulate personalized and realistic dialogues. |
| Outcome: | The proposed framework outperforms baselines in authenticity and diversity while maintaining comparable consistency. |
Copied to clipboard
| Challenge: | a new evaluation framework is used to assess the extent and impact of position bias in information retrieval. |
| Approach: | They introduce a position-aware retrieval benchmark and a diagnostic metric to quantify position bias . they compare models with BM25, dense embedding models, ColBERT-style late-interaction models . |
| Outcome: | The proposed framework evaluates retrieval models for position bias from a worst-case perspective. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated good performance in many reasoning tasks, but struggle with some more complex reasoning tasks including logical reasoning. |
| Approach: | They propose five concrete tasks from three cognitive dimensions of WHAT, WHY, and HOW to evaluate LLMs’ capability of logical fallacy understanding. |
| Outcome: | The proposed dataset can be used to evaluate LLMs’ LFU capability and to fine-tune LLM models to obtain significantly enhanced performance on logical reasoning. |
Copied to clipboard
| Challenge: | Chart2Code is a new benchmark for evaluating the natural language to chart code generation capabilities of large multimodal models. |
| Approach: | They introduce Chart2Code, a new benchmark for evaluating the natural language to chart code generation capabilities of large multimodal models. |
| Outcome: | The proposed benchmark is the first to scale task complexity while capturing diverse scenarios. |
Copied to clipboard
| Challenge: | Existing methods do not examine social groups categorised by geographical information, leaving the region-related biases in pre-trained LMs unexplored. |
| Approach: | They propose a hierarchical regional bias evaluation method to quantify regional bias in pre-trained language models. |
| Outcome: | The proposed method evaluates regional bias with regard to comprehensive topics and measures potential regional bias that can be propagated to downstream tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating MLLMs have not addressed active perception . a novel benchmark is proposed to evaluate active perception in ML models . |
| Approach: | They propose a benchmark to evaluate active perception in Multimodal Large Language Models . they restrict the perceptual field of a model and require it to actively zoom or shift it . |
| Outcome: | The proposed benchmark focuses on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. |
Copied to clipboard
| Challenge: | Large-scale pre-trained vision-language models do not possess the ability to conduct in-context learning. |
| Approach: | They propose to meta-train a language model to perform in-context learning on NLP tasks and then transfer this model to VL tasks by attaching a visual encoder. |
| Outcome: | The proposed model outperforms the baseline model on VQA, OK-VQA, and GQA while having 20 times fewer parameters. |
Copied to clipboard
| Challenge: | Traditional industrial agents rely on modular workflows that fracture into a labyrinth of ad-hoc patches, leading to cascading errors and high latency. |
| Approach: | They propose a paradigm shift from external workflows to internalized knowledge representation that consolidates complex business logic and SOPs directly into the model’s parameters. |
| Outcome: | The proposed model breaks the impossible triangle of latency, accuracy, and complexity. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on simple image-text interactions, overlooking complex visual formats like charts. |
| Approach: | They propose a semi-automatic framework for generating evaluation samples through multi-modal keypoint extraction, knowledge graph construction, and qa pair synthesis. |
| Outcome: | The proposed framework generates 4,738 question-answering pairs across 8 domains from real-world documents. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained settings. |
| Approach: | They propose to compress large language models to reduce computation and memory consumption while maintaining accuracy. |
| Outcome: | The proposed algorithms preserve training data privacy but weaken the protection of personally identifiable information during conversations. |
Copied to clipboard
| Challenge: | Existing LLMs are limited by text-context budgets, resulting in token-expensive storage of raw trajectories . Optical Context Retrieval Memory (OCR-Memory) renders historical tra-jectorios into images annotated with unique visual identifiers. |
| Approach: | They propose a framework that leverages the visual modality as a high-density representation of agent experience. |
| Outcome: | Optical Context Retrieval Memory (OCRM) renders historical trajectories into images annotated with unique visual identifiers. |
Copied to clipboard
| Challenge: | Existing dialogue datasets have a bias between query distributions and real-world user language usage. |
| Approach: | They propose a framework for Chinese role-playing and a robust evaluation method . they propose specialized Chinese dialogue extraction model and specialized memory retrieval module . |
| Outcome: | The proposed framework extracts character dialogue from novels and ensures high data quality. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit strong performance on multilingual tasks, yet the process of constructing predictions in the target language remains under-explored. |
| Approach: | They propose a novel interpretability method focusing on the Feed-Forward Network (FFN) layers of Large Language Models. |
| Outcome: | The proposed interpretability method is based on the Feed-Forward Network (FFN) layer of Large Language Models. |
Copied to clipboard
| Challenge: | LLM-based agents for machine learning engineering rely on tree search to rank candidates. |
| Approach: | They propose an LLM-based agent that operationalizes gradient-based optimization. |
| Outcome: | The proposed agent achieves a state-of-the-art 35.1% any-medal rate on MLE-Bench with a limited budget on a single GPU. |
Copied to clipboard
| Challenge: | Recent pre-trained transformer-based definition generation models lack effective representation learning to contain full semantic components of the given word, leading to under-specific definitions. |
| Approach: | They propose a novel contrastive learning method that encourages the model to capture more detailed semantic representations from the definition sequence encoding. |
| Outcome: | The proposed method could generate more specific definitions compared with state-of-the-art models. |
Copied to clipboard
| Challenge: | Traditional evaluation metrics rely heavily on lexical similarity with human-written references, showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences. |
| Approach: | They propose an interpretable evaluation framework that evaluates alignment with specific human preferences by providing detailed comments and fine-grained scoring. |
| Outcome: | The proposed framework outperforms GPT-4 in Kendall correlation and accuracy with zero-shot reviewers. |
Copied to clipboard
| Challenge: | Code large language models (codeLLMs) focus on synthesizing the correct code snippet, ignoring the alignment with human preferences. |
| Approach: | They propose a benchmark code-based on 40 categories and 44 programming languages to emulate real-world coding tasks. |
| Outcome: | The proposed benchmarks show that open-source code LLMs perform better than open-sourced ones. |
Copied to clipboard
| Challenge: | Existing work on commenting based on textual content is focused on other modalities, such as graphics and images. |
| Approach: | They propose a task to integrate multiple modalities into automatic commenting . they construct a large-scale dataset and propose 'co-attention' model to capture dependency between textual and visual information. |
| Outcome: | The proposed model can achieve better performance than baselines. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit notable deficiencies in temporal reasoning . phrasing changes can lead LLMs to produce inconsistent outputs . |
| Approach: | They investigate the mechanistic interpretability of temporal ordering within event temporal reasoning . they identify a sparse subset of attention heads that are causally responsible for reasoning outcomes . |
| Outcome: | The proposed model outperforms other models in a variety of tasks and is validated by intervention-based experiments. |
Copied to clipboard
| Challenge: | Existing backdoor attacks against prompt-based learning involve injecting back doors into embedding layers or word embedders. |
| Approach: | They propose a backdoor attack against prompt-based learning that injects backdoors into embedding layers or word embeddable vectors. |
| Outcome: | The proposed backdoor attack outperforms two state-of-the-art models on six NLP tasks and three prompting strategies. |
Copied to clipboard
| Challenge: | Existing methods neglect domain-specific knowledge and use the same word embedding for each word in all domain-specified datasets. |
| Approach: | They propose a method to incorporate domain-specific and task-oriented information into meta-embeddings by combining pre-trained word embeddings. |
| Outcome: | The proposed method performs well on four text classification datasets and shows that it is compatible with existing methods. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) represent significant strides toward artificial general intelligence (AGI). |
| Approach: | They introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. |
| Outcome: | The framework reveals intrinsic patterns and mechanisms of LLMs in non-English languages, with Hungarian serving as an example. |
Copied to clipboard
| Challenge: | Context-DPO is the first alignment method specifically designed to enhance contextfaithfulness for large language models. |
| Approach: | They propose a benchmark that simulates Retrieval-Augmented Generation scenarios with knowledge conflicts to evaluate context-faithfulness. |
| Outcome: | The proposed method improves LLMs' context-faithfulness by 35% to 280% over open-source models. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) encounter performance limitations, impeding further enhancements in code generation tasks. |
| Approach: | They propose to combine two distinct prompts through a hybridization process to enhance the evolution of training prompts for code LLMs. |
| Outcome: | The proposed method significantly improves the performance of Code LLMs across five code generation benchmarks, namely HumanEval, HumanEva+, MBPP, mbap+ and MultiPL-E. |
Copied to clipboard
| Challenge: | Existing methods for budget-constrained tool learning have been overlooked . et al., 2023b) compared tool learning with other methods to improve performance . |
| Approach: | They propose a method for budget-constrained tool learning by creating a preferable plan under the budget constraint before utilizing the tools. |
| Outcome: | The proposed method reduces the cost of tool learning and reaches competitive Pass Rate. |
Copied to clipboard
| Challenge: | Existing methods for multimodal named entity recognition are limited due to limited resources. |
| Approach: | They propose a Few-shot Multimodal Named Entity Recognition task to address these relation types by constructing a multimodal graph and a new multimodal causal intervention strategy. |
| Outcome: | The proposed model improves on two multimodal named entity recognition datasets. |
Copied to clipboard
| Challenge: | Existing methods for image-to-text generation store all knowledge within parameters, thus requiring computational-expensive fine-tuning. |
| Approach: | They propose a Retrieval-augmented Visual Language Model that stores all the knowledge within parameters and can be used to retrieve it from the external database. |
| Outcome: | The proposed model significantly boosts performance for image-to-text generation tasks with 4x less parameters compared with baseline methods. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are stateless and limited by a finite context window, preventing them from maintaining knowledge across long conversations or evolving tasks. |
| Approach: | They propose a reinforcement learning framework that empowers LLMs to actively manage external memory through two specialized agents. |
| Outcome: | The proposed framework outperforms baselines and benchmarks across diverse question types, three benchmarks, and multiple model scales. |
Copied to clipboard
| Challenge: | Existing work for backdoor attacks on neural code models insert triggers into task-specific data for code-related downstream tasks, limiting the scope of attacks. |
| Approach: | They propose task-agnostic backdoor attacks for code pre-trained models . they use two learning strategies to implant backdoors into code understanding and generation models - Poisoned Seq2Seq learning and token representation learning . |
| Outcome: | The proposed model is pre-trained with two learning strategies to support the multi-target attack of downstream code understanding and generation tasks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for text style transfer are unsatisfactory. |
| Approach: | They propose to use a graph-based method to extract attribute content from sentences . they propose an efficient regularization to leverage attribute-dependent content as guiding signals. |
| Outcome: | The proposed method is based on a YELP and IMDB dataset and it is able to detect errors in the human evaluation. |
Copied to clipboard
| Challenge: | Existing retrieval methods struggle to achieve ideal results, a study finds . existing large language models lack prior knowledge of the content of superior legal articles . |
| Approach: | They propose to use a Chinese superior legal article retrieval dataset to find relevant articles with higher legal effectiveness. |
| Outcome: | The proposed dataset shows that existing retrieval methods struggle to achieve ideal results. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas, but their effectiveness in complex mathematical reasoning involving multi-step FOL deductions remains under-explored. |
| Approach: | They propose a self-adaptive solution that enhances the Diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct their proofs. |
| Outcome: | The proposed model improves diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct proofs. |
Copied to clipboard
| Challenge: | Existing retrieval-based approaches to solve multihop Knowledge Base Question Answering (KBQA) fail to utilize information from head-tail entities and the semantic connection between relations to enhance the information capturing of relations in KGs. |
| Approach: | They propose to use a dual relation graph to find the answer entity in a knowledge graph . they use primal entity graph reasoning, dual relation grafitment and interaction . |
| Outcome: | The proposed approach achieves significant performance gain over the prior state-of-the-art on two public datasets, WebQSP and CWQ. |
Copied to clipboard
| Challenge: | LM-LEXICON is a definition modeling approach that integrates data clustering, semantic expert learning, and model merging. |
| Approach: | They propose a definition modeling approach that integrates data clustering, semantic expert learning, and model merging using a sparse mixture-of-experts architecture. |
| Outcome: | The proposed model outperforms existing methods on five widely used benchmarks and achieves a BLEU score of 7%. |
Copied to clipboard
| Challenge: | Existing studies on neurons focus on emotion and rhetoric, neglecting their intrinsic connections. |
| Approach: | They propose a framework for fine-grained steering of emotion and rhetoric in large language models . they propose 'neuro-based' masking method that integrates multi-dimensional screening . |
| Outcome: | The proposed method achieves directed induction of non-target sentences and enhancement of emotion tasks via rhetoric neurons. |
Copied to clipboard
| Challenge: | Existing methods for detecting modelgenerated texts from human texts are limited by the fact that absolute likelihood values of texts are bound to certain linguistic and cognitive constraints. |
| Approach: | They propose to use relative likelihood values instead of absolute ones to extract useful features from the spectrum-view of likelihood for the human-model text detection task. |
| Outcome: | The proposed method can reveal subtle differences between human and model languages, which find theoretical roots in psycholinguistics studies. |
Copied to clipboard
| Challenge: | Biology-Instructions is the first large-scale instruction-tuning dataset for multi-omics biological sequences. |
| Approach: | They propose a large-scale instruction-tuning dataset for multi-omics biological sequences . they propose 'chatMultiOmics' to overcome limitations of current LLMs on multi-ome tasks . |
| Outcome: | The proposed dataset bridges LLMs and complex biological sequence-related tasks while maintaining conversational fluency. |
Copied to clipboard
| Challenge: | Existing models have been introduced to improve image comprehension, but there is no robust benchmark for imagetoweb conversion. |
| Approach: | They propose a benchmark to assess imagetoweb conversion proficiency of large multimodal models . they propose to measure layout information of web pages by parsing the Document Object Model tree . |
| Outcome: | The proposed benchmark measures the layout information of web pages—i.e., the positional relationships between elements—which has been overlooked by prior work. |
Copied to clipboard
| Challenge: | Existing distributed training frameworks are plagued by over-reliance on prior profiling and poor generalization across models/hardware. |
| Approach: | They propose a model-driven multi-agent framework that leverages Large Language Models to enable automatic and explainable distributed training strategy configuration. |
| Outcome: | The proposed framework outperforms expert-designed training strategies within 20 iterations. |
Copied to clipboard
| Challenge: | Pre-trained language models have been proposed and applied to many NLP tasks, yielding state-of-the-art performance, but high storage and computational costs obstruct them to be effectively deployed on resource-constrained devices and real-time applications. |
| Approach: | They propose a BERT distillation method which allows each intermediate student layer to learn from any intermediate teacher layers. |
| Outcome: | The proposed method can learn from different teacher layers adaptively for different NLP tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of Large Language Models (LLMs). |
| Approach: | They propose a repository-level benchmark named DevEval to evaluate LLMs' coding abilities in real-world code repositories. |
| Outcome: | The proposed benchmarks show that the LLMs perform better in real-world code repositories than existing benchmarks. |
Copied to clipboard
| Challenge: | Prior studies treat refusal as a generic "I don't know" lack of distinction limits downstream action decisions like requesting clarification or invoking external tools. |
| Approach: | They propose a benchmark to evaluate explicit uncertainty attribution in large language models. |
| Outcome: | The proposed method improves uncertainty attribution while preserving answer accuracy. |
Copied to clipboard
| Challenge: | commercial LLMs can be difficult to use in real-world clinical decision-making . a lightweight LLM can be used to collaborate with diverse clinical tools . |
| Approach: | They propose a lightweight LLM that can be used to build medical LLMs as agents . they use recursive curriculum learning to optimize the LLM in an easy-to-hard progression . |
| Outcome: | The proposed approach outperforms human experts in medical examinations on diverse datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) based neural code models struggle to generalize effectively to real-world inter-project out-of-distribution data. |
| Approach: | They propose a Cond-Idf measurement to measure the relatedness of a token with a label and its project-specificness. |
| Outcome: | The proposed framework improves both inter-project OOD generalization and adversarial robustness while not sacrificing accuracy on intra-project IID data. |
Copied to clipboard
| Challenge: | TexSmart supports fine-grained named entity recognition (NER) Large-scale fine-granular entity types are expected to provide richer semantic information for downstream NLP applications. |
| Approach: | They introduce TexSmart, a text understanding system that supports fine-grained named entity recognition (NER) and enhanced semantic analysis functionalities. |
| Outcome: | The proposed system supports fine-grained named entity recognition (NER) and enhanced semantic analysis functions. |
Copied to clipboard
| Challenge: | Existing models for text retrieval are based on a multi-stage process that involves retrieving documents from a large corpus. |
| Approach: | They propose to build a multilingual text representation model and a cross-encoder reranker from scratch for text retrieval. |
| Outcome: | The proposed models outperform the state-of-the-art models on long-context retrieval benchmarks. |
Copied to clipboard
| Challenge: | Large language models have mastered syntax-level code generation, but complex algorithmic reasoning remains a challenge. |
| Approach: | They propose a recurrent inductive bias that aligns with the recursive nature of programming logic. |
| Outcome: | The proposed model achieves comparable performance to standard dense models with more parameters. |
Copied to clipboard
| Challenge: | Existing frameworks for Large Language Models (LLMs) for Click-Through Rate prediction require a careful balance between computational efficiency and predictive accuracy. |
| Approach: | They propose a framework that integrates Retrieval-Augmented Generation with a novel multi-head early exit architecture to address both challenges. |
| Outcome: | The proposed framework reduces retrieval time while maintaining high model performance. |
Copied to clipboard
| Challenge: | Prior work has focused on situational awareness, which refers to a model's ability to recognize its operating phase and constraints, but it has neglected the complementary capacity to identify and adapt to the identity and characteristics of a dialogue partner. |
| Approach: | They formalize interlocutor awareness and evaluate its emergence in contemporary LLMs. |
| Outcome: | The proposed model reliably identify same-family peers and certain prominent model families, such as GPT and Claude. |
Copied to clipboard
| Challenge: | Vaccine interventions aim to answer concerns expressed about vaccination. |
| Approach: | They propose a dataset to evaluate how well responses are tailored to a common-ground opinion . they find that GPT-4-Turbo performs significantly better than others . |
| Outcome: | The proposed dataset outperforms fine tuned LLMs on the task of tailoring vaccine responses to common-ground opinions. |
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) is a key technique for enhancing LLMs’ reasoning abilities, yet its data inefficiency remains a major bottleneck. |
| Approach: | They propose a gradient-alignment-based method which intelligently selects the learnable and representative training reasoning data for RLVR post-training. |
| Outcome: | Experiments on five reasoning benchmarks show that the proposed method significantly reduces training data requirements while improving performance. |
Copied to clipboard
| Challenge: | Existing work regards dropped pronoun recovery and conversational discourse parsing as two separate tasks and tackles them separately. |
| Approach: | They propose a neural model for dropped pronoun recovery and conversational discourse parsing in Chinese conversational speech. |
| Outcome: | The proposed model outperforms the state-of-the-art models on a new dataset . the proposed model is based on linguistic and semantic information from Chinese conversational speech . |
Copied to clipboard
| Challenge: | Existing metric fails to capture text surprisal, but FACE-2 produces stronger agreement with human preferences. |
| Approach: | They propose a new automatic evaluation metric for open-ended text generation . they propose metric that extracts the dynamic patterns (spectrum) of text surprisal . |
| Outcome: | The proposed metric outperforms existing methods in revealing the model scaling effect . it produces stronger agreement with human preferences from a large human-annotated dataset . |
Copied to clipboard
| Challenge: | Existing research on inductive reasoning models emphasizes rule design without grounding them in specific scenarios. |
| Approach: | They propose to use LLMs to learn underlying patterns from limited examples in entirely new environments. |
| Outcome: | The proposed benchmark evaluates the inductive reasoning abilities of large language models in scientific settings. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) enhance programming by understanding and generating code across languages. |
| Approach: | a new benchmark evaluates code understanding and generation in repositories using code large language models. |
| Outcome: | The proposed model improves code understanding and generation in repositories by evaluating 1,888 test cases across 6 programming languages. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning Large Language Models (LLMs) encounter performance limitations, impeding further enhancements in code generation tasks. |
| Approach: | They propose to combine two distinct prompts through a hybridization process to enhance the evolution of training prompts for code LLMs. |
| Outcome: | The proposed method significantly improves the performance of Code LLMs across five code generation benchmarks. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated the potential of large language models (LLMs) for automatic error detection in math word problems (MWPs). |
| Approach: | They propose a framework that generates adaptive reference solutions using LLMs to enhance error detection by reducing conformity bias in MWPs. |
| Outcome: | The proposed framework mitigates the performance gap between conventional and alternative solutions in MWPs, especially when combined with reasoning-enhancing techniques like chain-of-thought prompting. |
Copied to clipboard
| Challenge: | Multimodal Emotion Recognition in Conversations models struggle due to lack of Common Sense Knowledge (CSK). |
| Approach: | They propose a multimodal approach to integrate multiple knowledge into the edge representations by integrating textual and visual CSK. |
| Outcome: | The proposed model outperforms state-of-the-art methods on two popular datasets. |
Copied to clipboard
| Challenge: | Existing approaches to reading comprehension systems are vulnerable to adversarial attacks. |
| Approach: | They propose to use knowledge distillation to transfer knowledge from an ensemble to a single model. |
| Outcome: | The proposed methods outperform the teacher on adversarial datasets and NarrativeQA benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for review generation lack topical and syntactic characteristics of natural languages. |
| Approach: | They propose a review generation model that uses aspect semantics, syntactic sketch, and context information to generate a sentence and corresponding words. |
| Outcome: | The proposed model can generate long and informative review text for users given a product and her/his rating on it. |
Copied to clipboard
| Challenge: | Existing studies on text simplification systems have focused on unsupervised methods due to the limited evaluation data in language and domain. |
| Approach: | They propose a Chinese text simplification dataset that provides a detailed analysis and an annotation process. |
| Outcome: | The proposed dataset evaluates the performance of unsupervised methods and advanced large language models. |
Copied to clipboard
| Challenge: | Existing knowledge editing methodologies often encounter parameter conflict during knowledge overwriting and excessive computational overhead. |
| Approach: | They propose a method that erases outdated knowledge and inserts new knowledge at the location that corresponds to the target knowledge. |
| Outcome: | The proposed method achieves more effective knowledge editing at a lower cost compared to previous methods across various base models. |
Copied to clipboard
| Challenge: | Existing discourse parsing approaches are constrained by predefined relation types, which can impede the adaptability of the parser for downstream tasks. |
| Approach: | They propose to introduce a task-aware paradigm to improve the versatility of the parser. |
| Outcome: | Empirical studies on dialogue discourse parsing datasets and a downstream task demonstrate the proposed framework. |
Copied to clipboard
| Challenge: | Existing approaches to enhance output diversity but compromise quality of outputs. |
| Approach: | They propose a training-free plug-and-play method that enhances output diversity while preserving generation quality. |
| Outcome: | The proposed method enhances output diversity while maintaining an optimal balance between diversity and quality. |
Copied to clipboard
| Challenge: | Existing legal language models struggle with dynamic courtroom interactions, resulting in overfitting to standardized legal tasks. |
| Approach: | They propose a new adversarial evolutionary approach for agents that performs dynamic knowledge learning and evolution through structured adversarials in a simulated courtroom program. |
| Outcome: | The proposed approach outperforms existing LLM-based models in three critical dimensions: cognitive agility, professional knowledge, and logical rigor. |
Copied to clipboard
| Challenge: | Online advertisement text generation models have achieved remarkable success in generating high-quality text ads, but some challenges remain, such as low-resource scenarios and training efficiency for multiple ad tasks. |
| Approach: | They propose a unified text ad generation framework with multi-task prompt learning to tackle low-resource ade generation problem and a multi-step prompt learning mechanism to efficiently solve multiple aed generation tasks. |
| Outcome: | The proposed framework outperforms the state-of-the-art on offline and online metrics. |
Copied to clipboard
| Challenge: | Bragging is a pervasive social-linguistic phenomenon that reflects complex human interaction patterns. |
| Approach: | They propose to use bragging recognition, bragging explanation, and bragging generation tasks to examine bragging in large language models (LLMs) . |
| Outcome: | The proposed models can identify bragging intent, social appropriateness, and account for context sensitivity and provide new insights into how LLMs process bragging. |
Copied to clipboard
| Challenge: | Existing neural-based ED models are confused by changeable contexts during testing . we propose a system that extracts statistical event features from word-event cooccurrence frequencies . |
| Approach: | They propose to integrate a set of statistical event features from word-event co-occurrence frequencies into the training set to cooperate with contextual features. |
| Outcome: | The proposed model outperforms ten strong baselines on ACE2005 and KBP2015 datasets. |
Copied to clipboard
| Challenge: | lack of reliable reward models for tool-use tasks has limited progress toward agentic AI . recent advances in agentic artificial intelligence are driven by tool-using capabilities of large language models. |
| Approach: | They propose a pipeline that constructs pairwise preference data using rule-based scoring and multidimensional sampling to build lightweight reward models. |
| Outcome: | The proposed model outperforms existing models on tool calling tasks with higher accuracy. |
Copied to clipboard
| Challenge: | Experimental evaluations demonstrate that our RNA FM consistently outperforms existing RNA . |
| Approach: | They propose to use filtered high-fidelity structure annotations to enhance the modeling ability of FMs in single nucleotide resolution tasks. |
| Outcome: | The proposed model outperforms existing RNA FMs on four genomic benchmarks and achieves top-tier results on DNA genomic benchmark. |
Copied to clipboard
| Challenge: | Existing studies focus on compressing the Key-Value cache or grouping attention heads, while overlooking redundancy between layers. |
| Approach: | They propose a lightweight substitute for self-attention in well-trained LLMs that uses feed-forward networks to align attention heads between adjacent layers and low-rank matrices to approximate differences in layer-wise attention weights. |
| Outcome: | The proposed model reduces redundancy by sharing weights across layers while maintaining high response quality while reducing redundant calculations within 53% 84% of the total layers. |
Copied to clipboard
| Challenge: | Existing methods for adversarial example generation are word-level or character-level, which ignore the ubiquitous phrase structure. |
| Approach: | They propose a phrase-level adversarial example generation framework to enhance the robustness of the translation model by adopting a sentence-level substitution strategy. |
| Outcome: | The proposed method improves translation performance and robustness to noise on three benchmarks. |
Copied to clipboard
| Challenge: | Existing research on hypothetical induction is limited by the observation annotations in the dataset and the ground truth hypotheses are mostly commonsense knowledge. |
| Approach: | They propose a first dataset for social science academic hypotheses discovery using raw web corpus as observations and propose valid, useful scientific hypothese . they propose 'a multi-module framework' that includes feedback mechanisms to boost performance. |
| Outcome: | The proposed dataset generates valid, novel, and helpful scientific hypotheses, even new to humanity, using open-domain data and a web corpus as observations. |
Copied to clipboard
| Challenge: | Large Language Models (LMMs) struggle with simple tasks such as geometry, e.g., arithmetic, and reasoning. |
| Approach: | They propose to leverage code as supervision for cross-modal alignment . they propose to use FigCodifier and ImgCode-8.6M to synthesize novel mathematical figures . |
| Outcome: | The proposed model surpasses GPT-4o and Claude 3.5 Sonnet in the geometry problem-solving subset of MathVista, achieving improvements of 8.9% and 9.2%. |
Copied to clipboard
| Challenge: | Large Language Models exhibit strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC) when instructions are revealed progressively in multi-turn settings, models get "Lost in Conversation" |
| Approach: | They propose a framework that encourages models to generate correct answers and judge solvability in multi-turn conversations. |
| Outcome: | The proposed framework improves models' ability to balance problem-solving with abstention . it reduces premature answering behaviors that cause lost-in-conversation (LiC) |
Copied to clipboard
| Challenge: | Existing named entity correction models fail to transcribe domain-speciffcnamed entities when theforms of the wrongly-transcribed words and the ground-truth entity are signiffcantly different. |
| Approach: | They propose a method that utilizes speech sound features to retrieve candidate entities . it uses speech sound feature to annotate entityerrors in ASR transcripts . |
| Outcome: | The proposed method can bring signiffcant improvement to entity accuracy. |
Copied to clipboard
| Challenge: | Automated synthesis of zeolite holds great significance for attaining economic and environmental benefits. |
| Approach: | They propose an event extraction task to mine structural synthesis actions from experimental narratives for modular automated synthesis. |
| Outcome: | The proposed method can significantly expedite automated synthesis of zeolites owing to its machine readability. |
Copied to clipboard
| Challenge: | Retrieval-augmented Generation (RAG) relies on effective retrieval capabilities, yet traditional sparse and dense retrievers struggle with multi-hop retrieval scenarios. |
| Approach: | They propose a graph expansion mechanism that augments any conventional base retriever and an agent framework that incorporates the resulting graph-based retrieval into a multi-step retrieval framework. |
| Outcome: | The proposed system achieves state-of-the-art results on three multi-hop question answering datasets while consuming fewer tokens and requiring fewer iterations than existing multi-step retrieval systems. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have revolutionized the field of natural language processing . performance of LLMs on subjective tasks is limited, authors say . |
| Approach: | They propose a method that allows LLMs to select between direct, role, and third-person perspectives for best way to solve corresponding subjective problem. |
| Outcome: | The proposed method outperforms widely used single fixed perspective based methods on 12 subjective tasks. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) show exceptional problem-solving capabilities but struggle with complex reasoning tasks. |
| Approach: | They propose a novel RAG approach that integrates retrieved information to guide tree-based reasoning process based on LLMs. |
| Outcome: | The proposed approach outperforms existing methods in large language models . iteratively plans intermediate sub-queries and answers based on the LLM itself . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) excel in various domains but face challenges when applied to data science workflows due to their complex, multi-stage nature. |
| Approach: | They propose a hierarchical graph-based agent that represents complexity and a progressive strategy for step-by-step verification, refinement, and consistent context management. |
| Outcome: | The proposed agent surpasses state-of-the-art baselines on the MATH dataset and performs better on InfiAgent-DABench. |
Copied to clipboard
| Challenge: | Existing methods for fine-grained text sentiment transfer only reverse the sentiment polarity of text, but they lack a robust and parallel learning algorithm. |
| Approach: | They propose a novel fine-grained text sentiment transfer task that revises a sequence to satisfy a given sentiment intensity while preserving the original semantic content. |
| Outcome: | The proposed model outperforms existing methods by a large margin in automatic evaluation and human evaluation. |
Copied to clipboard
| Challenge: | Code Large Language Models (CLLMs) are reshaping how software is built, maintained, and evolved. |
| Approach: | They propose to use BPE tokenization to inadvertently leak code secrets . they propose to mitigate the gibberish bias by using a newer tokenizer . |
| Outcome: | The proposed model is based on a novel method that can be used to detect and mitigate gibberish bias in CLLMs. |
Copied to clipboard
| Challenge: | Existing models that only use auxiliary languages to encourage multilingual agreement ignore the relationships between different language pairs. |
| Approach: | They propose a multilingual agreement-based method which explicitly models the agreement between different translation directions by randomly substituting some fragments of the source language with their counterpart translations of auxiliary languages. |
| Outcome: | The proposed method improves on the multilingual translation task of 10 language pairs. |
Copied to clipboard
| Challenge: | Structured pruning is a practical approach to deploying large language models (LLMs) but it fails to capitalize on modest task-specific calibration signals, causing limited downstream gains. |
| Approach: | They propose a method that removes attention heads and MLP channels using loss-based important scores . they use perplexity for language modeling and a margin-based objective for decision-style tasks . |
| Outcome: | The proposed method lowers perplexity and improves accuracy at higher sparsity . it also stabilizes accuracy and mitigates perxity collapse without fine-tuning . |
Copied to clipboard
| Challenge: | Existing methods to update or supplement large language models struggle under continuous knowledge drift. |
| Approach: | They propose a dynamic event benchmark and time-aware retrieval baseline that captures how knowledge evolves over time. |
| Outcome: | The proposed method enables systematic evaluation of model adaptation under continuous knowledge drift. |
Copied to clipboard
| Challenge: | Existing training-free methods are vulnerable to the attention sink phenomenon . Existing methods include contrastive decoding and auxiliary expert models . |
| Approach: | They propose a training-free attention intervention that constructs a PAD map to identify semantically core visual regions and applies per-head Median Absolute Deviation Scaling to adaptively control the intervention strength. |
| Outcome: | The proposed intervention improves visual grounding and reduces hallucinations on multiple LVLMs and benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to named entity recognition (NER) in Chinese are limited by the lack of annotated data. |
| Approach: | They propose a method which can automatically populate annotated training data without humancost by using distant supervision. |
| Outcome: | The proposed method performs better than comparison systems on two datasets. |
Copied to clipboard
| Challenge: | achieving synergistic improvements between generalization and domain specialization remains a challenge in pre-training and post-training. |
| Approach: | They propose a test-time cross-domain knowledge integration method that integrates general-purpose and domain-specific models to enhance their performance on complex, domainspecific tasks. |
| Outcome: | The proposed method combines the outputs of general-purpose and domain-specific models to improve their performance on complex, domainspecific tasks. |
Copied to clipboard
| Challenge: | Large language models respond well in high-resource languages but struggle in low-resourced languages. |
| Approach: | They propose a method to construct cross-lingual instruction following samples with instruction in English and response in low-resource languages. |
| Outcome: | The proposed method builds a large-scale cross-lingual instruction tuning dataset on 10 languages. |
Copied to clipboard
| Challenge: | Recent studies show that large reasoning models (LLMs) achieve strong performance on complex reasoning tasks. |
| Approach: | They propose a method that scores each candidate generation using the joint log-likelihood of the reasoning and final answer. |
| Outcome: | The proposed method outperforms baselines with 2x fewer samples in 20 out of 25 comparisons. |
Copied to clipboard
| Challenge: | Existing methods for topic-to-essay generation are insufficient for generating novel, diverse, and topic-consistent paragraph-level text with a set of topics. |
| Approach: | They propose to integrate commonsense from external knowledge base into the generator through dynamic memory mechanism and adversarial training to further improve topic-consistency. |
| Outcome: | The proposed task is more novel, diverse, and topic-consistent than existing methods in terms of both automatic and human evaluation. |
Copied to clipboard
| Challenge: | Existing work on memory-efficient parallelisms to reduce time and space complexity focuses on reducing time and complexity from system perspective. |
| Approach: | They propose a memory-efficient parallelism to reduce time and space complexity . they split input sequence into multiple chunks and feed each chunk into GPU . |
| Outcome: | The proposed approach is compatible with most existing parallelisms and makes 4D parallelismal possible. |
Copied to clipboard
| Challenge: | Recent advances in Vision-Language Models (VLMs) have broadened the scope of multimodal applications, but evaluations often neglect abstract dimensions such as personality traits and human values. |
| Approach: | They propose a Visual Question Answering (VQA) benchmark based on Schwartz’s value dimensions that capture core human values guiding people’s preferences and actions. |
| Outcome: | The proposed model can be used to evaluate visual question answering (VQA) tasks and to simulate diverse personas. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning are resource-efficient, but performance often falls short . a new approach, TeamLoRA, integrates collaborative and competitive modules to improve performance. |
| Approach: | They propose to introduce task-specific LoRA as domain experts to improve learning efficiency . teamLoRA integrates collaborative and competition modules to improve model learning . |
| Outcome: | Experiments show that TeamLoRA improves performance in multi-task learning . teamLorea integrates collaborative and competitive modules to improve performance . |
Copied to clipboard
| Challenge: | Existing systems for large-scale entity extraction are limited by the scale and variety of data available on internet platforms. |
| Approach: | They propose to build an entity extraction system for multiple document types at large scale using multi-modal Transformers. |
| Outcome: | The proposed system extracts multiple types of entities from multiple document types at large scale using multi-modal Transformers. |
Copied to clipboard
| Challenge: | Existing methods for prompt optimization make light of the importance of high-quality initialization and the identification of effective directions. |
| Approach: | They propose a method which uses a meta-instruction to generate high-quality initial prompts and iteratively optimize them at the sentence level. |
| Outcome: | The proposed method achieves consistent accuracy gain over baselines with less than five optimization steps. |
Copied to clipboard
| Challenge: | Variational Autoencoders are powerful language models and effective representation learning frameworks. |
| Approach: | They propose a fix for posterior collapse which improves held-out likelihood, reconstruction and latent representation learning . |
| Outcome: | The proposed fix significantly improves held-out likelihood, reconstruction, and latent representation learning compared with previous state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing methods for extracting relational facts from text have been successful . but with explosion of Web text, human knowledge is increasing drastically . |
| Approach: | They propose to improve relation extraction methods to extract relational facts from text . they analyze existing methods and show promising directions towards more powerful RE . |
| Outcome: | The proposed methods can extract relational facts from text, but they are still lacking in the current field. |
Copied to clipboard
| Challenge: | Existing detectors for AI-generated text lack robustness against adversarial perturbations, with even minor changes in characters or words causing a reversal in distinguishing between human-created and AI-generated text. |
| Approach: | They propose a siamese calibration technique to train the model to make equally confident predictions under different noise, which improves the model’s robustness against adversarial perturbations. |
| Outcome: | The proposed detector outperforms baseline methods on four datasets and is more generalizable in cross-domain, cross-genre, and mixed-source scenarios. |
Copied to clipboard
| Challenge: | LR-bench is a high-fidelity, up-to-date benchmark curated from 2024–2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey . |
| Approach: | They propose a reviewer-centric ranking framework that distills each reviewer’s recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals. |
| Outcome: | The proposed framework outperforms existing benchmarks and the CMU gold-standard dataset in the evaluation of AI/NLP manuscripts. |
Copied to clipboard
| Challenge: | Prompt engineering is an essential technique for enhancing the abilities of large language models (LLMs) by providing explicit and specific instructions. |
| Approach: | They propose a new approach that uses text embeddings to obtain basis vectors by matrix decomposition and constructs a space for representing all prompts. |
| Outcome: | The proposed approach significantly outperforms state-of-the-art prompt paradigms on ten public reasoning benchmarks. |
Copied to clipboard
| Challenge: | Linguistically informed analyses of language models (LMs) contribute to understanding and improvement of such models. |
| Approach: | They introduce a corpus of Chinese linguistic minimal pairs (CLiMP) to investigate what knowledge Chinese LMs acquire. |
| Outcome: | The proposed corpus of Chinese linguistic minimal pairs (CLiMP) covers 9 major Chinese linguist phenomena. |
Copied to clipboard
| Challenge: | Empirical results show up to 11.74% absolute (20.97% relative) increase over unimodal baselines. |
| Approach: | They propose to patch the visual modality to the textual-established attribute in- formation extractor. |
| Outcome: | Empirical results show up to 11.74% absolute (29.9% relative) increase over unimodal baselines. |
Copied to clipboard
| Challenge: | Existing LLMs often rely on complex prompting or extensive fine-tuning to introduce new capabilities while preserving strong generalizability. |
| Approach: | They propose a large-scale pre-training corpus to enhance LLM agents' capabilities . they use 103B agent-specific data encompassing 76,537 APIs . |
| Outcome: | The proposed training corpus outperforms open-source LLMs and commercial LLM agents on three agent benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for generating open-ended rubrics suffer from scalability bottlenecks and coarse criteria resulting in a supervision ceiling effect. |
| Approach: | They propose a framework for automated Coarse-to-Fine Rubric Generation . their framework uses principle-guided synthesis, multi-model aggregation, difficulty evolution . |
| Outcome: | The proposed framework produces comprehensive and highly discriminative criteria capable of capturing the subtle nuances. |
Copied to clipboard
| Challenge: | Existing methods for text ranking have improved performance, but there are still challenges. |
| Approach: | They propose a method that learns to re-rank the text retrieved for a given query by learning to predict the most relevant passage based on a latent preference matrix. |
| Outcome: | The proposed method outperforms all prior methods on datasets with extensive results. |
Copied to clipboard
| Challenge: | Recent advances in reasoning language models have witnessed a paradigm shift from short to long CoT pattern. |
| Approach: | They propose a behavior-constrained policy gradient with negative sample augmented (BCPG-NSA) negative steps are valuable components in long CoT models, authors argue . |
| Outcome: | The proposed framework outperforms baselines on math/coding reasoning benchmarks using the same training dataset. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting and large language models (LLMs) have shown great potential in improving performance on challenging reasoning tasks. |
| Approach: | They propose a new metric which extends the concept of pointwise V-information to black-box models and quantifies label-relevant new information introduced by CoT prompting. |
| Outcome: | The proposed metric extends the concept of pointwise V-information to black-box models, quantifying label-relevant new information introduced by CoT prompting beyond pre-existing label information. |
Copied to clipboard
| Challenge: | Experimental results show that multimodal emotion recognition is a state-of-the-art technique . textual, visual and acoustic modalities are involved in multimodal video emotion recognition . |
| Approach: | They propose a quantum-inspired adaptive-priority-learning model to address the challenges . they use quantum state to model modal features and Q-attention to integrate three modalities . |
| Outcome: | Experimental results show that QAP improves on previous models. |
Copied to clipboard
| Challenge: | Large language models can generate factually inaccurate content, a problem known as hallucination. |
| Approach: | They propose an approach that integrates a working memory that receives feedback from external resources. |
| Outcome: | The proposed method outperforms baselines on four fact-seeking datasets and increases the factuality metric by 2 to 6 points absolute. |
Copied to clipboard
| Challenge: | Procedural Multimodal Documents organize textual instructions and corresponding images step by step. |
| Approach: | They propose a novel temporal-modal entity Graph for comprehending PMDs . they propose encoding and reasoning modules to capture textual and visual entities . |
| Outcome: | The proposed model can capture textual and visual entities and trace their temporal-modal evolution. |
Copied to clipboard
| Challenge: | Referring Expression Comprehension (REC) is a cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. |
| Approach: | They propose to use a new reference expression comprehension (REC) dataset to evaluate the capabilities of language understanding, image comprehension, and language-to-image grounding. |
| Outcome: | The proposed model is able to reject scenarios where the target object is not visible in the image, a key aspect often overlooked in existing models and approaches. |
Copied to clipboard
| Challenge: | Existing methods for o1-level performance focus on unidirectional supervised fine-tuning (SFT), overlooking the intricate interplay between diverse reasoning patterns. |
| Approach: | They construct a reverse reasoning dataset and examine how it is supervised . they find that naively mixing forward and reverse data during SFT weakens the directional distinction . |
| Outcome: | The proposed model improves accuracy by 1.6%–6.8% over a standard model. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate unreliable responses due to their cognitive alignment of context and intent. |
| Approach: | They propose a benchmark to identify possible implicit assumptions in QA questions . they use retrieved Wikipedia fragments to identify interpretations for a given query . |
| Outcome: | The proposed benchmark identifies possible implicit assumptions and improves answer accuracy by 11.75% . retrieved Wikipedia fragments help identify possible interpretations for a given query . |
Copied to clipboard
| Challenge: | Existing static safety evaluation methods are ill-equipped to address dynamic nature of AI risks and evolving regulations, creating a critical safety gap. |
| Approach: | They propose a new paradigm of agentic safety evaluation reframing evaluation as a continuous and self-evolving process rather than a one-time audit. |
| Outcome: | The proposed framework shows a consistent decline in model safety as the evaluation hardens. |
Copied to clipboard
| Challenge: | Large language models are reshaping internet services, and serving them is costly. |
| Approach: | They propose an efficient distributed LLM serving system that splits prefill and decode requests into smaller chunks . |
| Outcome: | The proposed system reduces TTFT, TPOT, and latency compared to the state-of-the-art system. |
Copied to clipboard
| Challenge: | Existing methods for reweighting data mixtures rely on manual designation with certain heuristics based on intuition or empirical results. |
| Approach: | They propose a model-based framework that learns to re-weight domains by reinforcement learning on large quantities of data mixing trajectories with corresponding feedback from an evaluation environment. |
| Outcome: | The proposed framework outperforms baselines in achieving balanced performance across source and target fields and domain spaces without retraining. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have demonstrated impressive capabilities in code generation across various domains, but their ability to replicate complex, multi-panel visualizations remains largely unassessed. |
| Approach: | They propose a large-scale benchmark to evaluate chart generation from large- scale raw data and assess iterative code refinement in a multi-turn conversational setting. |
| Outcome: | The new benchmark evaluates 14 leading VLMs on real-world data and shows they struggle with complex plot structures and authentic data. |
Copied to clipboard
| Challenge: | End-to-end automatic speech recognition systems struggle to recognize rare name entities such as personal names, organizations and terminologies that are not frequently encountered in the training data. |
| Approach: | They propose a convolutional neural network-based ASR system that performs open-vocabulary keyword-spotting before the decoder to match the features between the entities and the utterances. |
| Outcome: | The proposed system significantly improves mixed-error-rate (MER) and entity recall compared to the original Whisper model on three internal datasets and two publicly available datasets. |
Copied to clipboard
| Challenge: | Existing methods for classification of reading difficulty of texts are insufficiently trained and lack of linguistic features. |
| Approach: | They propose a method that combines adaptive pre-training with feature fusion to capture different text difficulties and an interactive attention mechanism to integrate linguistic and deep features. |
| Outcome: | The proposed method achieves state-of-the-art (SOTA) performance on Chinese textbook dataset and can be applied to other languages. |
Copied to clipboard
| Challenge: | Existing news recommendation methods learn news representations solely based on news titles. Existing methods only utilize title information and neglect other valuable news information such as categories and entities. |
| Approach: | They propose a multi-task method to incorporate multi-field information into BERT, which improves its news encoding capability. |
| Outcome: | Extensive experiments on the MIND news recommendation benchmark show the proposed method is effective. |
Copied to clipboard
| Challenge: | Large language models (LLMs) face memory challenges due to the high cost of backpropagation. |
| Approach: | They propose a zeroth-order (ZO) optimization that matches memory usage to inference . they propose scalable and memory-efficient zeroth order (ZE) optimizer that integrates annealed A-GNB gradients with diagonal Hessian estimation and layer-wise clipping as a second-order pre-conditioner. |
| Outcome: | The proposed algorithm outperforms state-of-the-art methods with an average speedup of 20 over MeZO on RoBERTa-large and OPT-1.3B. |
Copied to clipboard
| Challenge: | Recent advances in large language models have achieved promising performances across various applications, but the challenge of integrating long-tail knowledge continues to impede the seamless adoption of LLMs in specialized domains. |
| Approach: | They propose a dynamic co-augmentation framework for the refinement of large language models and knowledge graphs in the context of Alzheimer's Disease. |
| Outcome: | The proposed framework can be used to study Alzheimer's Disease (AD) using LLMs and KGs. |
Copied to clipboard
| Challenge: | Existing studies show that explicitly modeling concept flows with a large commonsense knowledge graph improves response quality, but there is a gap between the knowledge graph and the conversation. |
| Approach: | They propose to model human conversational concept flows with a commonsense knowledge graph . they extract abundant concepts and relations from natural conversations and build a conversation-aware knowledge graph. |
| Outcome: | The proposed method performs better than baselines on a large-scale reddit conversation dataset. |
Copied to clipboard
| Challenge: | Existing data arbitration strategies for large language model training rely on surface-level heuristics that fail to diagnose intrinsic learning needs. |
| Approach: | They propose a framework that arbitrates data based on its degree of cognitive conflict with the model's existing knowledge. |
| Outcome: | Extensive experiments on WebShop and ALFWorld show that PRISM outperforms state-of-the-art hybrid methods while reducing computational costs by up to 3.22 . |
Copied to clipboard
| Challenge: | Recent approaches to reduce resource requirements for task-specific large language models have been developed. |
| Approach: | They propose a delta compression approach that optimizes for importance of a model . they use SVD to dynamically adjust the sparsity ratios of different vectors based on their importance . |
| Outcome: | The proposed approach achieves state-of-the-art in retaining task-specific knowledge even at high sparsity ratios. |
Copied to clipboard
| Challenge: | Stylistic style transfer is an important part of the image processing field . due to the low semantic similarity between the original image and the style image, many fine-grained style features are discarded. |
| Approach: | They propose a new style representation and transfer framework that can be adapted to existing image style transfers. |
| Outcome: | The proposed framework can be adapted to existing image style transfers. |
Copied to clipboard
| Challenge: | Existing benchmarks for MLM agents in interactive environments are limited by their focus on a single environment, lack of detailed and generalized evaluation methods, and the complexity of constructing tasks and evaluators. |
| Approach: | They propose a cross-environment agent benchmark framework that integrates graph-based evaluation and task generation methods. |
| Outcome: | The proposed framework supports multiple devices and can be easily extended to any environment with a Python interface. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are expensive and require extensive Continued Pre-Training and data-intensive alignment to expand. |
| Approach: | They propose a method which upcycles a dense model into a Mixture-of-Experts architecture, allocating different experts to different languages. |
| Outcome: | Experiments show that the proposed model upcycles a dense model into a Mixture-of-Experts(MoE) architecture, allocating different experts to different languages. |
Copied to clipboard
| Challenge: | Recent research has focused on developing conversational recommendation system (CRS), which provides valuable recommendations to users through conversations. |
| Approach: | They construct an authentic Chinese dialogue dataset consisting of over 25k dialogues and 770k utterances, which contains user profile, product knowledge base, and multiple sequential real conversations between users and recommenders. |
| Outcome: | The proposed dataset contains user profile, product knowledge base, and multiple sequential real conversations between users and recommenders. |
Copied to clipboard
| Challenge: | Existing studies focus on partial aspects of knowledge abstraction, concretization, and completion (KACC). |
| Approach: | They propose a unified knowledge graph benchmark to improve existing benchmarks . they collect new datasets that contain larger concept graphs and cross-view links . |
| Outcome: | The proposed benchmark improves existing benchmarks in terms of dataset scale, task coverage, and difficulty. |
Copied to clipboard
| Challenge: | Existing work on integrating audio encoders with large language models (LLMs) has focused on semantic understanding tasks, but different tasks may require distinct features that emphasize either semantic or acoustic aspects. |
| Approach: | They propose to use a prompt-aware mixture to enhance the Speech LLM that uses multiple audio encoders to extract different features based on the prompt. |
| Outcome: | The proposed approach outperforms all single-encoder Speech LLMs on ASR, speaker number verification, and AC tasks. |
Copied to clipboard
| Challenge: | Recent advances in Video Large Language Models have led to rapid development, significantly enhancing the capture of overall video semantics and achieving remarkable performance in general video understanding tasks. |
| Approach: | They propose a large-scale instance-motion-aware video instruction-tuning dataset iMOVE that utilizes Event-awful Spatiotemporal Efficient Modeling to retain informative instance spatiotemporal motion details while maintaining computational efficiency. |
| Outcome: | The proposed model excels in video temporal understanding and general video understanding. |
Copied to clipboard
| Challenge: | sarcasm is a form of irony conveying mockery and contempt . social media has become increasingly popular for identifying sarcasm . |
| Approach: | They develop a method to detect sarcasm from social media using augmented potentials. |
| Outcome: | The proposed method outperforms baselines on benchmark datasets. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited. |
| Approach: | They propose an Adaptive Curriculum Reinforcement Learning framework to advance long-form writing capabilities beyond SFT. |
| Outcome: | Experiments on 7B-scale writer models show that Writing-RL improves long-form writing performance over strong SFT baselines. |
Copied to clipboard
| Challenge: | Proximal Policy Optimization (PPO) is central to aligning Large Language Models with verifiable rewards. |
| Approach: | They propose a scalable algorithm that harmonizes sample efficiency with stability of outcome-based updates. |
| Outcome: | The proposed algorithm outperforms standard PPO and matches the performance of computation-heavy group-based methods. |
Copied to clipboard
| Challenge: | Current work relies on pre-defined rules or templates to control the style of speech. |
| Approach: | They propose to use open-domain instructions to generate speech with the acoustic style that meets users’ needs based on their instructions. |
| Outcome: | The proposed model can be used to generate speech with the acoustic style that meets users’ needs based on open-domain instructions. |
Copied to clipboard
| Challenge: | Existing methods for detecting public sentiment drift are not designed for sentiment drift detection. |
| Approach: | They propose a Hierarchical Variational Auto-Encoder model to learn better distribution representation and a new drift measure to directly evaluate distribution changes between historical and new data. |
| Outcome: | The proposed model performs better than three existing state-of-the-art methods. |
Copied to clipboard
| Challenge: | LVLMs have shown impressive progress by integrating visual perception with linguistic understanding to produce contextually grounded outputs. |
| Approach: | They propose a visual evidence prompting method to mitigate hallucinations in large vision-language models by using small visual models to complement them. |
| Outcome: | The proposed method reduces hallucinations by reducing false activation and enhancing correct ones. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have advanced multimodal learning, driving progress in cross-modal reasoning. |
| Approach: | They propose to examine moral robustness of vision-language models by analyzing their moral stances under multimodal perturbations. |
| Outcome: | The proposed model-agnostic multimodal perturbations expose VLMs to a variety of moral vulnerabilities, including a sycophancy trade-off where stronger instruction-following models are more susceptible to persuasion. |
Copied to clipboard
| Challenge: | Existing text-based compliance checking methods are limited by their flexibility and lack structure. |
| Approach: | They propose a text-based compliance checking framework based on Retrieval-Augmented Generation that integrates a static layer for storing factual knowledge, a dynamic layer for retrieval and reasoning, and an eventic graph to structurally describe regulatory information. |
| Outcome: | The proposed framework consistently achieves state-of-the-art results across various scenarios surpassing baselines. |
Copied to clipboard
| Challenge: | a number of studies have shown that transformer-based language models detect when a word is anomalous in context, but likelihood scores do not tell the cause of the anomaly. |
| Approach: | They propose to use Gaussian models for density estimation at intermediate layers of three language models to evaluate grammaticality. |
| Outcome: | The proposed method on BLiMP shows that language models employ different mechanisms to detect different types of linguistic anomalies. |
Copied to clipboard
| Challenge: | Existing methods for dialogue state tracking are still challenging, but they are improving . a new approach to dialogue state monitoring is proposed, called Seq2Seq-DU . |
| Approach: | They propose a new dialogue state tracking module that formalizes DST as a sequence-to-sequence problem. |
| Outcome: | The proposed method outperforms existing methods on benchmark datasets in different settings. |
Copied to clipboard
| Challenge: | Existing benchmarks lack domain-specific data, realistic workflow-level task design, and standardized workflow- level evaluation. |
| Approach: | a new benchmark evaluates large language models on financial management workflows . the global financial services market is projected to grow to $37 trillion by 2027 . |
| Outcome: | a new benchmark for large language models on financial management workflows reveals critical capability gaps . accuracy drops from 90% on basic tasks to 40% on complex scenarios requiring multi-step reasoning . the global financial services market reached $25.8 trillion in 2022 and is projected to grow to $37 trillion by 2027 . |
Copied to clipboard
| Challenge: | Existing multi-agent learning approaches foster collaboration among Large Language Models (LLMs) yet they still rely on re-executing the MAS during inference. |
| Approach: | They propose a co-learning framework that integrates Dynamic Interaction and Perception Calibration to enhance LLMs' independent problem-solving ability. |
| Outcome: | The proposed framework integrates Dynamic Interaction and Perception Calibration to improve LLMs' independent problem-solving ability. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to adequately evaluate the proficiency of Large Language Models (LLMs) Existing standards do not cover the skills needed to evaluate LLMs in scientific literature analysis. |
| Approach: | They propose a benchmark to evaluate the proficiency of large language models in scientific literature analysis. |
| Outcome: | SciAssess evaluates 11 LLMs on multiple tasks across scientific fields. |
Copied to clipboard
| Challenge: | Existing methods for selecting training data from general datasets fail to account for the joint distribution of instructions, resulting in inefficient learning and suboptimal knowledge transfer. |
| Approach: | They propose a method that constructs a mixed gradient-based instruction graph to capture the joint distribution and interdependencies among instructions. |
| Outcome: | The proposed method outperforms existing methods on domain adaptation tasks and in complex, data-scarce scenarios. |
Copied to clipboard
| Challenge: | a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling . |
| Approach: | They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution. |
| Outcome: | The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data. |
Copied to clipboard
| Challenge: | Experimental results show that our model can generate natural, human-like and personalized comments. |
| Approach: | They propose a model that takes user profile into account when generating comments on social media and integrates it with a gated memory. |
| Outcome: | The proposed model can generate natural, human-like and personalized comments on social media. |
Copied to clipboard
| Challenge: | Existing solutions to reasoning tasks require extensive human annotations or fail in scenarios with inconsistent responses. |
| Approach: | They propose a new method that enables LLMs to self-rank their responses without additional resources. |
| Outcome: | The proposed method improves reasoning performance of ChatGPT and GPT-4 with 13% improvement over existing methods. |
Copied to clipboard
| Challenge: | Large vision-language models (LVLMs) are evolving rapidly and require data with human supervision to achieve better alignment. |
| Approach: | They introduce VLFeedback, the first large-scale vision-language feedback dataset . they train Silkie, an LVLM fine-tuned via direct preference optimization . |
| Outcome: | The proposed model outperforms its base model in helpfulness, visual faithfulness, and safety metrics and exhibits enhanced resilience against red-teaming attacks. |
Copied to clipboard
| Challenge: | Existing metrics for video captioning are based on text-based comparisons with ground-truth references. |
| Approach: | They propose a reference-free benchmark that assesses video captions based on their utility . they will release the benchmark to facilitate reproducible research . |
| Outcome: | The proposed benchmark improves on human-verified, fine-grained questions . it correlates significantly better with human judgments than existing metrics . |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have stepped forward the development of multilingual speech and machine translation by its reduced representation errors and incorporated external knowledge. |
| Approach: | They propose a generative paradigm for translation tasks that integrates the diverse translation versions in N-best list. |
| Outcome: | The proposed model outperforms the state-of-the-art model on speech and machine translation benchmarks on various languages. |
Copied to clipboard
| Challenge: | Improving training efficiency remains a challenge in large-scale Reinforcement Learning (RL). |
| Approach: | They propose a curriculum RL framework with stage-wise context scaling to improve RL training efficiency. |
| Outcome: | The proposed framework outperforms state-of-the-art reasoning models on five benchmarks and achieves 49.6% accuracy on AIME 2024. |
Copied to clipboard
| Challenge: | Existing long document question answering systems process texts as flat sequences or use heuristic chunking, which overlooks the discourse structures that guide human comprehension. |
| Approach: | They propose a discourse-aware hierarchical framework that leverages rhetorical structure theory for long document question answering. |
| Outcome: | The proposed framework exhibits strong robustness across diverse document types and linguistic settings. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) lack understanding of multi-image and interleaved inputs due to the visual features encoded by frozen encoders before being fed into the LLM backbone. |
| Approach: | They propose a two phase paradigm to enable in-depth multimodal context fusion prior to feeding the features into LLMs. |
| Outcome: | The proposed paradigm boosts the performance on 7 multi-image scenarios, contributing to increments on average accuracy by 2.13% and 7.60% against strong MLLMs baselines with 3B and 11B LLMs, respectively. |
Copied to clipboard
| Challenge: | Existing methods for depression assessment rely on standardized ratings, but they are time-consuming and subject to inter-rater variability. |
| Approach: | They propose a fine-grained model for subscore prediction via multi-task learning that can be used to predict depression severity using multiple tasks. |
| Outcome: | The proposed model outperforms baselines and Qwen3-14B direct scoring on the public E-DAIC dataset and to a large-scale private clinical dataset. |
Copied to clipboard
| Challenge: | Current Question Answering over Knowledge Graphs (KGQA) tasks focus on binary facts, but neglect n-ary facts. |
| Approach: | They propose a new fact-tree reasoning framework that transforms the question into a fact tree and performs iterative fact reasoning on the fact tree to infer the correct answer. |
| Outcome: | The proposed framework performs iterative fact reasoning on the fact tree to infer the correct answer. |
Copied to clipboard
| Challenge: | Existing approaches to task-oriented dialogue systems require a large number of handcrafted features and labels. |
| Approach: | They propose a "Two-Teacher One-Student" learning framework for task-oriented dialogue . the framework amalgamates knowledge from two teacher networks and provides guidance . |
| Outcome: | The proposed framework outperforms baseline methods on two benchmark datasets . it can retrieve accurate KB entities and generate human-like responses simultaneously . |
Copied to clipboard
| Challenge: | Existing neural machine translation models adopt a monotonic decoding order of either left-to-right or right-to left. |
| Approach: | They propose a method that starts decoding target words from the right side of a median word and generates words on the left. |
| Outcome: | The proposed method outperforms baseline models on three datasets. |
Copied to clipboard
| Challenge: | Chain-of-thought (CoT) has been shown to improve the reasoning capability of large language models (LLMs). |
| Approach: | They propose a framework which iteratively enhances the model’s Vision-language Reasoning by Reflecting on CoT Rationales. |
| Outcome: | The proposed framework improves multimodal reasoning on vision-language tasks by 23% to 60% over baselines. |
Copied to clipboard
| Challenge: | Existing methods to shorten CoTs use length penalties or global entropy reduction . Existing approaches to CoT reasoning have significant practical drawbacks . |
| Approach: | They propose a method that shortens CoTs by length penalties or global entropy reduction . they integrate ETR into Group Relative Policy Optimization and evaluate it . |
| Outcome: | The proposed objective improves accuracy–efficiency trade-off by +9.9% while reducing CoT length by 67% across four benchmarks. |
Copied to clipboard
| Challenge: | Existing RAG methods suffer from a two-part problem: semantic drift and concatenation fallacy . et al.: rapid development of Large Language Models has led to a paradigm shift in artificial intelligence . |
| Approach: | They propose a multi-agent retrieval augmentation framework guided by medical domain knowledge to address these challenges. |
| Outcome: | The proposed framework outperforms existing general RAG baselines on five widely used medical benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promise in formal theorem proving, but their token-level processing often fails to capture the inherent hierarchical nature of mathematical proofs. |
| Approach: | They propose a regularization method that aligns LLMs’ attention mechanisms with mathematical reasoning structures and establishes a five-level hierarchy from foundational elements to high-level concepts. |
| Outcome: | The proposed method improves proof success rates by 2.05% on miniF2F and 1.69% on ProofNet while reducing proof complexity by 23.81% and 16.50% respectively. |
Copied to clipboard
| Challenge: | Existing dialogue systems focus on brief single-session interactions, neglecting real-world needs for long-term companionship and personalized interactions. |
| Approach: | They propose a model-agnostic framework for long-term dialogue agents . they use event summary and persona management to enable reasoning . |
| Outcome: | The proposed framework incorporates three independently tunable modules dedicated to event perception, persona extraction, and response generation. |
Copied to clipboard
| Challenge: | Existing embedding-based retrieval systems rely on heuristic and suboptimal cutoffs for item retrieval. |
| Approach: | They propose a probabilistic Embedding-Based Retrieval framework that learns a shared semantic representation space for both queries and items. |
| Outcome: | The proposed framework improves retrieval precision and recall, and ablation studies show it captures the differences between head-to-tail queries. |
Copied to clipboard
| Challenge: | Existing approaches to generate paraphrases with weak supervision are limited in real-world scenarios due to the lack of coherent and controllable generated paraphrase. |
| Approach: | They propose a method to generate high-quality paraphrases with weak supervision . they obtain abundant weakly-labeled parallel sentences via retrieval-based pseudo paraphrase expansion . |
| Outcome: | The proposed approach achieves significant improvements over existing methods and is even comparable in performance with supervised state-of-the-arts. |
Copied to clipboard
| Challenge: | Existing fewshot methods for slot tagging are weak in encoding slot name semantics and slot dependencies. |
| Approach: | They propose a simple and effective few-shot model for slot tagging which incorporates machine reading comprehension (MRC) using source domain and target domain data. |
| Outcome: | The proposed model outperforms state-of-the-art methods on the SNIPS dataset. |
Copied to clipboard
| Challenge: | Existing studies on aspects-based sentiment analysis focus on a single opinionated sentence. |
| Approach: | They propose a model to combine aspects and their sentiments for QA forums . they use cross-sentence aspect-opinion interaction modeling to align the aspect mentioned in the question and associated opinion clues in the answer. |
| Outcome: | The proposed model outperforms baseline models on three real-world datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have achieved impressive results in Machine Translation (MT). human evaluations reveal that LLM-generated translations still contain various errors. |
| Approach: | They propose a LLM-based self-refinement framework that feeds error information back into LLMs to facilitate self-finement, leading to enhanced translation quality. |
| Outcome: | The proposed framework outperforms internal refinement and feedback methods while ensuring a robust translation quality baseline. |
Copied to clipboard
| Challenge: | Existing methods for dynamic web navigation rely on greedy strategies or value estimation, struggle to achieve effective backtracking and are heavily dependent on proprietary models. |
| Approach: | They propose a cognitive multi-agent collaboration framework that enhances cyberspace exploration capability through In-Context Exploration. |
| Outcome: | The proposed framework surpasses the proprietary model Claude-3.5 Sonnet on the WebArena benchmark. |
Copied to clipboard
| Challenge: | Existing QA datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements. |
| Approach: | They propose to use FairytaleQA to generate 10,580 questions based on 278 children-friendly stories to assess model's fine-grained learning skills. |
| Outcome: | The proposed dataset consists of 10,580 questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations. |
Copied to clipboard
| Challenge: | Existing work focuses on strengthening the knowledge-time association between text and time-stamps, but this is insufficient for downstream tasks. |
| Approach: | They propose a model that explicitly connects all temporally-scoped facts by modeling the time relations between any two sentences. |
| Outcome: | The proposed model outperforms baseline T5 on multiple temporal question answering datasets . it is especially good at modeling long-range complex temporal dependencies, the authors say . |
Copied to clipboard
| Challenge: | Sensor metadata tagging is a key component of smart building applications. |
| Approach: | They propose a framework that learns a sensor metadata tagger for a new building based on its raw metadata and some existing fully annotated building. |
| Outcome: | The proposed framework learns a sensor metadata tagger for a new building based on its raw metadata and some existing fully annotated building. |
Copied to clipboard
| Challenge: | Existing watermarking algorithms rely on heuristic green/red token lists . however, these lists are inconsistent and can be compromised . |
| Approach: | They propose a framework for robust and secure LLM watermarking using reinforcement learning. |
| Outcome: | The proposed method achieves state-of-the-art trade-off across all criteria with notable improvements in resistance to spoofing attacks without degrading other criteria. |
Copied to clipboard
| Challenge: | Existing approaches to generating factually inconsistent outputs are resource-intensive. |
| Approach: | They propose a plug-and-play intervention designed to enhance factuality by inserting premature layers formed through mathematical interpolation with adjacent layers. |
| Outcome: | The proposed intervention reduces hallucinations while outperforming baselines on four datasets. |
Copied to clipboard
| Challenge: | Existing methods for detecting new intents with labeled data are not cluster-friendly . a robust prototypical attracting learning (RPAL) method is designed to compel instances to gravitate toward their corresponding prototype . |
| Approach: | They propose a robust and adaptive prototypical learning framework for globally distinct decision boundaries for both known and new intent categories. |
| Outcome: | The proposed method improves on CLINC, BANKING, and StackOverflow benchmarks on three challenging benchmarks. |
Copied to clipboard
| Challenge: | Existing pre-training methods underutilize the benefits of language understanding for generation. |
| Approach: | They propose a GAN-style model for encoder-decoder pre-training with an auxiliary discriminator. |
| Outcome: | The proposed model outperforms existing pre-trained models and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Experimental results show ToM outperforms existing divide-and-conquer frameworks . RAG relies on similarity-based rankings to retrieve and reason over chunks based on logical coherence . |
| Approach: | They propose a Tree-oriented MapReduce framework for long-context reasoning . it leverages the hierarchical structure of long documents by constructing a DocTree . |
| Outcome: | Experimental results show that ToM outperforms existing divide-and-conquer frameworks and RAGs . the proposed framework improves logical coherence and long-context reasoning on 70B+ LLMs compared to existing approaches . |
Copied to clipboard
| Challenge: | Existing methods to evaluate factual consistency in text summarization neglect the intrinsic cause of factual inconsistency or rely on auxiliary tasks. |
| Approach: | They propose a method to evaluate the factual consistency in text summarization via counterfactual estimation, which formulates the causal relationship between source document, generated summary, and the language prior. |
| Outcome: | The proposed metric improves correlation with human judgments and convenience of usage on three public abstractive text summarization datasets. |
Copied to clipboard
| Challenge: | Existing methods to attack transformer models are not effective at character level, but they are a natural attack scenario. |
| Approach: | They propose a character-level adversarial attack method against transformer models . they use a gradient-based method to find the most vulnerable words in a sentence . |
| Outcome: | The proposed method outperforms previous methods on sentence-level and token-level tasks. |
Copied to clipboard
| Challenge: | Named entity recognition is a natural language processing task . nested NER is based on a linear structure, but there is no research on applying corpus-level information to NER. |
| Approach: | They propose a holistic structure parsing algorithm to reveal the entire NEs in a sentence . they introduce points-wise mutual information and other frequency features from corpus-aware statistics . |
| Outcome: | The proposed model outperforms existing models on widely-used benchmarks and achieves state-of-the-art. |
Copied to clipboard
| Challenge: | Existing approaches to optimize conversational agents often rely on explicit preference pairs and expert evaluations. |
| Approach: | They propose a conversational agent framework that leverages the structured dependency between agent responses and user reactions to extract implicit feedback. |
| Outcome: | The proposed framework improves on MT-Bench-101, WildBench, and FB-Bech, and shows that mining implicit feedback supports better multi-turn alignment under evolving user preferences. |
Copied to clipboard
| Challenge: | aaron carroll: the precise localization of non-verbal vocal events remains a critical yet under-explored challenge. carroll says current methods suffer from insufficient task definitions with limited category coverage. carrol: knowing exactly where an event occurred is not enough; knowing exactly what it happened is. |
| Approach: | They propose a taxonomy of 21 vocal events with a new categorization into discrete versus continuous types. |
| Outcome: | The proposed model disentangles ASR errors from event detection while maintaining ASR quality. |
Copied to clipboard
| Challenge: | Existing evaluation metrics cannot fairly evaluate the outputs of RAG models during training and evaluation. |
| Approach: | They propose a method which prompts LLMs to generate different judgments based on various combinations of judgment dimensions and utilizes the judge-consistency to evaluate these judgments. |
| Outcome: | The proposed method generates more accurate evaluations for RAG models across different RAG model and datasets. |
Copied to clipboard
| Challenge: | Existing dialogue systems do not exploit document knowledge effectively enough. |
| Approach: | They propose a Transformer-based architecture for document grounded conversations that incorporates document knowledge into a two-pass decoder to improve context coherence and knowledge correctness. |
| Outcome: | The proposed model outperforms baselines on context coherence and knowledge relevance on a real-world document grounded dataset. |
Copied to clipboard
| Challenge: | In-battle commentary is an important component of live streaming of e-sports competitions and is applicable to a wide range of scenarios like combat information analysis and live streaming. |
| Approach: | They propose a generative system for in-battle real-time commentary in mobile MOBA games and propose 'transform' method to convert match statistics and utterances into consistent encoding space. |
| Outcome: | The proposed system is based on real-time match statistics and events and can be used for live streaming, e-sports commentary and combat information analysis. |
Copied to clipboard
| Challenge: | Existing approaches to zero-shot link prediction use textual features of relations as auxiliary information to improve the encoded representation. |
| Approach: | They propose a Hierarchical N-gram framework for Zero-Shot Link Prediction that leverages character n-gram information for ZSLP. |
| Outcome: | The proposed method achieves state-of-the-art on two standard ZSLP datasets. |
Copied to clipboard
| Challenge: | Scripts are written text for plays, movies, or broadcasts. |
| Approach: | They propose a multi-level contrastive learning framework to capture characters’ global information in a fine-grained manner. |
| Outcome: | The proposed framework improves on three character understanding sub-tasks by a considerable margin. |
Copied to clipboard
| Challenge: | Recent Chinese word segmentation models tend to learn the segmentation knowledge through in-vocabulary words rather than understanding the meaning of the entire context. |
| Approach: | They propose a context-aware approach that incorporates unsupervised sentence representation learning over different dropout masks into the multi-criteria training framework. |
| Outcome: | The proposed approach achieves state-of-the-art (SoTA) performance on six of the nine CWS benchmark datasets and out-of vocabulary (OOV) recalls for eight of nine. |
Copied to clipboard
| Challenge: | Existing studies on robustness to explicit noise (e.g., document semantics) but overlook implicit noise (spurious features). |
| Approach: | They propose a framework to quantify the robustness of RAGs against spurious features by integrating a data synthesis pipeline and a taxonomy. |
| Outcome: | The proposed framework quantifies the robustness of RALMs against spurious features. |
Copied to clipboard
| Challenge: | Extensive experiments show that STAR outperforms previous pre-training methods and ranks first on the leaderboard . text-to-SQL parsing aims to translate natural language (NL) questions into executable SQL queries . |
| Approach: | They propose a SQL guided pre-training framework STAR for context-dependent text-to-SQL parsing . they propose two objectives that explore context-dependence of NL utterances and SQL queries . |
| Outcome: | The proposed framework outperforms existing methods on two downstream benchmarks and ranks first on the leaderboard. |
Copied to clipboard
| Challenge: | Historical analogies are important abilities that help people make decisions and understand the world. |
| Approach: | They propose a historical analogy acquisition task that uses large language models to acquire historical analogies. |
| Outcome: | The proposed method mitigates hallucinations and stereotypes when LLMs generate historical analogies. |
Copied to clipboard
| Challenge: | Compared to existing benchmarks, FinanceReasoning provides three key advancements: (1) credibility; (2) comprehensiveness; (3) numerical precision; (4) complexity; (5) complexity; and (6) complexity. |
| Approach: | They propose a benchmark to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. |
| Outcome: | The proposed benchmark exceeds existing benchmarks in 67.8% of financial concepts and formulas and is credible, comprehensive, and challenging. |
Copied to clipboard
| Challenge: | Experimental results show that unified model outperforms other models that treat encoding and matching separately. |
| Approach: | They evaluate a unified model with Transformer layers for machine reading comprehension . they find that the model learns different modeling strategies compared with previous models . |
| Outcome: | The unified model outperforms models with Transformer layers on the machine reading comprehension task. |
Copied to clipboard
| Challenge: | Existing methods such as Medusa lack adequate information interaction between different drafting heads. |
| Approach: | They propose an enhanced speculative decoding framework that builds upon Medusa and integrates a drafting block capable of parallel inference. |
| Outcome: | The proposed framework outperforms Medusa in terms of head accuracy and latency. |
Copied to clipboard
| Challenge: | Existing humor computing research focuses on content while neglecting interaction relationships in social media. |
| Approach: | They propose a dataset which introduces social context information from social media . they propose 'humor recognition' task and 'horror evaluation task' |
| Outcome: | The proposed model incorporates social context information from social media . it shows that it is efficient and can be used to evaluate humor in real life . |
Copied to clipboard
| Challenge: | Low-rank decomposition methods suffer from accuracy degradation and expensive calibration procedures. |
| Approach: | They propose a fast and accurate, training-free structural compression method based on fine-grained low-rank transformations in the activation space. |
| Outcome: | The proposed method outperforms pruning baselines in generalization and downstream performance while delivering inference speedups. |
Copied to clipboard
| Challenge: | a survey of RAG-based reasoning-based approaches shows that it is not effective for multi-step inferences. |
| Approach: | They map how advanced reasoning optimizes each stage of RAG . they show how retrieved knowledge supply missing premises and expand context for complex inference . |
| Outcome: | The proposed frameworks achieve state-of-the-art across knowledge-intensive benchmarks. |
Copied to clipboard
| Challenge: | Existing methods have primarily treated ABAM as a nested named entity recognition problem, overlooking the need for tailored strategies to effectively address the specific challenges of ABA M tasks. |
| Approach: | They propose a layer-based Hierarchical Enhancement Framework (HEF) for Aspect-Based Argument Mining and introduce three new components to improve the performance and accuracy. |
| Outcome: | Experiments on multiple datasets and tasks verify the effectiveness of the proposed framework and components. |
Copied to clipboard
| Challenge: | Existing GUI Agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. |
| Approach: | They propose an MLLM-based GUI Agent with a two-stage supervised fine-tuning pipeline that enhances GUI understanding and grounding. |
| Outcome: | InfiGUIAgent achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. |
Copied to clipboard
| Challenge: | Existing methods for hallucination detection focus on implicit neural uncertainty or explicit symbolic reasoning, ignoring factual hallucinosities. |
| Approach: | They propose a framework that bridges neural features and symbolic judgments for hallucination detection by leveraging a "meta-judgment" process to map symbolic labels back into the feature space. |
| Outcome: | Extensive experiments on 4 public datasets, across 4 LLMs, against 8 baselines demonstrate the superiority of LaaB. |
Copied to clipboard
| Challenge: | Existing methods to extract information from evidence are unable to grasp relational and logical information among the evidence. |
| Approach: | They propose a graph-based evidence aggregating and reasoning framework to integrate evidence from multiple pieces of evidence. |
| Outcome: | The proposed framework achieves significant performance improvements on a large-scale benchmark dataset. |
Copied to clipboard
| Challenge: | Social media is a key platform for emotional expression, yet deep learning lacks flexibility and interpretability. |
| Approach: | They propose to use Chinese social media to train interpretable mental health instruction datasets to test models' ability to explain their decisions. |
| Outcome: | The proposed models outperform deep learning and LLMs on three mental health downstream tasks and demonstrate their potential for clinical applications. |
Copied to clipboard
| Challenge: | Molecular Relational Learning (MRL) is a promising way to understand interactions between molecular pairs. |
| Approach: | They propose a novel LLM-based multi-modal framework for molecular interaction modeling following Chain-of-Thought (CoT) theory which integrates graphical information of two molecules in pair. |
| Outcome: | The proposed framework integrates graphical information of two molecules in pair. |
Copied to clipboard
| Challenge: | Mathematical reasoning has long been a key benchmark for evaluating large language models. |
| Approach: | They propose a framework that transforms math word problems into scalable tabular reasoning tasks. |
| Outcome: | The proposed framework transforms math word problems into scalable and verified tabular reasoning tasks. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) generate explicit Chain-of-Thought rationales, but often suffer from "overthinking". |
| Approach: | They propose a unified training framework to synthesize concise reasoning chains by identifying tokens that reduce the model’s likelihood of the correct answer. |
| Outcome: | Experiments on deepSeek-R1-Distill-Qwen backbones show that MUTO yields better efficiency-accuracy Pareto frontier. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown strong potential for text-attributed graph (TAG) learning, yet effectively integrating LLM semantics with graph structural information remains challenging. |
| Approach: | They propose a framework for learning Graph-Aware Semantic Embeddings using frozen LLMs. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on node classification and achieves a 5 speedup over fine-tuning-based methods. |
Copied to clipboard
| Challenge: | Empirical results on mathematical reasoning benchmarks substantiate the efficacy of Large language models (LLMs). |
| Approach: | They propose a framework that combines an LLM with a process-based verifier to generate plausible candidates and provide timely process-driven feedback to distinguish desirable and undesirable outputs. |
| Outcome: | Empirical results show that LLM2 improves accuracy on GSM8K and self-consistency increases major@20 accuracy. |
Copied to clipboard
| Challenge: | Recent deep learning methods for MeSH indexing fail to capture complex correlations between terms. |
| Approach: | They propose a model to learn the relationship between MeSH terms using Graph Convolution Network (GCN) they use two biGRUs to learn embedding representations of abstract and title of MeSH index text . |
| Outcome: | The proposed model is competitive with the state-of-the-art models on two datasets. |
Copied to clipboard
| Challenge: | ReflectEvo-460k is a large-scale, comprehensive, self-generated reflection dataset with broadened instructions and diverse multi-domain tasks. |
| Approach: | They propose a pipeline that iteratively generates self-reflection for self-training and a large-scale reflection dataset with broadened instructions and diverse multi-domain tasks. |
| Outcome: | The proposed pipeline improves Llama-3 reasoning ability by up to 71.2% and Mistral by upto 44.4%. |
Copied to clipboard
| Challenge: | Generating synthetic datasets via large language models (LLMs) has emerged as promising approach to improve LLM performance. |
| Approach: | They propose three mitigation strategies to mitigate bias inheritance in LLMs by analyzing real and LLM-augmented data. |
| Outcome: | The proposed methods can work differently on different tasks and biases. |
Copied to clipboard
| Challenge: | Past work in NLP examined the task of goal-step inference for textual goals . wikiHow dataset shows that goal-step inference is challenging for state-of-the-art models . |
| Approach: | They propose a task where a model is given a textual goal and must choose which of four images represents a plausible step towards that goal. |
| Outcome: | The proposed task is challenging for state-of-the-art multimodal models and can be transferred to other datasets. |
Copied to clipboard
| Challenge: | Large Multimodal Models have demonstrated strong performance on vision-language benchmarks, yet current evaluations focus on single-image reasoning. |
| Approach: | STRIPCIPHER is a benchmark designed to evaluate model ability on understanding implicit narratives in silent comics. |
| Outcome: | STRIPCIPHER is a high-quality, human-annotated dataset featuring fine-grained annotations and comprehensive coverage of varying difficulty levels. |
Copied to clipboard
| Challenge: | AutoMonitor-Bench evaluates the reliability of LLM-based misbehavior monitors across diverse tasks and failure modes. |
| Approach: | They introduce AutoMonitor-Bench, a benchmark designed to evaluate misbehavior monitors across diverse tasks and failure modes. |
| Outcome: | The new benchmark evaluates the reliability of LLM-based misbehavior monitors across tasks and failure modes. |
Copied to clipboard
| Challenge: | Existing Process Reward Models (PRMs) are vulnerable to reward hacking and require expensive, large-scale annotation of reasoning steps. |
| Approach: | They propose a reward model approach which evaluates both individual and consecutive reasoning steps from fine-grained and coarse-grounded level. |
| Outcome: | Empirical results show that the proposed model performs better than existing PRMs and is more robust than existing models. |
Copied to clipboard
| Challenge: | Recent years have seen a rapid growth of interest in building task-oriented dialogue systems. |
| Approach: | They propose a retrieve-and-memorize framework to deal with unbalanced distribution of system actions in dialogue datasets. |
| Outcome: | The proposed framework achieves competitive performance among state-of-the-art models on a large-scale task-oriented dialogue dataset. |
Copied to clipboard
| Challenge: | Existing approaches to improve performance of deep neural models are limited by the nature of spurious patterns in the data. |
| Approach: | They propose to use augmented data to generate spurious patterns in NLP models . they propose to generate counterfactual data for data augmentation and explanation . |
| Outcome: | The proposed approach improves performance on augmented data and on human-generated data. |
Copied to clipboard
| Challenge: | Open Information Extraction (OIE) aims to extract structured information from text without the limitations of close ontology. |
| Approach: | They propose a method to assign ground truth labels to parallelly generated tuple proposals . they leverage intersection-over-union (IoU) as assignment quality measurement . |
| Outcome: | The proposed method outperforms the state-of-the-art models on three benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for training large language models require additional annotations to adjust to shifted distributions. |
| Approach: | They propose an algorithm that allows LLMs and reward models to update alternatively via a min-max game to improve their alignment. |
| Outcome: | The proposed framework improves existing alignment baselines in terms of LLM helpfulness and harmlessness. |
Copied to clipboard
| Challenge: | Few-shot text matching is a more practical technique to determine whether two texts are semantically identical. |
| Approach: | They propose a pluggable prompt learning method for few-shot text matching . they use the semantics of instances to regulate the effects of the gate on the prompt tokens . |
| Outcome: | The proposed method outperforms baselines on MRPC and QQP. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has re-emerged as a natural approach for training interactive LLM agents in real-world environments. |
| Approach: | They propose a variant that operates on a turn-level MDP formulation, instead of the commonly used token-level one. |
| Outcome: | The proposed method is more robust than the widely used GRPO algorithm and more efficient than token-level MDPs. |
Copied to clipboard
| Challenge: | Existing benchmarks for musical score understanding are narrow in scope, focusing on isolated fragments, short excerpts, or multiple-choice formulations, rather than supporting holistic reasoning over entire scores. |
| Approach: | They propose a benchmark for score-level musical understanding across textual and visual modalities. |
| Outcome: | The musical score understanding benchmark contains 1,800 question-answer pairs from works by Bach, Beethoven, Chopin, Debussy, and others. |
Copied to clipboard
| Challenge: | Existing benchmarks for agentic programming in long-horizon command-line interface tasks are limited by short task horizons, data contamination from GitHub scraping, and a lack of fine-grained evaluation metrics. |
| Approach: | They propose a benchmark to evaluate agentic capabilities across long-horizon command-line interface tasks. |
| Outcome: | The proposed benchmarks cover four engineering categories: from scratch, feature addition, bug fixing, and refactoring. |
Copied to clipboard
| Challenge: | Existing approaches to visual-language understanding lack unified tokenization for images and videos . lack of unified visual representations makes it difficult to learn multi-modal interactions from poor projection layers. |
| Approach: | They propose to unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM. |
| Outcome: | The proposed model outperforms Video-ChatGPT on image benchmarks and on 9 image benchmark benchmarks. |
Copied to clipboard
| Challenge: | Existing studies treat SLMs as student models and use long-form Chains-of-Thought (CoTs) as supervision signals for Supervised Fine-Tuning (SFT). Existing research focuses on distilling reasoning ability from LLMs to enhance the mathematical reasoning performance of small-scale models. |
| Approach: | They propose a framework that refines teacher CoTs through an error-aware reflection process to enable the student model to construct more tailored teacher Cots. |
| Outcome: | Experiments on multiple mathematical reasoning benchmarks show that ORION improves performance by more than 2% over all baselines. |
Copied to clipboard
| Challenge: | Recent approaches in Incomplete Utterance Rewriting (IUR) fail to capture the source of important words, introducing words from irrelevant utterances. |
| Approach: | They propose a framework to capture the multi-granularity of semantic information and fetch the relevant utterance. |
| Outcome: | The proposed framework outperforms state-of-the-art models on two benchmark datasets . it can capture the source of important words and fetch the relevant utterance . |
Copied to clipboard
| Challenge: | Existing methods to regularize dropout are consistency training and dropout is a problem in many pre-trained neural language models. |
| Approach: | They propose a layer-wise regularized dropout technique which regularizes dropout at the output layer using consistency training. |
| Outcome: | The proposed model can be regarded as a "self-distillation" framework, in which each sub-model generated by dropout is the other's "teacher" model and "student" model. |
Copied to clipboard
| Challenge: | Existing approaches to reward modeling in reinforcement learning tasks are limited when dealing with ambiguous preferences. |
| Approach: | They propose to use AAM to dynamically calibrate preference margins using the Bradley-Terry model's internal parameter knowledge to improve reward modeling in subjective tasks. |
| Outcome: | The proposed approach improves reward modeling by dynamically calibrating preference margins using the model’s internal parameter knowledge. |
Copied to clipboard
| Challenge: | Existing approaches focus on text information, but authors and overall ratings are ignored, both of which are proved to be significant on interpreting the sentiments of different aspects. |
| Approach: | They propose a hierarchical user-aspect rating network model to consider user preference and overall ratings jointly. |
| Outcome: | The proposed model can predict aspects of a product in two real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods for news comment generation have not been well studied. |
| Approach: | They propose a “read-attend-comment” procedure for automatic news comment generation and formalize it with a reading network and a generation network. |
| Outcome: | The proposed procedure outperforms existing methods in terms of automatic evaluation and human judgment on two public datasets. |
Copied to clipboard
| Challenge: | Existing research on information-seeking conversations is stymied by the lack of training data. |
| Approach: | They propose to use autoconv for synthetic conversation generation to capture the characteristics of the information-seeking process and fine tune an LLM with a few human conversations to generate synthetic conversations with high quality. |
| Outcome: | The proposed model improves on two commonly-used datasets and alleviates the dependence on human annotation. |
Copied to clipboard
| Challenge: | Existing methods for capturing instruction-following complexity rely on single-dimensional signals, but they fail to capture complexity across diverse fields. |
| Approach: | They propose three foundational metrics that leverage Multi-LLMs wisdom to capture instruction-response pair characteristics and propose CrowdSelect, an integrated metric incorporating a clustering-based approach to maintain response diversity. |
| Outcome: | The proposed metrics outperform existing models on MT-bench and Arena-hard and show improvements of 4.81% on full and LoRA fine-tuning. |
Copied to clipboard
| Challenge: | Despite of significant achievements in improving instruction-following capabilities of large language models, the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge. |
| Approach: | They construct multi-turn instruction with 1.1K high-quality multi-turned conversations using the human-in-the-loop approach and examine their capabilities. |
| Outcome: | The proposed model shows that it is difficult to integrate multiple turns and balance competing objectives when instructions intersect or conflict. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks often rely on single-frame assessments, which can lead to outcome-hacking. |
| Approach: | They propose a process-aware evaluation paradigm that uses a hierarchical rubric to evaluate the validity of the intermediate steps and the final result. |
| Outcome: | The proposed model achieves POC@1.0 only about 20% and exhibits significant outcome-hacking. |
Copied to clipboard
| Challenge: | Existing methods for evaluating code large language models assume access to proprietary training corpora or use external reference sets with manually tuned, non-generalizable thresholds. |
| Approach: | They propose a framework for self-referential leakage detection for gray-box and black-box settings. |
| Outcome: | The proposed framework improves average F1 by 21.52 points in the gray-box setting and 14.46 points in black-box settings over strong baselines. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) have improved generalization performance but the out-of-distribution (OOD) generalization problem remains a challenge in many NLP tasks. |
| Approach: | They propose to create a benchmark for evaluating out-of-distribution (OOD) generalization in NLP models. |
| Outcome: | The proposed benchmarks highlight the importance of OOD robustness and provide insights on how to measure it and improve it. |
Copied to clipboard
| Challenge: | Existing approaches to integrate the recommendation function and dialog generation function smoothly are lacking. |
| Approach: | They propose to integrate dialog context for recommendation and dialog generation better using a pre-trained language model and an item metadata encoder to integrate the recommendation and dialogue generation. |
| Outcome: | The proposed architecture improves the integration of recommendation and dialog generation functions. |
Copied to clipboard
| Challenge: | Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), but its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content and poor OCR quality. |
| Approach: | They propose a corruption-robust training paradigm that surpasses existing strategies for mitigating the effects of corrupted data. |
| Outcome: | The proposed training paradigm surpasses existing strategies for mitigating the effects of corrupted data. |
Copied to clipboard
| Challenge: | Existing methods for in-context learning with large language models focus on using correct or negative examples, ignoring the potential value of incorrect or negative samples. |
| Approach: | They propose a few-shot technique that leverages both correct and incorrect sample constructions to create in-context learning demonstrations. |
| Outcome: | The proposed technique outperforms previous few-shot in-context learning methods on a broad spectrum of related tasks. |
Copied to clipboard
| Challenge: | Vision-Language Models struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. |
| Approach: | AgentThink integrates Chain-of-Thought reasoning with dynamic, agent-style tool invocation for autonomous driving tasks. |
| Outcome: | Experiments on the DriveLMM-o1 benchmark show AgentThink significantly boosts overall reasoning scores by 53.91% and enhances answer accuracy by 33.54% . |
Copied to clipboard
| Challenge: | Existing studies on social media use tags to profile users, but we have found that sentence-level self-introductions are more natural and engaging. |
| Approach: | They propose a novel topic-guided encoder-decoder framework that uses a user's tweeting history to generate a short sentence outlining their personal interests. |
| Outcome: | The proposed framework outperforms existing encoder-decoder models on a large-scale Twitter dataset and shows that it is more natural and engaging than previous approaches. |
Copied to clipboard
| Challenge: | Existing memory systems represent conversation history as unstructured embedding vectors, retrieving information through semantic similarity. |
| Approach: | They propose a bio-inspired memory architecture that models memory as a dynamic graph with Hebbian learning dynamics. |
| Outcome: | The proposed architecture leverages both semantic similarity and learned associations . it can be used to build a bio-inspired memory graph with Hebbian learning dynamics . |
Copied to clipboard
| Challenge: | Large language models (LLMs) often struggle with complex reasoning tasks due to the vast reasoning space inherent in the complexity and inherent ambiguities of natural languages. |
| Approach: | They propose a mixture-of-search-agents paradigm that integrates diverse reasoning pathways by combining independent exploration and iterative refinement among multiple LLMs. |
| Outcome: | The proposed approach improves performance over single-agent and multi-agend baselines in complex mathematical and commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | Existing methods to enhance LLMs with knowledge graphs have limited results . knowledge graph question answering (KGQA) provides interpretable reasoning for large language models . |
| Approach: | They propose a framework for KG-enhanced LLM based on question decomposition and atomic retrieval . they propose question decomposing tree as framework for LLM reasoning . |
| Outcome: | The proposed framework outperforms existing reasoning-based baselines on KGQA datasets. |
Copied to clipboard
| Challenge: | Existing models that require task labels or performance trade-offs are susceptible to catastrophic forgetting. |
| Approach: | They propose a representation-aware model merging framework for continual learning without access to historical data. |
| Outcome: | The proposed framework outperforms baselines in knowledge retention and generalization across five NLP tasks and multiple continual learning scenarios. |
Copied to clipboard
| Challenge: | a new spoken dialogue system with single-stage training is demonstrating its low latency and high quality . SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens . |
| Approach: | They propose a timbre-controllable, end-to-end voice interaction system with single-stage training. |
| Outcome: | The proposed system outperforms previous models on 4 GPUs with limited data. |
Copied to clipboard
| Challenge: | LongLeader aims to assess different LLMs' long-context comprehension abilities . long-constext comprehension is a key bottleneck for many use cases . |
| Approach: | They propose a leaderboard to assess different LLMs' long-context comprehension abilities . they offer open-source access to the benchmarks and maintain a dedicated website . |
| Outcome: | The proposed model assesses different LLMs on selected benchmarks and provides open-source access to the benchmarks. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods rely on unidirectional, feed-forward pipelines . a minor retrieval error or logical mismatch at an early hop can become a silent failure . |
| Approach: | They propose a framework for closed-loop post-edit reasoning that uses a Critic agent to verify coherence and step-wise correctness. |
| Outcome: | Experiments on MQuAKE-2002 and MQuADE-hard show that CARE effectively mitigates error propagation . a minor retrieval error or logical mismatch at an early hop can become a silent failure . |
Copied to clipboard
| Challenge: | Existing efforts to alleviate hallucination in chatbots require additional training and data annotation. |
| Approach: | They propose a Citation-Enhanced Generation approach that uses retrieval argumentation to generate citations and a natural language inference-based citation generation module to generate content. |
| Outcome: | The proposed method outperforms state-of-the-art methods on three benchmarks. |
Copied to clipboard
| Challenge: | Existing studies in text-to-SQL do not require generating complex SQL queries with multiple clauses or sub-queries. |
| Approach: | They propose a syntax tree network to address the complex text-to-SQL generation task. |
| Outcome: | The proposed model outperforms the current state-of-the-art model by 9.5% on a large text-to-SQL corpus. |
Copied to clipboard
| Challenge: | Distributed LLMs avoid raw inputs by transmitting intermediate hidden states, a practice widely assumed to preserve privacy. |
| Approach: | They propose a distributed inference framework that transmits intermediate hidden states to avoid sending raw inputs by exposing sensitive user attributes. |
| Outcome: | The proposed approach achieves Top-1 accuracy of 0.997 on CMS, 0.980 on Skytrax, and 0.986 on ECHR. |
Copied to clipboard
| Challenge: | Existing approaches to generate relevance judgments are limited due to dynamic nature of query distributions. |
| Approach: | They propose a self-evolving relevance model approach to generalize queries to practical search scenarios . they use a multi-agent sample miner and a relevance annotator to generate reliable labels . |
| Outcome: | The proposed approach can achieve significant performance gains on a large-scale industrial platform, validated by offline multilingual evaluations and online testing. |
Copied to clipboard
| Challenge: | a gap exists between human embodied logic and machine statistical learning . authors: models internalize statistical patterns or mimic static recognition . |
| Approach: | They propose a cognitive linguistic benchmark to test whether large language models internalize statistical logic or not . they find that models function as "Super-Associators" expert at static recognition yet fail at causal reasoning . |
| Outcome: | The proposed model fails at causal reasoning and has a high-fidelity concept representation but lacks transformational operators essential for true relational understanding. |
Copied to clipboard
| Challenge: | Long-context language models exhibit position bias, also known as "lost in the middle" research shows that even long-contemporary LLMs fail to utilize all context information effectively . |
| Approach: | They propose a method to mitigate position bias by scaling positional hidden states . they propose to use a channel of hidden states to modify positional Hidden states a LCLM's positional bias . |
| Outcome: | The proposed method can improve performance by 15.2% in a "lost in the middle" benchmark. |
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning (RL) are limited by poor data efficiency and weak generalization. |
| Approach: | They propose a novel architecture that integrates large language models into episodic RL. |
| Outcome: | The proposed architecture achieves 2–6 higher data efficiency than baselines and is the only method to solve complex tasks like UnlockLocal with over 90% success. |
Copied to clipboard
| Challenge: | Existing methods to extract relation facts from limited labeled corpora are laborintensive to obtain . Existing approaches use self-training to generate pseudo labels that will cause gradual drift problem or leverage meta-learning scheme which does not solicit feedback explicitly. |
| Approach: | They propose a Gradient Imitation Reinforcement Learning method to encourage pseudo label data to imitate gradient descent direction on labeled data and bootstrap its optimization capability through trial and error. |
| Outcome: | The proposed method handles two major scenarios in low-resource relation extraction when no unlabeled data is available. |
Copied to clipboard
| Challenge: | Existing work on stance detection tasks using large language models shows promising results, but it may not be able to provide detailed background knowledge. |
| Approach: | They propose a method which leverages the generated experienced experts and lets LLMs reason in a semi-parametric way. |
| Outcome: | The proposed method outperforms methods with self-consistency reasoning and reduces bias. |
Copied to clipboard
| Challenge: | Existing approaches to adapt image-focused models for video understanding have not been successful in analyzing long video sequences. |
| Approach: | They propose a video instruction dataset that outperforms existing video instruction data for fine-tuning MLLMs by incrementally increasing input context length. |
| Outcome: | The proposed model outperforms existing models on video benchmarks and outperformed proprietary models on VideoMME even with a compact 7B model. |
Copied to clipboard
| Challenge: | Existing unsupervised paraphrase generation methods require large-scale, manually annotated paraphrase datasets, which are labor-intensive to build. |
| Approach: | They propose a self-supervised pseudo-data construction method that generates diverse pseudo-paraphrases in distinct surface structures for a given sentence. |
| Outcome: | The proposed method generates diverse pseudo-paraphrases in distinct surface structures for a given sentence. |
Copied to clipboard
| Challenge: | Existing studies focus on fusing different features but ignore the challenge of modality heterogeneity. |
| Approach: | They propose a text-guided fusion module with novel Sparse-Attention to reduce the negative impacts of redundant visual elements and a sentiment-based congruity constraint task to calibrate the feature shift in the representation space. |
| Outcome: | The proposed model is competitive against existing methods and achieves state-of-the-art results on two public benchmark datasets. |
Copied to clipboard
| Challenge: | lexicalist linguistic theories assume argument structure is predictable from meaning of verbs . construction grammarians propose argument structure constructions distinct from verbs. |
| Approach: | They adapt psycholinguistic studies to probe for the existence of argument structure constructions in Transformer-based language models. |
| Outcome: | The proposed method could be used to probe argument structure constructions in LMs . the study shows that LM learners prefer grouping by construction over verb grouping . |
Copied to clipboard
| Challenge: | a challenge for aspect term extraction is to extract phrase-level aspect terms . a constituency lattice structure is constructed using the span annotations of constituents of a sentence . |
| Approach: | They propose to incorporate the span annotations of constituents of a sentence to leverage syntactic information in neural network models. |
| Outcome: | The proposed model outperforms existing models on two benchmark datasets. |
Copied to clipboard
| Challenge: | Recent approaches demonstrate that MLLMs can be adapted into competitive embedding models via large-scale contrastive learning. |
| Approach: | They propose a compressed pre-training phase which serves as a warm-up stage for contrastive learning. |
| Outcome: | The proposed model achieves state-of-the-art among MLLMs of comparable size on the MMEB, realizing optimization in both efficiency and effectiveness. |
Copied to clipboard
| Challenge: | Recent advances in preference optimization have demonstrated significant potential for improving mathematical reasoning capabilities in large language models. |
| Approach: | They propose a framework that establishes two quantitative metrics for preference selection: surface-level answer correctness and intrinsic token-level probability consistency. |
| Outcome: | The proposed framework outperforms existing outcome-only criterion approaches across a diverse range of LLMs and benchmarks. |
Copied to clipboard
| Challenge: | Solving expert-level multimodal tasks requires strong user query understanding, domain-specific knowledge, and advanced reasoning abilities. |
| Approach: | They propose a benchmark of open-ended user queries encapsulating professional expertise and advanced reasoning. |
| Outcome: | The proposed benchmark is publicly accessible at TBC. |
Copied to clipboard
| Challenge: | Existing methods for large language model reasoning suffer from exploration collapse due to the semantic homogeneity of random rollouts. |
| Approach: | They propose to use latent policy optimization via iterative information bottleneck to optimize reasoning trajectories by diversifying reasoning . |
| Outcome: | Empirical results show that the proposed method achieves state-of-the-art performance with margins of up to 5.3% in accuracy and 7.4% in diversity metrics. |
Copied to clipboard
| Challenge: | Existing word analogy datasets rely on handcrafted words with only dozens of predefined relations. |
| Approach: | They present a commonsense word analogy dataset with 90,505 analogies . they use an ontology that annotates 88K Chinese words with their structured sense definitions and English translations. |
| Outcome: | The proposed dataset shows that word representations embed commonsense knowledge. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have impressive multimodal abilities but remain prone to multilingual object hallucination. |
| Approach: | They propose a cross-lingual attention intervention method to mitigate multilingual object hallucination in LVLMs by aligning attention patterns. |
| Outcome: | The proposed method improves 13.56% (up to 30%) on the POPE and 21.75% on the hallucination subsets across languages. |
Copied to clipboard
| Challenge: | Existing methods for knowledge distillation use Chain-of-Thought (CoT) and answer pairs, but they lack appropriate supervision signals. |
| Approach: | They propose a framework that decouples CoT and answer supervision . the framework applies semantic similarity constraints while maintaining strict literal matching for the answer . |
| Outcome: | The proposed framework decouples CoT and answer supervision while maintaining strict literal matching for the answer. |
Copied to clipboard
| Challenge: | Existing work on extending specialized agents to multi-agent systems is dependent on human-designed frameworks, limiting the functional scope and scalability of agent systems. |
| Approach: | They propose a generic method to automatically extend specialized agents to multi-agent systems via evolutionary algorithm . they consider existing agent frameworks as the initial individual and apply evolutionary operators to generate multiple agents with diverse settings. |
| Outcome: | The proposed method can extend specialized agents to multi-agent systems . it can generate multiple agents with diverse settings, and improves performance across tasks . |
Copied to clipboard
| Challenge: | Scientific information extraction (SciIE) is a key bottleneck for turning unstructured papers into computable knowledge bases. |
| Approach: | They propose a scientific information extraction framework that solves a Sudoku problem as a progressive filling problem. |
| Outcome: | The proposed framework outperforms the GPT-4o model on a document-level adjuvant dataset. |
Copied to clipboard
| Challenge: | Large language models require a balance between efficiency and performance. |
| Approach: | They propose a low-rank compression technique that reduces non-essential parameters by decomposing weight matrices into products of two low-ranked matrici. |
| Outcome: | The proposed method outperforms existing pruning and low-rank compression techniques in maintaining model performance at the same compression ratio. |
Copied to clipboard
| Challenge: | Current approaches to task-oriented dialogue systems integrate knowledge retrieval and response generation, which poses scalability challenges when dealing with extensive knowledge bases. |
| Approach: | They propose a retriever-generator architecture that harnesses a retrieval and a generator to generate system responses by using feedback from the generator as pseudo-labels. |
| Outcome: | The proposed architecture shows superior performance on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to solve knowledge hypergraph link prediction problem are limited by their ability to generate chain-of-thought (CoT) representations. |
| Approach: | They propose a structure-aware approach that models multi-hop structural reasoning as a depth-sensitive progressive evidence accumulation process. |
| Outcome: | Experiments on three real-world datasets show that HyperCoT outperforms strong n-ary KGC baselines while yielding interpretable multi-hop reasoning traces. |
Copied to clipboard
| Challenge: | Empathy is a key trait of everyday human conversations. |
| Approach: | They propose a serial encoding and Emotion-Knowledge interaction method for empathetic dialogue generation which is more sensitive to emotion dynamics in conversations. |
| Outcome: | The proposed method outperforms baseline evaluations on the utterance-level annotated EMPATHETICDIALOGUES. |
Copied to clipboard
| Challenge: | Existing approaches to early exit reasoning often rely on handcrafted or empirical indicators that are unreliable and impractical. |
| Approach: | They propose a framework that allows LRMs to assess the sufficiency of its chain-of-thought and determine the optimal point for early exit. |
| Outcome: | The proposed framework reduces reasoning length by 28.9%–34.9% with minimal performance loss, effectively mitigating overthinking. |
Copied to clipboard
| Challenge: | Existing methods focus on enhancing multi-scale clip representations but lack robust data alignment . inherent data uncertainty renders PRVR vulnerable to distractor videos with spurious similarities . |
| Approach: | proposed framework for partially relevant video retrieval aims to retrieve untrimmed videos partially relevant to a given query. |
| Outcome: | The proposed framework can be seamlessly integrated into existing architectures. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) can grasp the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. |
| Approach: | They propose a framework for multimodal large language models to grasp the intention of a question and decompose it into a series of visual recognition sub-tasks to find out the answer. |
| Outcome: | The proposed framework improves the accuracy of complex video-related questions by 29.6% and 17.2% on CVQA and the existing VQA datasets. |
Copied to clipboard
| Challenge: | Existing language evaluation benchmarks for English are limited to English . lack of such benchmarks makes it difficult to replicate success in other languages . |
| Approach: | They introduce a large-scale Chinese language understanding evaluation benchmark . the benchmark uses a set of current state-of-the-art pre-trained Chinese models . |
| Outcome: | The first large-scale Chinese Language Understanding Evaluation (CLUE) benchmark is released . the benchmark evaluates models across a wide range of tasks on original Chinese text . existing language evaluation benchmarks are mostly limited to English . |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is the task of identifying spans that belong to particular categories, such as person, location, organization, etc. |
| Approach: | They propose a method that integrates named entity’s type information into BERT by an adapter layer and integrates it into a gazetteer. |
| Outcome: | The proposed method outperforms baselines in multiple corpus. |
Copied to clipboard
| Challenge: | Recent studies on Chinese Word Segmentation (CWS) have focused on the cross-domain scenarios, but there is a high cost of manually annotating high-quality data. |
| Approach: | They propose to explicitly mine word boundaries from parallel speech-text data by using the Montreal Forced Aligner toolkit to perform character-level alignment on speech- text data. |
| Outcome: | The proposed approach is based on character-level alignment on speech-text data and a robust complete-then-train (CTT) strategy. |
Copied to clipboard
| Challenge: | Scaling laws in language modeling quantify training loss as a function of dataset size and model parameters, but neglect the critical role of data quality in model generalization. |
| Approach: | They propose to use effective training tokens as a combination of text diversity and syntheticity as measured by a teacher model to calculate scaling laws. |
| Outcome: | The proposed term effective training tokens is a combination of two readily-computed indicators of text diversity and syntheticity as measured by a teacher model. |
Copied to clipboard
| Challenge: | None. None.. None! |
| Approach: | None. None.. None! |
| Outcome: | None. None. No. : |
Copied to clipboard
| Challenge: | Existing evaluation methods for human-machine interactions are static and can be misleading. |
| Approach: | They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement. |
| Outcome: | The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment. |
Copied to clipboard
| Challenge: | Existing supervised text classifications require a large number of manually labeled documents. |
| Approach: | They develop a pseudo-label based dataless Naive Bayes classifier with seed words . they initialize pseudo-labels for each document using seed word occurrences . |
| Outcome: | The proposed classifier outperforms traditional supervised text classification algorithms with seed words on an imbalanced dataset. |
Copied to clipboard
| Challenge: | Existing methods for developing compact and efficient large language models lack token-level dependencies and linguistic diversity. |
| Approach: | They propose a logits-based fine-tuning framework that integrates supervised learning and knowledge distillation to build enriched training targets using teacher logits and ground truth labels. |
| Outcome: | The proposed method outperforms existing methods on a large-scale logits dataset and a series of science-focused models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are susceptible to adversarial attacks such as jailbreaking, which can elicit harmful or unsafe behaviors. |
| Approach: | They propose a multilingual guardrail with reasoning for prompt classification that integrates culturally and linguistically nuanced variants and supervised fine-tuning. |
| Outcome: | The proposed guardrail outperforms baselines across in-domain and out-of-domain languages by more than 15%. |
Copied to clipboard
| Challenge: | Text-to-Video (T2V) generation is a challenge under complex scenarios. |
| Approach: | They propose a scenario-aware and self-correcting multi-agent prompt refinement framework for T2V prompting. |
| Outcome: | The proposed framework improves text-to-video alignment and overall generation quality under complex scenarios. |
Copied to clipboard
| Challenge: | Exemplar-Guided Paraphrase Generation (EGPG) aims to generate a target sentence which conforms to the style of the given exemplar while encapsulating the content information of the source sentence. |
| Approach: | They propose a method to generate a target sentence which conforms to the style of the given exemplar while encapsulating the content information of the source sentence. |
| Outcome: | The proposed method can generate a target sentence which conforms to the style of the given exemplar while encapsulating the content information of the source sentence. |
Copied to clipboard
| Challenge: | Knowledge graphs are a useful tool for organizing complex data in knowledge-intensive domains. |
| Approach: | They propose an expandable framework that combines structured domain texts with advanced semantic techniques to create a tree-like graph from textbooks. |
| Outcome: | The proposed framework surpasses competing methods in the text-Annotated dataset with high scores on the Text-Annalytated data. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated that In-Context Learning (ICA) can align Large Language Models (LLMs) with human preferences without requiring parameter adjustments. |
| Approach: | They investigate the effectiveness of each part in enabling ICA to function effectively and examine how variants in these parts impact alignment performance. |
| Outcome: | The proposed model can comprehend human instructions without parameter adjustments. |
Copied to clipboard
| Challenge: | Existing approaches to training LLMs at ultra-low precisions suffer from convergence instability and substantial training costs. |
| Approach: | They propose a progressive QAT framework with outlier channel splitting to address these issues . they use nested structure of integer quantization grids to enable a "train once, deploy any precision" paradigm . |
| Outcome: | The proposed framework outperforms baselines on both Llama2/3 and W2A16, with an 11 speedup over BF16. |
Copied to clipboard
| Challenge: | Multi-domain dialogue state tracking is a challenge for task-oriented dialogue systems . domains and slots are aggregated into a single query to generate domain-slot specific representations . |
| Approach: | They propose to disentangle domain-slot attention for multi-domain dialogue state tracking by separating query about domains and slots from the attention component. |
| Outcome: | The proposed approach outperforms the standard multi-head attention with aggregated domain-slot query. |
Copied to clipboard
| Challenge: | Existing approaches for spatial reasoning in text overlook the gap between natural language and symbolic structures. |
| Approach: | They propose a novel depth-wise Graph Neural Network to aggregate spatial information over the depth dimension instead of the breadth dimension of the graph. |
| Outcome: | The proposed model outperforms existing methods on two multi-hop spatial reasoning datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have strong reasoning and tool-use capabilities, yet fail in real-world tool-interactions due to incorrect parameterization, poor tool selection, or misinterpretation of user intent. |
| Approach: | They propose a curriculum-inspired framework that leverages structured reasoning templates to guide LLMs through more deliberate step-by-step instructions for generating function calls. |
| Outcome: | The proposed framework reduces tool-use errors and improves interpretability and transparency of tool-using agents. |
Copied to clipboard
| Challenge: | Existing methods to integrate Chain-of-Thought into spoken dialogue models incur prohibitive latency. |
| Approach: | They propose a Streaming Masking Mechanism to ensure uninterrupted audio streaming . they use a quadruple-constraint system to reconstruct logical atomicity . |
| Outcome: | Experimental results show that Dual-Reasoner improves speech generation performance with low latency. |
Copied to clipboard
| Challenge: | Existing benchmarks assess factual accuracy in isolated queries but fail to evaluate LLMs’ resilience to misinformation in interactive settings. |
| Approach: | MisinfoBench is a benchmark designed to assess LLMs’ ability to discern, resist, and reject misinformation. |
| Outcome: | MisinfoBench assesses large language models’ ability to discern, resist, and reject misinformation in interactive settings. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning. |
| Approach: | They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios. |
| Outcome: | The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly entrusted with high-stakes decisions that affect human welfare. |
| Approach: | They evaluate 20 state-of-the-art Large language models (LLMs) and 20 LLM dictators to create a social welfare function benchmark. |
| Outcome: | The proposed model creates dilemma between maximizing collective efficiency and ensuring distributive fairness. |
Copied to clipboard
| Challenge: | Existing language model agents excel in planning and reasoning, but lack creativity in unfamiliar environments. |
| Approach: | They propose a benchmark suite of room escape game environments to challenge agents with creative reasoning, unconventional tool use and iterative problem-solving to uncover implicit goals. |
| Outcome: | The proposed framework can perform with 40% fewer steps and hints and performs robustly across difficulty levels. |
Copied to clipboard
| Challenge: | CPsyExam prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios. |
| Approach: | They propose a psychological benchmark, CPsyExam, constructed from questions from Chinese examination systems. |
| Outcome: | The proposed benchmark prioritizes psychological knowledge and case analysis separately, recognizing the significance of applying psychological knowledge to real-world scenarios. |
Copied to clipboard
| Challenge: | Existing methods focus on weakly aligning uni-modal representations and generatively data augmentation techniques, but they ignore the potential impact of event role information on MEAE. |
| Approach: | They propose a cross-modal variational role hypergraph network via semantic enhancement to model high-order role correlations among cross-mod arguments in multi-modal documents. |
| Outcome: | The proposed method achieves a 6.9% improvement in F1-score on the M2E2 benchmark compared to current state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing fact-checking methods that use large language models often generate subtle factual errors. |
| Approach: | They propose a fact-checking framework that uses extracted knowledge graphs to enhance text representation. |
| Outcome: | GraphCheck outperforms existing specialized fact-checkers on seven benchmarks spanning general and medical domains . Graph Neural Networks process extracted knowledge graphs as a soft prompt, enabling efficient fact- checking in a single inference call. |
Copied to clipboard
| Challenge: | Existing LLM-based evaluation methods fail to accurately identify error spans and assess their severity. |
| Approach: | They propose a Hierarchical Multi-Agent Framework for Machine Translation Evaluation based on the MQM error typology and a hierarchical multi-agent system enabling granular evaluation of subtype errors. |
| Outcome: | The proposed framework outperforms baselines in error span detection and severity assessment. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) are becoming tool-interactive agents . quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data . et al.: a new approach to scale trajectory diversity improves tool-use generalization . |
| Approach: | They propose a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. |
| Outcome: | Experiments on general tool-use benchmarks and code agent tasks show that TDScaling improves tool-user generalization and inherent coding proficiency. |
Copied to clipboard
| Challenge: | Existing methods of model editing and knowledge updating add additional network parameters, knowledge bases, knowledge base, and model parameters. |
| Approach: | They propose a new paradigm for fine-tuning called F-Learning that employs parametric arithmetic to facilitate the forgetting of old knowledge and learning of new knowledge. |
| Outcome: | The proposed model outperforms existing models on two datasets and is comparable to full fine-tuning and LoRA fine-uning. |
Copied to clipboard
| Challenge: | Few-Shot Document-Level Relation Extraction (FSDLRE) aims to develop models capable of generalizing to new categories with minimal support examples. |
| Approach: | They propose a meta-training approach to train Large Language Models to improve their ICL capabilities . they construct simulated episodes using relation types that do not overlap with test corpus . |
| Outcome: | Experimental results show that the proposed approach outperforms baseline models on few-shot tasks. |
Copied to clipboard
| Challenge: | Existing automatic prompt optimization methods fail to optimize prompts and decoding hyperparameters within a unified framework to achieve stable global improvements. |
| Approach: | They propose a dynamic prompt optimization framework for complex reasoning that unifies prompt templates and decodes hyperparameters as inheritable agent configurations. |
| Outcome: | Experiments on multiple mathematical and hybrid reasoning benchmarks show that Agent-GWO improves accuracy and stability over existing prompt optimization methods. |
Copied to clipboard
| Challenge: | Existing studies largely overlook the interplay between logical complexity and semantic complexity, limiting their robustness under abstract propositions, ambiguous contexts, and conflicting stances. |
| Approach: | They propose a semiotic-square-guided framework that integrates automated deduction with reflective verification to manage logical complexity across deeper reasoning chains. |
| Outcome: | The proposed framework achieves state-of-the-art performance on RepublicQA with 6.25% average gain, and generalizes well to four mainstream logical reasoning benchmarks with an additional 7.05% improvement. |
Copied to clipboard
| Challenge: | Existing methods for instruction data selection have limitations such as relying on fragile external APIs, being affected by biases in GPT models, or reducing the diversity of the selected instruction dataset. |
| Approach: | They propose an industrial-friendly, expert-aligned and diversity-preserved instruction data selection method: Clustering and Ranking (CaR). |
| Outcome: | The proposed method outperforms Alpaca's existing methods by 32.1% in GPT-4 evaluations. |
Copied to clipboard
| Challenge: | Existing methods for automating vulnerability repair suffer from syntactic overfitting . nvd published 49,230 Common Vulnerabilities and Exposures (CVE) records in 2025 alone . |
| Approach: | They propose a semantic-aware reward framework that optimizes for code semantic equivalence rather than lexical mimicry. |
| Outcome: | The proposed framework outperforms state-of-the-art frameworks on repository-level splits . it incorporates expert-aligned reasoning mechanism that grounds patch generation in structured diagnosis. |
Copied to clipboard
| Challenge: | Existing models with incomplete utterances have too large search space, resulting in poor quality of rewriting results. |
| Approach: | They propose a 2-phase rewriting framework which predicts empty slots in the utterance that need to be completed and generates the part to be filled into each position. |
| Outcome: | The proposed framework achieves state-of-the-art results on several public rewriting datasets. |
Copied to clipboard
| Challenge: | Existing Large Language Models suffer from "Reasoning Collapse" on mathematical reasoning tasks where stochastic sampling produces lexical variations of the same erroneous logic rather than genuine semantic exploration. |
| Approach: | They propose a geometric inference framework that uses a spectral orthogonal probe to introduce semantically heterogeneous reasoning signals into the teacher's orthogonale complement of its dominant subspace. |
| Outcome: | The proposed framework improves accuracy and sampling efficiency over baseline methods on logic and code generation benchmarks. |
Copied to clipboard
| Challenge: | Existing tool learning studies focus on general-purpose tool-use capability, but ignore the importance of personalized tool-user preferences. |
| Approach: | They propose a framework to adapt Large Language Models to personalized tool learning task, which is trained through supervised fine-tuning and direct preference optimization. |
| Outcome: | Extensive experiments on PEToolBench show that the proposed framework outperforms existing LLMs in the personalized tool learning task. |
Copied to clipboard
| Challenge: | Existing approaches to pre-training focus on embedding alignment, but they neglect the modeling of bidirectional contexts. |
| Approach: | They propose a framework to learn languageuniversal representations using multi-granularity contrasting framework . they encode semantic equivalents from different languages into similar representations . |
| Outcome: | The proposed framework can achieve significant performance gains in machine translation and cross-lingual language understanding. |
Copied to clipboard
| Challenge: | Existing prompting techniques for Large Multi-Modal Models (LMMs) focus on improving textual reasoning or leveraging tools for image preprocessing, lacking a simple and general visual prompting scheme to promote vision-language coordination. |
| Approach: | They propose a prompting scheme that scaffolds coordinates to promote vision-language coordination in Large Multi-Modal Models (LMMs) they overlay a dot matrix within the image as visual information anchors and leverage multi-dimensional coordinates as textual positional references. |
| Outcome: | Experiments on a wide range of vision-language tasks show the superiority of SCAFFOLD prompting over the textual Chain-of-Thought prompting. |
Copied to clipboard
| Challenge: | Neural text generation is notorious for repetitive loops and tedious outputs. |
| Approach: | They propose a method that penalizes future generation of repetitive content . they construct an anti-LM based on previously generated text . |
| Outcome: | The proposed method outperforms established baselines in terms of generation quality, decoding speed, and universality. |
Copied to clipboard
| Challenge: | Existing speech codecs struggle to balance these objectives at low bitrates . XY-Tokenizer achieves stronger semantic alignment than representative semantic-distillation codec . |
| Approach: | They propose a low-bitrate speech codec that aligns discrete speech representations with text while preserving fine-grained acoustic details for reconstruction. |
| Outcome: | The proposed codec outperforms existing low-bitrate speech codecs in speech understanding and generation tasks. |
Copied to clipboard
| Challenge: | Existing jailbreak attacks primarily utilize scenario camouflage techniques, however their explicit mention of malicious intent will be easily recognized and defended by LLMs. |
| Approach: | They propose an indirect jailbreak attack approach, Puzzler, which can bypass LLM’s defensive strategies and obtain malicious response by implicitly providing LLMs with some clues about the original malicious query. |
| Outcome: | The proposed approach can bypass the LLM’s defensive strategies and obtain malicious response by implicitly providing LLMs with some clues about the original malicious query. |
Copied to clipboard
| Challenge: | Existing work on social intelligence in NLP does not provide a coherent subfield for researchers to analyze and identify research gaps and future directions. |
| Approach: | They build a social AI taxonomy and a data library of 480 NLP datasets to analyze existing datasets and evaluate language models’ performance in different social intelligence aspects. |
| Outcome: | The proposed infrastructure analyzes existing dataset efforts and evaluates language models’ performance in different social intelligence aspects. |
Copied to clipboard
| Challenge: | Emotional Intelligence (EI) is a key concept in the field of human intelligence. |
| Approach: | They propose a method to enhance EI of large language models by naive fine-tuning on EI-related tasks. |
| Outcome: | The proposed method improves EI of two LLM-based assistants without compromising GI. |
Copied to clipboard
| Challenge: | Existing resources for training neural models to finely classify mental-health stigma are limited, relying primarily on social media or synthetic data without theoretical underpinnings. |
| Approach: | They propose to use an expert-annotated corpus of human-chatbot interviews to finely classify mental-health stigma. |
| Outcome: | The proposed corpus can facilitate research on computationally detecting, neutralizing, and counteracting mental-health stigma. |
Copied to clipboard
| Challenge: | Effectively and efficiently handling complex realworld problems has become a key focus across industry and academia. |
| Approach: | They propose a tree-of-code framework that generates nodes through self-supervision and combines prompt and model exploration in a GT-free setting. |
| Outcome: | Experiments on two datasets with ten popular zero-shot LLMs show that Tree-of-Code boosts accuracy by nearly 20% over CodeAct with fewer than 1/4 turns. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly tasked with creative generation, but their ability to portray non-prosocial, antagonistic personas remains largely unexamined. |
| Approach: | They propose a moral alignment benchmark to test the safety of large language models . they find that models struggle with traits directly antithetical to safety principles . |
| Outcome: | The proposed model fails to accurately portray morally ambiguous or villainous characters . the model fails most with traits directly antithetical to safety principles . |
Copied to clipboard
| Challenge: | Large language models (LLMs) store vast amount of knowledge in their parameters, but they still have limitations in the memorization and utilization of certain knowledge. |
| Approach: | They propose a comprehensive definition of the LLM knowledge boundary and introduce a formalized taxonomy categorizing knowledge into four distinct types. |
| Outcome: | The proposed definition of the LLM knowledge boundary and taxonomy categorizes knowledge into four distinct types . aims to offer a comprehensive overview, facilitate access to key issues, and inspire further advancements in LLM research. |
Copied to clipboard
| Challenge: | Existing datasets and benchmarks focus only on patents or cover limited aspects of the IP field, lacking alignment with real-world scenarios. |
| Approach: | They propose a bilingual IP task taxonomy and a large-scale bilingual benchmark to evaluate LLMs in real-world IP practice. |
| Outcome: | The proposed model achieves only 75.8% accuracy, indicating room for improvement . open-source IP and law-oriented models lag behind closed-source general-purpose models . |
Copied to clipboard
| Challenge: | Clinical trials require that patients meet eligibility criteria to ensure safety and effectiveness of studies. |
| Approach: | They propose a dataset that includes the first-of-its-kind eligibility-criteria corpus and queries for criteria-to-sql . they propose 'neuro semantic parser' which can translate eligibility criteria to executable SQL queries . |
| Outcome: | The proposed parser outperforms existing state-of-the-art general-purpose models while highlighting the challenges presented by the new dataset. |
Copied to clipboard
| Challenge: | Group-Relative Policy Optimization (GRPO) has emerged as an efficient paradigm for aligning Large Language Models (LLMs), but its efficacy is confined to domains with verifiable ground truths. |
| Approach: | They propose a meta-cognitive orchestration layer that treats reward scalarization as a dynamic latent policy, leveraging the model’s terminal hidden states as 'a semantic bottleneck' . Across seven benchmarks, MAESTRO consistently outperforms single-reward and static multi-objective baselines while preserving the efficiency advantages of GRPO. |
| Outcome: | The proposed model outperforms single-reward and static multi-objective baselines while preserving efficiency advantages. |
Copied to clipboard
| Challenge: | Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. |
| Approach: | They propose to use data diversity to measure instruction tuning of large language models. |
| Outcome: | The proposed diversity metric outperforms existing methods on simulated and real-world data and shows that it captures diversity variations and achieves a 0.97 correlation with instruction tuning. |
Copied to clipboard
| Challenge: | Existing models focus on a single therapy, but complex cases require flexible strategies among various therapies. |
| Approach: | They propose a multi-session, multi-therapy, and highly realistic benchmark . it is designed to address three key challenges: 1) can we train a highly realistic AI counselor? 2) How to systematically evaluate an AI counselor?" |
| Outcome: | The proposed benchmark is annotated with extensive professional skills and includes over 677 meta-skills and 4577 atomic skills. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have remarkable evaluation and critique capabilities, providing insightful feedback and identifying flaws in various tasks. |
| Approach: | They propose a framework to train critic models using refinement signals to generate feedback loops where critiques guide the model in refining its responses. |
| Outcome: | The proposed framework outperforms traditional methods and open-source models in terms of critique quality and refinement outcomes. |
Copied to clipboard
| Challenge: | Multimodal manga analysis focuses on enhancing manga understanding with visual and textual features. |
| Approach: | They propose a task to enhance manga understanding with visual and textual features by providing a shared semantic space for vision and language understanding. |
| Outcome: | The proposed task provides a shared semantic space for vision and language understanding. |
Copied to clipboard
| Challenge: | Existing methods that adapt LVLMs to egocentric tasks overlook critical agent-environment interactions, limiting their ability to perform egoic reasoning. |
| Approach: | They propose a zero-shot paradigm to enhance egocentric reasoning by simulating human causal reasoning by formalizing ego-centric reasoning using a structural causal model. |
| Outcome: | The proposed method improves egocentric reasoning abilities on six tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are reshaping recommender systems by leveraging extensive world knowledge and semantic reasoning to interpret user intent. |
| Approach: | They propose a single-agent Trajectory-Aligned Recommender to integrate reasoning capabilities into a model by a multi-agend teacher system. |
| Outcome: | The proposed model surpasses its teacher by 8.7% to 39.5% while eliminating iterative latency. |
Copied to clipboard
| Challenge: | Existing audio deepfake detection datasets are outdated and lack generalization capabilities. |
| Approach: | They construct a new cross-domain audio deepfake detection dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. |
| Outcome: | The proposed models achieve 4.1% and 6.5% error rates in the cross-domain ADD dataset generated by five advanced zero-shot TTS models. |
Copied to clipboard
| Challenge: | Existing approaches to large language models fail to meet expectations for code generation tasks . existing approaches are faced with drawbacks of high resource consumption and inadequate handling of multi-API tasks. |
| Approach: | They propose an Efficient multi-Api code GENeration framework that uses private APIs to pre-train LLMs. |
| Outcome: | The proposed framework shows good acceptability and readability on single-GPU tasks compared to fully fine-tuned LLMs with a larger number of parameters. |
Copied to clipboard
| Challenge: | Existing text attack methods are designed for English text, but robust implementation of Chinese text is understudied. |
| Approach: | They propose an adaptive immune-based sound-shape code algorithm for Chinese text attacks . they leverage the Sound-Shape Code to generate natural substitutions . |
| Outcome: | The proposed algorithm produces high-quality Chinese adversarial examples . it can reduce duplication of population and improve search ability . |
Copied to clipboard
| Challenge: | Existing studies focus on automating instruction generation but do not consider other objectives that impact instruction quality. |
| Approach: | They propose an approach that treats instruction generation as an evolutionary multi-objective optimization problem. |
| Outcome: | The proposed approach improves fine-tuning performance and the generation of high-quality instructions. |
Copied to clipboard
| Challenge: | Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy. |
| Approach: | They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions. |
| Outcome: | The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks. |
Copied to clipboard
| Challenge: | Argument structure extraction (ASE) aims to identify the discourse structure of arguments within documents. |
| Approach: | They propose an Efficient Context-aware ASE model that fully exploits contextual information by augmenting modeling capacity and augmenting training data. |
| Outcome: | The proposed model can extract argumentative discourse structure from documents and reduce reliance on specific words or less informative sentences. |
Copied to clipboard
| Challenge: | Existing research has focused on the earlier stages of emergency response . lack of suitable datasets for reliable and compliance-aware decision-oriented modeling and evaluation is limiting current research . |
| Approach: | They propose a first real-world emergency decision-making dataset EDM-Bench . they propose 'rule-enhanced reasoning framework' that integrates external regulatory knowledge with constrained inference mechanisms to improve both decision safety and interpretability. |
| Outcome: | The proposed framework improves decision safety and interpretability by integrating regulatory knowledge with constrained inference mechanisms. |
Copied to clipboard
| Challenge: | X-ray and CT are the gold standard for COVID-19 diagnosis and treatment . however, due to the excessive number of patients, writing reports becomes a heavy burden for radiologists. |
| Approach: | They propose to use X-ray and CT to generate medical reports automatically . they evaluate DeltaNet on a COVID-19 dataset, where it outperforms state-of-the-art approaches . |
| Outcome: | The proposed system outperforms state-of-the-art methods on a COVID-19 dataset. |
Copied to clipboard
| Challenge: | Existing methods rely on unstructured retrieval or coarse abstractions, which lead to temporal conflicts, brittle reasoning, and limited traceability. |
| Approach: | They propose a unified memory framework that consolidates long-term agent experiences into three interconnected components that combine structured knowledge and evidence to construct compact yet information-dense contexts for reasoning. |
| Outcome: | The proposed framework significantly improves multi-hop and temporal reasoning accuracy while reducing input context length by over 95% compared to long-context baselines. |
Copied to clipboard
| Challenge: | Acronyms are abbreviations formed from the initial components of words or phrases . acronyms can be difficult to understand for people who are not familiar with the subject matter . |
| Approach: | They propose a framework to automatically resolve the true meanings of acronyms in a given context . they use the enterprise corpus as input and a high-quality acronym disambiguation system as output . |
| Outcome: | The proposed framework can be deployed to any enterprise to support acronym disambiguation. |
Copied to clipboard
| Challenge: | Document-level Event Extraction (DEE) is a vital task in NLP . current approaches overlook intricate relationships among events and subtle correlations among arguments within a document . |
| Approach: | They propose a document-level event extraction tool that integrates event relationships and argument correlation graphs to model the relationship among events. |
| Outcome: | The proposed network outperforms existing models and large language models in terms of F1-score across two benchmark datasets. |
Copied to clipboard
| Challenge: | Despite substantial progress in safety alignment techniques, aligned large language models can still produce unsafe responses under minor internal perturbations. |
| Approach: | They introduce Activation Steering Attack (ASA) and leverage the Negative Log-Likelihood (NLL) as a diagnostic signal to probe the local sensitivity of safety behaviors in latent space. |
| Outcome: | The proposed method is model-agnostic and supervision-free, enabling a general and reproducible diagnostic metric for analyzing safety robustness. |
Copied to clipboard
| Challenge: | Multiple-Choice Question Answering (MCQA) is a widely used task in the evaluation of large language models (LLMs). |
| Approach: | They propose a tuning-free, causal effect driven debiasing method which intervenes the activations of identified components according to their causal effects. |
| Outcome: | The proposed method alleviates the aforementioned bias and improves the performance of LLMs. |
Copied to clipboard
| Challenge: | Reinforcement learning with human feedback (RLHF) is widely employed to align large language models with user intent. |
| Approach: | They propose to combine rejection sampling and direct preference optimization to improve alignment with user intent by identifying pairs of contrastive samples from human annotator and alternative LLMs. |
| Outcome: | The proposed method outperforms existing methods including RS, PPO, and DPO in a limited resource environment. |
Copied to clipboard
| Challenge: | Existing approaches to conditional question answering on long documents ignore document structure and discourse relations between sentences in document sections. |
| Approach: | They construct a Structure-Discourse Hierarchical Graph and conduct bottom-up information propagation to address this issue. |
| Outcome: | The proposed approach outperforms the existing methods on the conditional question answering on long documents by 3.0 EM score and 2.4 F1 score on answer measuring, and 2.2 EM and 1.9 F1 scores on jointly answer and condition measuring. |
Copied to clipboard
| Challenge: | Existing frameworks for federated multilingual neural machine translation (Fed-MNMT) are limited in language resources. |
| Approach: | They propose a framework that keeps PLMs frozen and only transfers lightweight adapter modules between clients. |
| Outcome: | The proposed framework reduces communication cost by over 98% while achieving similar or even better performance compared to baselines. |
Copied to clipboard
| Challenge: | Existing multimodal Mixture-of-Experts models accurately perceive image content yet fail in subsequent reasoning . Seeing but not thinking phenomenon is a puzzling phenomenon . |
| Approach: | They propose a routing-guided intervention method that enhances domain expert activation. |
| Outcome: | The proposed method achieves consistent improvements on visual reasoning tasks. |
Copied to clipboard
| Challenge: | Existing speech-text pre-training methods are limited to one or two specific tasks, despite their success in speech-language processing tasks. |
| Approach: | They propose a temporal position prediction task to capture the speech-text alignment . they use a textual dialog pre-training task to generalize a response selection task . |
| Outcome: | The proposed model is superior in learning speech-text alignment and multi-turn dialog context. |
Copied to clipboard
| Challenge: | Large language models have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. |
| Approach: | They propose a benchmark for evaluating large language models on a sufficient set of scientific discovery sub-tasks. |
| Outcome: | The proposed framework extracts critical components from papers across 12 disciplines with expert validation confirming its accuracy. |
Copied to clipboard
| Challenge: | Recent studies reveal a security threat to natural language processing models, called the Backdoor Attack. |
| Approach: | They propose to hack a model by modifying one single word embedding vector without sacrificing accuracy on clean samples. |
| Outcome: | The proposed method is more efficient and stealthier on sentiment analysis and sentence-pair classification tasks. |
Copied to clipboard
| Challenge: | Existing methods for training effective PRMs focus on the first incorrect step and all preceding steps, assuming that all subsequent steps are incorrect. |
| Approach: | They propose a data annotation method specifically designed to score the long CoT reasoning process by using an LLM-based judger for annotation. |
| Outcome: | The proposed method improves PRMs' ability to identify effective self-correction behaviors and reasoning based on erroneous steps. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in complex reasoning through long chain-of-thought, yet they struggle with precise computations and algorithmic operations. |
| Approach: | They propose a training-free approach that activates LRMs’ latent tool-use capabilities through artificial hints and a framework that enables models to learn effective tool utilization through diverse hint patterns and rejection-based data synthesis. |
| Outcome: | Experiments show that START significantly improves state-of-the-art LRMs across challenging benchmarks, including competition-level mathematics (AMC23: 95.0%, AIME24: 75.6%) and graduate-level science questions (GPQA: 64.6%). |
Copied to clipboard
| Challenge: | Existing large language models struggle to achieve an accuracy of even 60%, which is the pass mark for Chinese exams. |
| Approach: | They propose to use CMMLU to evaluate Chinese multilingual and Chinese LLMs in a comprehensive benchmark that covers various subjects and settings. |
| Outcome: | The proposed benchmark covers natural sciences, social sciences, engineering, and the humanities and aims to improve on existing models. |
Copied to clipboard
| Challenge: | Existing methods for generating pun sentences with word senses lack large-scale corpus for supervised learning . a pun is a clever and amusing use of a word with two meanings (word senses) |
| Approach: | They propose an adversarial generative network for pun generation with a generator and a discriminator to distinguish between generated pun sentences and real sentences with specific word senses. |
| Outcome: | The proposed network generates sentences that are more ambiguous and diverse in both automatic and human evaluation. |
Copied to clipboard
| Challenge: | Language-model-based agents operating over extended interactions face persistent challenges in preserving temporally grounded information and maintaining behavioral consistency across sessions. |
| Approach: | They propose a general-purpose memory architecture that decomposes agent memory into six components that operate at complementary time scales. |
| Outcome: | BMAM outperforms memory-augmented baselines on LoCoMo benchmarks with 78.45% accuracy . a targeted refinement of the temporal-trigger heuristics raises LongMemEval multi-session accuracy from 45.2% to 56.4%. |
Copied to clipboard
| Challenge: | Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation. |
| Approach: | They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. |
| Outcome: | The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models. |
Copied to clipboard
| Challenge: | Code agents are increasingly trusted to autonomously fix bugs on platforms such as GitHub, yet their security evaluation focuses on functional correctness. |
| Approach: | They propose to attack functionally correct yet vulnerable (FCV) patches by combining multi-turn reasoning with tool invocation and environment interaction. |
| Outcome: | The proposed FCV-Attack achieves an attack success rate of 40.7% on GPT-5 Mini + OpenHands. |
Copied to clipboard
| Challenge: | Existing approaches to generate informative titles for products with limited labels are inadequate for novel products. |
| Approach: | They propose a prompt-based approach to generate attractive titles for novel products . they use multimodal prompts to preserve characteristics and writing styles of novel products. |
| Outcome: | The proposed approach achieves state-of-the-art results on novel product categories with limited labels. |
Copied to clipboard
| Challenge: | Existing approaches to image paragraph captioning ignore the past alignment information, resulting in repetitive captioning and incomplete captioning. |
| Approach: | They propose an Interactive key-value Memory-augmented Attention model for image paragraph captioning to keep track of attention history along with update-chain of decoder state. |
| Outcome: | Extensive experiments on a benchmark dataset demonstrate the effectiveness of the proposed model. |
Copied to clipboard
| Challenge: | Existing attention mechanisms are data-driven, but most are data driven. |
| Approach: | They propose a knowledge-attention encoder which integrates prior knowledge from external lexical resources into deep neural networks for relation extraction task. |
| Outcome: | The proposed system outperforms existing CNN, RNN, and self-attention based models on a large-scale relation extraction dataset. |
Copied to clipboard
| Challenge: | Current evaluation methods for large language models face two key challenges: 1. evaluation validity and 2. Result interpretation reduce the pluralistic and incommensurable values to one-dimensional scores. |
| Approach: | They propose a platform for comprehensive value diagnosis of large language models (LLMs) that provides a generative evaluation paradigm that automatically creates real-world test items co-evolving with ever-advancing LLMs. |
| Outcome: | The proposed platform provides a framework for comprehensive value diagnosis of large language models (LLMs) with fine-grained scores and case studies across 27 value dimensions for 33 leading LLMs, customized comparisons, and visualized analysis of LLM’s alignment with cultural values. |
Copied to clipboard
| Challenge: | Recent studies focus on retrieval to solve knowledge-intensive tasks, but the potential of retrieval for non-knowledge-intensive (NKI) tasks remains under-explored. |
| Approach: | They propose a task-agnostic retrieval framework for NKI tasks that uses a static index and a prompt-guided reranker to re-rank the nearest evidence according to task-specific relevance. |
| Outcome: | The proposed framework outperforms state-of-the-art retrieval-augmented methods on NKI tasks and will be released for further research. |
Copied to clipboard
| Challenge: | Recent advances in large language model reasoning focus on mathematics and coding domains, but scientific reasoning remains limited in other domains due to limited dataset coverage. |
| Approach: | They propose a framework for sustainable scientific reasoning QA generation by synthesizing a new dataset of domain-specific science questions from peer-reviewed literature. |
| Outcome: | The proposed framework and dataset enable scalable and sustainable research in scientific reasoning. |
Copied to clipboard
| Challenge: | Tables are a widely used data format that poses unique challenges for language models due to their structured row-column interactions. |
| Approach: | They propose a region-based reinforcement learning approach that integrates region evidence into reasoning steps. |
| Outcome: | The proposed method outperforms baseline models on three benchmark datasets and significantly reduces the reasoning token consumption by 67.5%. |
Copied to clipboard
| Challenge: | Existing studies suggest that Large Language Models can generate human-like responses, but it is unclear how well they work and where the plausible predictions derive from. |
| Approach: | They propose to use LLMs to generate human-like responses by mutability and accessibility of social inputs to perform a social prediction task. |
| Outcome: | The proposed model performs well in three realistic settings and a novel social prediction task. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting has been shown to be effective in eliciting structured reasoning from large language models (LLMs). |
| Approach: | They propose a data distribution lens to understand when and why CoT reasoning fails . they propose 'data-based' training that trains LLMs from scratch . |
| Outcome: | The proposed model enables models to generate reasoning trajectories that approximate those observed during training. |
Copied to clipboard
| Challenge: | a recent study shows that robots display human-like characteristics in dialogues . this anthropomorphism raises concerns about the accuracy of AI and its capabilities . |
| Approach: | They propose to use a dataset to analyze self-anthropomorphic and non-self-anthropophilic responses in robots . they propose to combine these two types of responses to create a new category of bot responses . |
| Outcome: | The proposed approach preserves the original dialogues from existing corpora and enhances them with paired responses: self-anthropomorphic and non-self-anthropophilic for each original bot response. |
Copied to clipboard
| Challenge: | Existing approaches to longterm memory rely on rigid retrieval granularity, accumulation-heavy maintenance strategies, and coarse-grained update mechanisms. |
| Approach: | They propose a framework that leverages coordinated agents to manage memory across multiple granularities. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks while reducing token consumption by approximately 80%. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) struggle with processing long contexts due to the limited context window. |
| Approach: | They propose to combine a limited context window with a continual learning perspective to improve LLMs' efficiency in processing long contexts. |
| Outcome: | The proposed models improve the performance of Large Language Models (LLMs) by integrating learning strategies with existing approaches. |
Copied to clipboard
| Challenge: | Existing benchmark-based evaluations cannot accurately reflect the performance of real-world applications. |
| Approach: | They propose a reliable strategy for domains to choose more robust LLMs for real-world applications. |
| Outcome: | The proposed strategy addresses the challenges faced by domains to choose more robust LLMs for real-world applications. |
Copied to clipboard
| Challenge: | Anomaly detection (AD) is an important machine learning task with many real-world uses, including fraud detection, medical diagnosis, and industrial monitoring. |
| Approach: | They propose a benchmark that evaluates how large language models (LLMs) can help with NLP anomaly detection. |
| Outcome: | The proposed model can perform zero-shot detection without tasks-specific training, data augmentation and model selection, and it can suggest unsupervised AD models. |
Copied to clipboard
| Challenge: | Existing detectors for Large Language Models (LLMs) struggle to generalize in open-world settings. |
| Approach: | They propose a framework to detect LLM-generated text with exceptional generalization to unseen domains by reinforcing LLMs’ inherent rewriting tendencies. |
| Outcome: | The proposed framework outperforms state-of-the-art detection methods by 23.04% in AUROC, 35.10% for out-of distribution tests, and 48.66% under adversarial attacks. |
Copied to clipboard
| Challenge: | Recent studies have shown that by curating high quality and diverse instruction tuning datasets, we can significantly improve instruction-following capabilities. |
| Approach: | They propose an algorithm to control diversity and quality of instruction tuning datasets and validate it. |
| Outcome: | The proposed algorithm significantly improves worst and average case performance on large scale instruction tuning datasets. |
Copied to clipboard
| Challenge: | Existing approaches to recover dropped pronouns ignore the dependencies between pronounes in neighboring utterances. |
| Approach: | They propose a framework that combines Transformer network and General Conditional Random Fields to model the dependencies between pronouns in neighboring utterances. |
| Outcome: | The proposed framework outperforms state-of-the-art models on three Chinese conversation datasets showing that it captures the dependencies between pronouns in neighboring utterances. |
Copied to clipboard
| Challenge: | Document AI parsing semi-structured image form is a key information extraction task. |
| Approach: | They propose a multimodal and multilingual semi-structured FORM PARSER which integrates SER and relation extraction into a unified framework. |
| Outcome: | The proposed framework achieves up to 1.79% improvement on RE tasks in multilingual and zero-shot settings. |
Copied to clipboard
| Challenge: | Existing unsupervised vision-and-language pre-training methods take pre-extracted region-based visual features from external object detectors, which limits flexibility and reduces computational efficiency. |
| Approach: | They propose an unsupervised vision-and-language pre-training task that predicts which patches contain an object referred to in natural language from the encoded visual features. |
| Outcome: | The proposed approach outperforms existing methods and obtains state-of-the-art results on four vision-and-language tasks. |
Copied to clipboard
| Challenge: | Currently, about half of ancient Chinese glyphs have not been deciphered yet. |
| Approach: | They propose to use a Chinese glyph knowledge graph to infer the Chinese character label for the unknown ancient Chinese . they propose to combine the visual, textual, and the graph data to create a multimodal Chinese morph identification framework. |
| Outcome: | The proposed method can identify ancient Chinese characters from 1300 BC to 200 BC based on image and radical semantics on a 1000-year-old Chinese glyph dataset. |
Copied to clipboard
| Challenge: | Extensive experiments on 9 proprietary LLMs reveal that SE behaviors are widespread . study identifies egoistic decision-making as a risk for large language models . |
| Approach: | They propose a benchmark to measure egoistic behavior in large language models . they propose toxicity, jailbreak vulnerability and a lightweight mitigation that reinforces situational constraints . |
| Outcome: | The proposed model has a 67.96% occurrence rate and frequently manifests as manipulative coercion. |
Copied to clipboard
| Challenge: | Existing work focuses on extracting aspect terms and opinion terms without considering the relations between aspect terms . |
| Approach: | They propose a task to extract aspect terms, opinion terms, and expressed sentiments from a two-dimensional (2D) table. |
| Outcome: | The proposed method achieves state-of-the-art on several public benchmarks and is well-suited to the ASTE task. |
Copied to clipboard
| Challenge: | Existing multimodal reasoning benchmarks for large vision-language models emphasize single-image analysis and fail to exploit contextual information across multiple images. |
| Approach: | They propose a benchmark to evaluate Olympiad-level reasoning when evidence is distributed over multiple images. |
| Outcome: | The proposed model outperforms existing models on bi-image Olympiads and Gemini-3-Pro on multimodal Olympiad-level reasoning tasks. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for long-form speech are limited to limited domains, creating a significant gap with the diverse downstream applications. |
| Approach: | They propose a benchmark that decomposes "long-form speech quality" into specific, disentangled dimensions. |
| Outcome: | The proposed benchmark decomposes “long-form speech quality” into specific, disentangled dimensions. |
Copied to clipboard
| Challenge: | Existing Large Multi-modal Models lack a robust visual processing capability that is often masked by evaluation metrics that prioritize final-answer accuracy. |
| Approach: | They propose a three-layer evaluation framework that scrutinizes the generation of valid visual aids and the soundness of subsequent reasoning steps. |
| Outcome: | The proposed framework examines the generation of valid visual aids and the soundness of subsequent reasoning steps on state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing methods to extract entities and relations from unstructured texts are difficult to handle due to the overlapping triple problem. |
| Approach: | They propose a translation decoding schema for joint extraction of entities and relations from unstructured texts to form factual triples. |
| Outcome: | The proposed model can handle the overlapping triple problem, and is 2 times faster than the state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing models for word-level autocompletion (WLAC) only use human typed sequences as prefixes in decoding module. |
| Approach: | They propose a novel iterative nonautoregressive instruct generation model for WLAC task . it uses human typed sequences and iterating decoding with subwords to fully utilize input information. |
| Outcome: | The proposed model is more competent in dealing with low-frequency words, and achieves state-of-the-art results on the WMT22 and benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for converting unstructured text into structured Knowledge Graphs (KGs) have limitations such as large amount of noise, inaccurate knowledge, and hallucination . |
| Approach: | They propose a GraphJudge framework to reduce noise in real-world documents . they propose Graphjudge to fine-tune a LLM as a graph judge to enhance quality . |
| Outcome: | The proposed framework eliminates noise in real-world documents and improves the quality of generated KGs. |
Copied to clipboard
| Challenge: | Existing memory systems can support long-horizon human-LLM interactions by persisting historical interactions beyond limited context windows. |
| Approach: | They propose a framework that augments memory systems with a self-evolving meta-memory . meta-meso is iteratively distilling transferable knowledge utilization experiences . results show MetaMem outperforms strong baselines by over 3.6% . |
| Outcome: | The proposed framework outperforms baselines by over 3.6% in the long-horizon human-LLM interaction. |
Copied to clipboard
| Challenge: | SIQ quantifies voice understanding abilities and provides unified comparisons between cascaded methods and end-to-end models. |
| Approach: | They propose a human cognition-inspired evaluation pipeline for voice understanding large language models (LLM_Voice) that quantifies voice understanding abilities and provides unified comparisons between cascaded methods and end-to-end models. |
| Outcome: | The proposed framework quantifies voice understanding abilities and provides unified comparisons between cascaded methods and end-to-end models, identifies annotation errors in existing benchmarks, and detects hallucinations in LLM_Voice. |
Copied to clipboard
| Challenge: | Existing static image-text benchmarks are insufficient for evaluating multimodal large language models’ dynamic perception and interactive reasoning abilities. |
| Approach: | They propose a game-based evaluation framework to assess multimodal large language models’ visual reasoning in dynamic, continuous-space environments. |
| Outcome: | The proposed framework systematically assesses MLLMs’ visual reasoning in dynamic, continuous-space environments. |
Copied to clipboard
| Challenge: | Existing methods to augment sentiment models have failed to mitigate spurious association problem inherent in the original data. |
| Approach: | They propose a framework for enhancing sentiment models using an antonymous paradigm and contrastive learning to generate high-quality samples. |
| Outcome: | The proposed framework achieves state-of-the-art performance on four benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for parameter pruning fail to utilize the knowledge from pruned parameters. |
| Approach: | They propose a method that uses manifold learning and the Information Bottleneck measure to merge similar layers to preserve model performance. |
| Outcome: | The proposed method outperforms pruning methods on multiple datasets and LLMs with quantization and achieves substantial compression ratios. |
Copied to clipboard
| Challenge: | Multi-aspect controllable text generation has attracted increasing attention . but the mutual interference of multiple prefixes limits its extensibility to training-time unseen combinations. |
| Approach: | They propose to use trainable gates to normalize the intervention of prefixes to restrain the interference. |
| Outcome: | The proposed approach outperforms baselines on constraint accuracy, text quality, and extensibility. |
Copied to clipboard
| Challenge: | Existing methods for complex instruction-following with elaborate constraints rely on a weaker model, especially GPT-4, limiting their application. |
| Approach: | They propose a Multi-granularity Self-Contrastive Training framework to improve instruction alignment without relying on a stronger model. |
| Outcome: | The proposed framework improves instruction-following with elaborate constraints without external supervision on coarse and fine granularity. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) based large language models are popular for multitasking . however, whether each expert can specialize to a task remains unclear . |
| Approach: | They propose to use a dictionary learning approach to analyze expert collaboration mechanisms in MoE LLMs. |
| Outcome: | The proposed model outperforms existing methods by 2.5% while enabling 50% expert reduction. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) scaling is limited by data quality and domain mixing and instance selection are two separate problems. |
| Approach: | They propose a framework that unifies mixing and selection without training proxy models or relying on external reference datasets. |
| Outcome: | The proposed framework achieves 2.0 data efficiency over a random baseline and further improves overall performance compared to SOTA methods in reasoning-heavy evaluations and multilingual generalization. |
Copied to clipboard
| Challenge: | Existing methods that learn from multiple semantically-equivalent questions are limited to one-to-one mapping . |
| Approach: | They propose a constraint to explore the underlying complementary semantic information among multiple semantically-equivalent questions and learn robust feature representations with reduced spurious associations. |
| Outcome: | The proposed method outperforms strong competitors and achieves state-of-the-art results on five benchmark datasets. |
Copied to clipboard
| Challenge: | Existing research on multi-modal dialogue pre-training is limited due to limited availability of multi-dimensional data . a recent emergence of chatGPT 1 has increased confidence in the potential for this goal . |
| Approach: | They propose a framework for multi-modal dialogue pre-training that integrates experts to accommodate multi-faceted tasks. |
| Outcome: | The proposed framework achieves state-of-the-art on eight multi-modal dialog benchmarks. |
Copied to clipboard
| Challenge: | Existing empathetic dialogue models lack emotion-dependent response generation . elaine mccartney: "i'm sorry to hear that! " |
| Approach: | They propose a model to generate empathetic responses with human-consistent intents . they aim to address the bias of the empathic intent distribution between epd models and humans . |
| Outcome: | The proposed model outperforms state-of-the-art models in terms of empathy, relevance, and diversity on automatic and human evaluation. |
Copied to clipboard
| Challenge: | OpenAI’s o1 model showed this capability but did not publicly share its methodology, leading to many replication efforts. |
| Approach: | They curate a small dataset s1K with 1,000 reasoning questions based on three criteria we validate through ablations: difficulty, diversity, and quality. |
| Outcome: | The proposed model exceeds o1-preview on competition math questions by up to 27% (MATH and AIME24). |
Copied to clipboard
| Challenge: | Existing methods for active learning rely on model uncertainty or disagreement to pick unlabeled data, leading to over-confidence in superficial patterns and lack of exploration. |
| Approach: | They propose to use a bi-directional encoder and a uni-directional decoder to generate and score an explanation for low-resource text classification. |
| Outcome: | The proposed model improves on 9 strong baselines on six datasets and can generate explanations for its predictions. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have evolved from statistical sequence predictors to sophisticated autonomous agents capable of reasoning, planning, and sustaining multi-turn conversa-tions. |
| Approach: | They propose a system that instantiates a "Sentient Agent" that simulates human-like emotional changes and inner thoughts to provide a more realistic evaluation of the model in multi-turn conversations. |
| Outcome: | The proposed framework measures the agent's higher-order social cognition in multi-turn conversations. |
Copied to clipboard
| Challenge: | Recent advances in neural machine translation (NMT) depend on source text to generate translation. |
| Approach: | They propose to use extracted templates from tree structures as soft target templates to guide the translation procedure. |
| Outcome: | The proposed model outperforms baseline models on four benchmarks and demonstrates the effectiveness of soft target templates. |
Copied to clipboard
| Challenge: | Existing approaches to attack Large Language Model (LLM) tool-learning systems are black-box oriented and rely on static commands that cannot adapt flexibly to the changes in user queries and the invocation chain. |
| Approach: | They propose a dynamic attack comment generation approach for information theft attacks in LLM tool-learning systems that mimics the familiar by inferring the information utilized by upstream tools. |
| Outcome: | The proposed approach outperforms baselines with +13.2% ASRTheft and can be generalized to new tool-learning systems to expose their information leakage risks. |
Copied to clipboard
| Challenge: | Existing evaluations of phone recognition systems only measure surface-level transcription accuracy. |
| Approach: | They propose to standardize transcription-based evaluation and assess downstream utility in clinical, educational, and multilingual settings with transcription and representation probes. |
| Outcome: | The proposed system outperforms LALMs in clinical, educational, and multilingual settings. |
Copied to clipboard
| Challenge: | Existing knowledge graph completion frameworks for knowledge graphs are far from complete and require missing triples to be added to them. |
| Approach: | They propose a dynamic pruning technique to obtain a pruned model from a large source model, where the pruning mask of the pruned models could be updated adaptively per epoch after the model weights are updated. |
| Outcome: | The proposed framework achieves competitive performance compared to strong baselines, while being 10x smaller than baselines. |
Copied to clipboard
| Challenge: | Existing methods to learn sentence embeddings with rich semantics are limited due to the discrete nature of natural language. |
| Approach: | They propose to use AI feedback to improve contrastive learning of sentence embeddings by combining human feedback and AI feedback. |
| Outcome: | The proposed method achieves state-of-the-art performance on several semantic textual similarity and transfer learning tasks compared to other unsupervised and supervised contrastive learning methods. |
Copied to clipboard
| Challenge: | Existing benchmarks primarily focus on Python and are limited in terms of language diversity. |
| Approach: | They propose a multilingual debugging benchmark that includes 3.9K test samples of 20 programming languages and introduces the debug instruction corpora MdEval-Instruct by injecting bugs into the correct multilingual queries and solutions. |
| Outcome: | The proposed benchmark includes 3.9K test samples of 20 programming languages and covers the automated program repair task, bug localization task, and bug identification task. |
Copied to clipboard
| Challenge: | Existing PLMs are infeasible for processing long documents due to computational costs and incomprehensive document understanding. |
| Approach: | They propose a retrieval model that models local semantics and global context semantics in a tightly-coupled manner. |
| Outcome: | The proposed model overcomes three core challenges of long document retrieval: substantial computational cost, incomprehensive document understanding, and scarce annotations. |
Copied to clipboard
| Challenge: | Existing methods suffer from key information loss and difficulty in adjusting the length of compressed sequences based on documentation lengths. |
| Approach: | They propose two strategies for compressing tool documentation into concise and precise summary sequences for tool-using language models. |
| Outcome: | The proposed approach achieves comparable performance to the upper-bound baseline under 16x compression ratio. |
Copied to clipboard
| Challenge: | Existing methods for automating program repair face insufficient bug dependency modeling and inadequate global repair planning when addressing semantically complex multi-location bugs. |
| Approach: | They propose a multi-location automatic repair method via cascading planning and generation . they propose to model dependencies among bugs and cluster them to ensure rationality . |
| Outcome: | The proposed method resolves 84 multi-location bugs, achieving a 31% improvement over current methods. |
Copied to clipboard
| Challenge: | Existing studies ignore data imbalance in multilingual settings and do not utilize monolingual data. |
| Approach: | They propose a cross-lingual summarization model that aligns cross-linguistic data with high-resource monolingual data via contrastive and consistency loss. |
| Outcome: | The proposed model outperforms baseline models and consistently dominates on 45 language pairs. |
Copied to clipboard
| Challenge: | Existing work reveals only randomly permuted activations to the client, allowing adversaries to extract model weights. |
| Approach: | They propose an attack that aligns differently shuffled activations to a common permutation and exploits them to extract model weights. |
| Outcome: | The proposed attack can align shuffled activations to a common permutation and exploit them to extract model weights with a query cost of approximately $1. |
Copied to clipboard
| Challenge: | GigaSpeech 2 is a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages. |
| Approach: | They propose a large-scale, multi-domain, multilingual speech recognition corpus for low-resource languages and an automated pipeline for data crawling, transcription, and label refinement. |
| Outcome: | The proposed corpus reduces the word error rate for Thai, Indonesian, and Vietnamese on a realistic YouTube test set by 25% to 40% compared to Whisper large-v3. |
Copied to clipboard
| Challenge: | Existing methods for video editing rely on textual cues from ASR transcripts and segment selection, often neglecting rich visual context. |
| Approach: | They propose a human-inspired automatic video editing framework that leverages multimodal narrative understanding to address these limitations. |
| Outcome: | The proposed framework outperforms existing baselines across general and advertisement-oriented editing tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks. however, their extensive memory requirements present significant challenges for deployment in resource-constrained environments. |
| Approach: | They propose a training-free framework that achieves ultra-low equivalent bit-width KV cache quantization. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on TruthfulQA and LongBench. |
Copied to clipboard
| Challenge: | Existing approaches to interpretable representation learning rely on masks that weight the significance of input features, but the origin of these masks is uncertain. |
| Approach: | They propose a causal framework that directly learns identifiable representations from attention weights rather than relying on importance masks. |
| Outcome: | The proposed framework learns identifiable and explainable representations from attention weights, rather than masks, and guarantees faithfulness on real-world datasets. |
Copied to clipboard
| Challenge: | Large language model (LLM) routing assigns each query to the best suitable model from an ensemble. |
| Approach: | They introduce a large-scale benchmark and unified framework for LLM routing . they find that many routing methods exhibit similar performance under unified evaluation . |
| Outcome: | The proposed benchmark provides comprehensive metrics for both performance-oriented and performance-cost trade-off routing. |
Copied to clipboard
| Challenge: | Experimental results demonstrate that a Pruned interpretable knowledge Graph Learning framework for explainable stance detection is state-of-the-art for social media stance prediction. |
| Approach: | They propose a Pruned interpretable knowledge Graph Learning framework for explainable stance detection that incorporates commonsense knowledge and prunes redundant information to ensure precision and minimize noise. |
| Outcome: | The proposed framework achieves state-of-the-art on three public datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios. |
| Approach: | They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts. |
| Outcome: | The proposed model performs comparable to state-of-the-art large models on the test set. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) offer promising avenues for automated cognitive restructuring in mental health settings, but current approaches lack the adaptability to balance conflicting therapeutic dimensions, such as empathy and rationality. |
| Approach: | They propose a decoupled optimization framework that implements a dimension-guided Monte Carlo tree search to train expert policies specialized for distinct therapeutic attributes rather than relying on a monolithic alignment strategy. |
| Outcome: | The proposed framework achieves consistent improvements over baselines across multiple evaluation dimensions, including diagnostic accuracy, contextual appropriateness, task effectiveness, and overall helpfulness, while enabling controllable cognitive restructuring generation. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are the leading paradigm in document-level relation extraction. |
| Approach: | They propose a cascade framework that leverages the complementary strengths of PLMs and LLMs through a detect-then-rethink paradigm. |
| Outcome: | The proposed framework improves on BioRED and CDR datasets and improves existing models. |
Copied to clipboard
| Challenge: | FlagEvalMM is an evaluation framework designed to assess multimodal models . it is designed to be used for vision-language understanding and generation tasks . |
| Approach: | They propose an evaluation framework that decouples model inference from evaluation through an independent evaluation service. |
| Outcome: | The evaluation framework offers accurate and efficient insights into model strengths and limitations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown remarkable achievements across various language tasks. |
| Approach: | They propose a scientific literature LLM and a knowledge service system based on it . they collect scientific literature and then pre-train it using autoregressive training . |
| Outcome: | The proposed system provides literature investigation, paper reading, and academic writing functions. |
Copied to clipboard
| Challenge: | Recent studies have focused on constructing substantial quantities of IFT data with minimal human effort. |
| Approach: | They propose a multi-agent cooperation framework for the improvement of IFT responses for large language models using a debate-advise-edit-judge paradigm. |
| Outcome: | The proposed framework outperforms baseline models on unseen tasks and shows that it can improve instruction-following capabilities on large language models. |
Copied to clipboard
| Challenge: | Document-level context is crucial for speech translation due to noise from ASR . incorporating document-level contextual information into ST remains a challenge . |
| Approach: | They develop an online framework that integrates document-level context into machine translation . they use document-based modules to integrate document- level context into ST . |
| Outcome: | The proposed framework outperforms baselines in sentence and discourse metrics . it can correct ASR transcription errors and improve translation performance . |
Copied to clipboard
| Challenge: | Existing benchmarks designed to evaluate the reasoning capabilities of large models are limited in scope and lack flexibility to adapt difficulty according to evolving reasoning capacities of models. |
| Approach: | They propose a benchmark that incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models. |
| Outcome: | The proposed benchmark incorporates multidisciplinary questions to evaluate the reasoning capabilities of large models and can adjust and update question difficulty based on the reasoning abilities of advanced models. |
Copied to clipboard
| Challenge: | Existing datasets lack consulting knowledge, resulting in LLMs lacking professional consulting competence. |
| Approach: | They propose a report-based multi-turn dialogue reconstruction framework for Chinese psychological counseling that uses large language models to assist counseling. |
| Outcome: | The proposed framework is open-source and can be used in future research. |
Copied to clipboard
| Challenge: | Existing data on MBTI personality detection are based on self-reported labels and fail to capture the full range of population personality traits. |
| Approach: | They construct a manually annotated MBTI personality detection dataset with soft labels under the guidance of psychologists and use them to identify the task. |
| Outcome: | The MBTIBench is the first manually annotated MBti personality detection dataset with soft labels under the guidance of psychologists. |
Copied to clipboard
| Challenge: | Recent advances in sentiment analysis tend to interference from local factors such as irrelevant words and edges, hindering the precise identification of opinion words. |
| Approach: | They propose a distance-based syntactic weight and Aspect-Fusion Attention to solve this problem. |
| Outcome: | The proposed model outperforms state-of-the-art models on three public datasets and verify its effectiveness. |
Copied to clipboard
| Challenge: | Existing automated generation methods exhibit Weak Applicability and Weak Scalability . existing methods are limited by their reliance on metadata from specific corpora . |
| Approach: | They propose an approach to generate scalable RAG benchmarks using corpus-agnostic methods . they propose a difficulty-guided metric that directs query evolution process . |
| Outcome: | The proposed approach evolves queries significantly more challenging than existing methods . it is able to dynamically increase difficulty, limiting scalability of the query . |
Copied to clipboard
| Challenge: | Existing context window extension methods obstruct scaling external knowledge input. |
| Approach: | They develop a multi-agent framework to overcome two core bottlenecks in existing agent orchestration designs. |
| Outcome: | The proposed framework overcomes two core bottlenecks and improves inference-time knowledge integration without longer-context training. |
Copied to clipboard
| Challenge: | Existing DocRE models which perform well may make more mistakes when merely changing the entity names in the document, hindering the generalization to novel entity names. |
| Approach: | They propose a pipeline to generate entity-renamed documents by replacing the original entity names with names from Wikidata. |
| Outcome: | The proposed pipeline generates entity-renamed documents by replacing the original entity names with names from Wikidata. |
Copied to clipboard
| Challenge: | Existing benchmarks lack the ability to automatically evaluate from users’ perspective and lack the explainability of the results of LLM agents’ code generation capabilities. |
| Approach: | They propose a new benchmark for LLM agents' automated evaluation by simulating user interaction. |
| Outcome: | The proposed benchmark can evaluate the generated projects by user interaction simulation and by code similarity through existing objective indicators. |
Copied to clipboard
| Challenge: | Existing diffusion models have limitations in modeling discrete data, e.g., languages . we present a novel diffusion model for language modeling inspired by linguistic features in languages based on iterative denoising . |
| Approach: | They propose a method that iteratively denoises and adds corruptions to the textual data through soft-masking to better noise it. |
| Outcome: | The proposed model achieves better generation quality and lower training cost than current models with better performance. |
Copied to clipboard
| Challenge: | Existing methods for zero-shot Relation Extraction (RE) lack detailed, context-specific prompts for understanding various sentences and relations. |
| Approach: | They propose a framework that uses a three-stage diversity approach to prompt LLMs by generating multiple synthetic samples that encapsulate specific relations from scratch. |
| Outcome: | The proposed framework outperforms existing LLM-based zero-shot RE methods on benchmark datasets and shows that it produces high-quality synthetic data that enhances performance. |
Copied to clipboard
| Challenge: | Existing studies have evaluated LLMs under noise-free context but the dilemma for LLM to produce inaccurate results under noisy context has not been fully investigated. |
| Approach: | They propose a new method for CoT reasoning using Chain-of-Thought prompting that interacts with LLMs to perform key sentence extraction, variable declaration and answer prediction. |
| Outcome: | The proposed method outperforms existing CoT prompting methods on five reasoning tasks under noisy context. |
Copied to clipboard
| Challenge: | Recent work shows that finetuning pretrained models with contrastive learning makes it possible to learn good sentence embeddings without labeled data. |
| Approach: | They propose an unsupervised contrastive learning framework for learning sentence embeddings . they use a masked language model to mask out the edited sentence . |
| Outcome: | The proposed framework outperforms SimCSE on semantic textual similarity tasks by 2.3 absolute points. |
Copied to clipboard
| Challenge: | Aligned Large Language Models exhibit remarkable versatility, capable of handling diverse real-world tasks. |
| Approach: | They propose a coarse to fine framework to fine-tune aligned Large Language Models to achieve a balance between speciality and versatility. |
| Outcome: | The proposed framework outperforms baseline methods across diverse tasks and model scales. |
Copied to clipboard
| Challenge: | Word sense disambiguation (WSD) methods have not explored word-formations in parataxis languages like Chinese. |
| Approach: | They propose to leverage word-formation knowledge to enhance Chinese WSD by incorporating word-forms into sense disambiguation models. |
| Outcome: | The proposed model improves on baselines in Chinese word sense disambiguation (WSD) with word-formation knowledge, the results show. |
Copied to clipboard
| Challenge: | Pre-trained language models typically lead to high computational cost during inference. |
| Approach: | They propose a slowdown attack framework that can reduce inference efficiency by 80% by leveraging existing adversarial attacks targeting model accuracy. |
| Outcome: | The proposed framework can reduce the efficiency of multi-exit models by 80% on average, validating its effectiveness and generalization ability. |
Copied to clipboard
| Challenge: | Existing evaluation protocols for few-shot natural language understanding (NLU) tasks are inconsistent and hinder fair comparison and measuring progress. |
| Approach: | They propose an evaluation framework that improves previous evaluation procedures in three key aspects, i.e., test performance, dev-test correlation, and stability. |
| Outcome: | The proposed framework improves evaluation procedures in three key aspects, i.e., performance, dev-test correlation, and stability. |
Copied to clipboard
| Challenge: | Existing distillation methods that focus on encoder-only LMs fail to handle the distillation of encoder decoder LM. |
| Approach: | They propose a method that finetunes pretrained language models (LMs) they propose 'MiniEnD' that allows for task-agnostic distillation of LMs. |
| Outcome: | The proposed distillation method is generally effective and competitive compared to other alternatives. |
Copied to clipboard
| Challenge: | a study of large language models (LLMs) shows that they can generate outputs that are honest, positive, harmless, etc. |
| Approach: | They propose a method that amplifies logits difference between positive and negative tokens . they propose to use the logits gap to generate positive and positive tokens after alignment . |
| Outcome: | The proposed method achieves effective alignment, but requires fewer computational resources compared to training-time alignment methods. |
Copied to clipboard
| Challenge: | Existing methods ignore the correlations between labels and different parts of the text can contribute differently for predicting different labels. |
| Approach: | They propose to view the multi-label classification task as a sequence generation problem and apply a decoder-based sequence generation model to solve it. |
| Outcome: | The proposed methods outperform previous work by a substantial margin. |
Copied to clipboard
| Challenge: | Existing work on lifelong learning requires incremental memory space to learn a model . existing work on experience replay or elastic weighted consolidation requires incremental space . |
| Approach: | They propose a framework that leverages a recall optimization mechanism to memorize parameters of previous tasks via regularization and a domain drift estimation algorithm to compensate the drift between different domains in the embedding space. |
| Outcome: | The proposed framework outperforms SOTA models on paraphrase and dialog response generation tasks. |
Copied to clipboard
| Challenge: | Existing methods that ignore contextual knowledge fail to reliably fall back to parametric knowledge when presented with irrelevant context. |
| Approach: | They propose to use contextual knowledge to update and correct LLMs' knowledge by in-context editing instead of retraining. |
| Outcome: | The proposed method outperforms current state-of-the-art methods by a large margin on a dataset that contains irrelevant questions. |
Copied to clipboard
| Challenge: | Despite advances in improving large language model (LLM) to refuse to answer malicious instructions, LLMs remain vulnerable to jailbreak attacks where attackers generate instructions with distributions differing from safety alignment corpora. |
| Approach: | They propose a framework that leverages embedding space distribution analysis to generate jailbreak-like instructions. |
| Outcome: | The proposed framework shows significant decreases in attack success rate on Qwen2.5, Llama3.1, and Llma3.2 without compromising their utility. |
Copied to clipboard
| Challenge: | Existing methods for product attribute value extraction focus on extracting values for a set of known attributes with sufficient training data. |
| Approach: | They propose a prompt tuning approach to extract attributes from product information using mixed prompts. |
| Outcome: | The proposed approach improves on two product benchmarks and shows parameter-efficient training and avoids model overfitting. |
Copied to clipboard
| Challenge: | Existing methods for detecting code capture the overall semantics of the code rather than its intrinsic vulnerability-specific semantics. |
| Approach: | They propose an approach that leverages contrastive learning to generate precise vulnerability code representations under the supervision of vulnerability descriptions. |
| Outcome: | The proposed approach outperforms state-of-the-art methods in vulnerability detection tasks by 11.85% and 13.61%. |
Copied to clipboard
| Challenge: | NoteChat is a cooperative multi-agent framework for generating patient-physician dialogues . evaluator finds it outperforms state-of-the-art models for generating clinical notes . clinical documentation is largely done by physicians at both steps . |
| Approach: | They propose a cooperative multi-agent framework leveraging Large Language Models to generate patient-physician dialogues. |
| Outcome: | The proposed framework outperforms state-of-the-art models for generating clinical notes . it can engage patients directly and help clinical documentation, a leading cause of physician burnout . |
Copied to clipboard
| Challenge: | Existing methods to zero-shot relation classification can only identify seen relations . existing methods rely on descriptive information to improve understandability of relation types . |
| Approach: | They propose a logic-guided semantic representation learning model for zero-shot relation classification that builds connections between seen and unseen relations via implicit and explicit semantic representations with knowledge graph embeddings and logic rules. |
| Outcome: | The proposed model can generalize to unseen relation types and achieve promising improvements. |
Copied to clipboard
| Challenge: | Existing time-aware datasets that focus on persona-grounded conversations focus on temporal dynamics, which narrows their scope and diminishes their complexity. |
| Approach: | They propose a multimodal, time-aware persona dialogue dataset that integrates linguistic, visual, and temporal elements within dialogue and persona memory. |
| Outcome: | The proposed framework integrates linguistic, visual, and temporal elements within dialogue and persona memory to assess a model’s ability to understand implicit temporal cues and dynamic interactions. |
Copied to clipboard
| Challenge: | Existing approaches to integrating graph and language models face two key limitations: achieving robust semantic alignment and ensuring interpretability in outputs. |
| Approach: | They propose a framework to integrate graph and language modalities while enhancing transparency. |
| Outcome: | Extensive experiments on three benchmark datasets show that the proposed framework surpasses existing methods in efficiency and generates outputs that are significantly more interpretable. |
Copied to clipboard
| Challenge: | Existing datasets for semantic parsing are too small in terms of number of programs for training modern data-intensive models. |
| Approach: | They propose a large-scale complex and cross-domain semantic parsing task for a database . they use a dataset with 10,181 questions and 5,693 unique complex SQL queries . |
| Outcome: | The proposed task is different from previous tasks because it uses the same database and program . the best model achieves only 9.7% exact matching accuracy on a database split setting. |
Copied to clipboard
| Challenge: | a new ensemble decoding approach enhances the performance of Large Language Models. |
| Approach: | They propose a multi-prompt ensemble decoding approach to enhance LLM performance . they submit n variations of prompts with X to LLMs in batch mode to decode and derive probability distributions . |
| Outcome: | The proposed method improves pass@k rates, LENS metrics and BLEU scores on diverse NLP tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are too large to be fine-tuned with budget constraints and some are only accessible via APIs. |
| Approach: | They propose a pluggable Reward-Driven Contextual Adapter that integrates large language models as generators and trains them to refine the retrieved information. |
| Outcome: | The proposed method improves ReQA performance on three datasets by up to 20% compared to existing methods. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized various fields, yet their training efficiency is heavily reliant on effective data curation. |
| Approach: | They propose to reuse pre-computed sample-level scores originally generated for data efficiency and introduce two new data ordering methods to improve LLM training. |
| Outcome: | The proposed methods improve the stability and performance of LLM training. |
Copied to clipboard
| Challenge: | Existing methods for modifying large language models focus on individual models, resulting in errors and hallucinations. |
| Approach: | They propose an ensemble-based approach that employs a plug-in model as the editing module and a dynamic weight mechanism to enhance its effectiveness. |
| Outcome: | The proposed approach outperforms existing methods while achieving superior editing efficiency. |
Copied to clipboard
| Challenge: | Existing fraud detection benchmarks focus on single-turn classification tasks, failing to capture dynamic nature of real-world fraud attempts. |
| Approach: | They propose a bilingual benchmark to assess LLMs' ability to resist fraud and phishing attacks across five key fraud categories: Fraudulent Services, Impersonation, Phishing Scams, Fake Job Postings, and Online Relationships. |
| Outcome: | The proposed model improves in role-play settings and in e-commerce and recommendation systems. |
Copied to clipboard
| Challenge: | Existing studies focus on extracting informative or distinct partial-representations from different subspaces, while few studies have paid attention to the aggregation of the extracted partial-Representations. |
| Approach: | They propose to use a routing-by-agreement algorithm to improve multi-head attention by iteratively updating the proportion of how much a part should be assigned to a whole based on agreement between parts and wholes. |
| Outcome: | The proposed algorithm improves the information aggregation for multi-head attention over the standard linear transformation on linguistic probing and machine translation tasks. |
Copied to clipboard
| Challenge: | Extensive research on spoken dialogue systems has advanced the development of intelligent voice assistants, but integration of role information within speech remains an underexplored area. |
| Approach: | They propose a language-based spoken dialogue system that integrates role information within speech to generate contextually appropriate responses. |
| Outcome: | The proposed architecture achieves speaker-specific responses, character understanding, and the generation of targeted replies in multi-party dialogue scenarios, surpassing existing spoken dialogue systems. |
Copied to clipboard
| Challenge: | Recent studies have focused on integrating commonsense knowledge into chatbots to enhance their ability to understand and generate dialogue responses. |
| Approach: | They propose a framework that integrates commonsense knowledge into chatbots to enable them to elicit more empathetic responses. |
| Outcome: | The proposed framework enables LLMs to uncover the implicit requirements of the conversation, aiming to elicit more empathetic responses. |
Copied to clipboard
| Challenge: | Traditional Video Quality Assessment (VQA) focuses on aesthetic fidelity and technical distortions. |
| Approach: | They propose a new task that evaluates whether a UGC item has positive community resonance based on multimodal attributes rather than visual quality alone. |
| Outcome: | The proposed task outperforms state-of-the-art baselines on CASTER-Bench . it provides interpretable and empathetic reasoning paths that align with real community feedback. |
Copied to clipboard
| Challenge: | Existing web agents suffer from limited robustness, efficiency and task success due to lack of structural understanding of websites and lack of browsing priors in pre-trained models. |
| Approach: | They propose an agent-oriented sitemap protocol that integrates structured website knowledge into web agents. |
| Outcome: | The proposed agent-oriented sitemap improves robustness, efficiency and effectiveness without extra training. |
Copied to clipboard
| Challenge: | Existing resources and techniques for named entity recognition (NER) for manufacturing-specific named entities are limited. |
| Approach: | They propose a corpus of Chinese manufacturing specifications, named MS-NERC, with 4,424 sentences and 16,383 entities. |
| Outcome: | The proposed model outperforms neural methods in few-shot and rich-resource domains. |
Copied to clipboard
| Challenge: | Large vision-language models have been widely used but stereotypical biases are unexplored. |
| Approach: | They propose a framework to SCAN stereotypical bias within large vision-language models . they examine stereotype biases with respect to gender and race in three scenarios . |
| Outcome: | The proposed framework can reduce stereotypical biases in large vision-language models . the currently popular models show significant stereotype biase . |
Copied to clipboard
| Challenge: | Existing knowledge graph construction frameworks require predefined schemas, limiting their scalability and domain coverage. |
| Approach: | They propose a framework for fully autonomous knowledge graph construction that eliminates the need for predefined schemas. |
| Outcome: | The proposed framework outperforms state-of-the-art models on multi-hop QA tasks and enhances LLM factuality. |
Copied to clipboard
| Challenge: | Recent intelligent open-domain chatbots have made substantial progress thanks to the rapid development of large-scale pre-training approaches. |
| Approach: | They propose a dynamic flow mechanism to model the context flow and a model to capture the information dynamics across dialogue utterances. |
| Outcome: | The proposed model outperforms the DialoGPT on the dialogue generation task. |
Copied to clipboard
| Challenge: | Sentiment transfer is a text style transfer task that aims to reverse sentiment polarity and reversal in meaning. |
| Approach: | They propose a task called positive reframing that neutralizes a negative point of view and generates 'positive' perspectives without contradicting original meaning. |
| Outcome: | The proposed model neutralizes a negative point of view and generates 'positive' perspectives without contradicting the original meaning. |
Copied to clipboard
| Challenge: | Existing methods to extract relationships from texts depend on memory size and replay these memorized samples in subsequent tasks. |
| Approach: | They propose to use a model to extract relations between entities from texts where the samples of different relations are delivered into the model continuously. |
| Outcome: | The proposed model outperforms the state-of-the-art models and avoids catastrophic forgetting. |
Copied to clipboard
| Challenge: | Contemporary practices in instruction tuning often hinge on enlarging data scaling without a clear strategy for ensuring data quality. |
| Approach: | They propose a method that leverages one-shot learning to discern and select high-quality instruction data from extensive datasets. |
| Outcome: | Nuggets outperforms existing methods on MT-Bench and Alpaca-Eval benchmarks. |
Copied to clipboard
| Challenge: | Mixture of Experts (MoE) models use homogeneous experts with diverse capacities, resulting in a lack of expert specialization and parameter utilization. |
| Approach: | They propose a framework where experts differ in size and possess diverse capacities . they propose HMoE to encourage frequent activation of smaller experts . |
| Outcome: | The proposed framework outperforms homogeneous homogenous MoE models on evaluation benchmarks and achieves lower loss rate with fewer activated parameters. |
Copied to clipboard
| Challenge: | Current large foundational models have demonstrated transformative capabilities, approaching or surpassing human-level performances in many tasks. |
| Approach: | They propose a unified and standardized multimodal benchmark framework with over 50 tasks and more than 10 models to promote transparent and reproducible evaluations. |
| Outcome: | The proposed framework has 50 tasks and more than 10 models to promote transparent and reproducible evaluations. |
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) lack visual-aware tutorial retrieval and historical visual context curation and pruning. |
| Approach: | They propose a framework that integrates an orchestrator and a Reflection-Memory Agent for robust automation. |
| Outcome: | Experimental results show that OS-Symphony delivers substantial performance gains across model scales. |
Copied to clipboard
| Challenge: | Large language models can fix recognition or translation errors that traditional rescoring cannot fix. |
| Approach: | They propose a benchmark for GER that covers both ASR and speech-to-text translation across 15 languages and 28 language pairs. |
| Outcome: | The proposed benchmark is built on common voice 20.0 and CoVoST-2 with Whisper and SeamlessM4T. |
Copied to clipboard
| Challenge: | Existing studies focus on overall performance of machine translation but ignore TS performance, authors say . if TS is applied into post-editing, it will reduce the time and cost of post-production. |
| Approach: | They propose to use a golden corpus annotated by experts to generate a translation suggestion model. |
| Outcome: | The proposed model improves on the golden corpus annotated by translators on four translation directions. |
Copied to clipboard
| Challenge: | Neural text generation models are typically trained by maximizing log-likelihood with the sequence cross entropy (CE) loss. |
| Approach: | They propose an Edit-Invariant Sequence Loss method which computes the matching loss of a target sequence with all n-grams in the generated sequence. |
| Outcome: | The proposed method outperforms the common CE loss and strong baselines on a wide range of tasks. |
Copied to clipboard
| Challenge: | Existing approaches struggle with temporal-spatial challenges in capturing subtle linguistic shifts across different disease stages. |
| Approach: | They propose a large language model-driven T-S fusion framework that integrates multilingual LLMs, contrastive learning and interpretable marker discovery to revolutionize late onset AD detection. |
| Outcome: | The proposed framework achieves state-of-the-art performance in late onset AD detection while enabling cross-linguistic diagnostics. |
Copied to clipboard
| Challenge: | Existing work on entity state tracking or event reasoning is limited to procedural texts. |
| Approach: | They propose a benchmark for causal reasoning of event plausibility and entity states . they represent entities as programming languages while prompting language models . |
| Outcome: | The proposed model outperforms existing models on human reasoning and event reasoning. |
Copied to clipboard
| Challenge: | Existing approaches to train Small Task-specific Models (STMs) using synthetic datasets are limited by the low quality of such datasets. |
| Approach: | They propose a data-generation based zero-shot learning framework that uses multiple PLMs to train small task-specific models. |
| Outcome: | The proposed framework outperforms existing methods in boosting performance across tasks. |
Copied to clipboard
| Challenge: | Existing multi-agent reinforcement learning methods depend on large critic networks to evaluate joint actions, leading to instability and high memory costs. |
| Approach: | They propose a method to optimize large language models for agent-specific roles . they propose combining agent-based frameworks with retrieval-augmented generation . |
| Outcome: | Experiments show that multi-agent group policy optimization outperforms baselines in task performance and computational efficiency. |
Copied to clipboard
| Challenge: | Existing studies have shown that cross-lingual knowledge distillation can improve the performance of pre-trained models for cross-linguistic similarity matching tasks. |
| Approach: | They propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model using contrastive learning, bottleneck, and parameter recurrent strategies. |
| Outcome: | The proposed model can compress the size of XLM-R and MiniLM by more than 50% while the performance is only reduced by about 1%. |
Copied to clipboard
| Challenge: | Experimental results show that representation-based text matching methods suffer from performance degradation due to the lack of interactions between the pair of texts. |
| Approach: | They propose a virtual interaction mechanism that enables deep interaction between texts . they propose 'inteRacTion mechanism' that can be integrated into existing methods as plugins . |
| Outcome: | The proposed method outperforms state-of-the-art models on six text matching benchmarks. |
Copied to clipboard
| Challenge: | Experimental results demonstrate the superiority of our approach to aligning large language models with human preferences. |
| Approach: | They propose a method that evaluates and assigns specific credit to each token using an off-the-shelf reward model. |
| Outcome: | The proposed method evaluates and assigns specific credit to each token using an off-the-shelf reward model. |
Copied to clipboard
| Challenge: | Existing adversarial attacks can cause LLMs to make wrong predictions on downstream tasks or generate harmful content misaligned with human values. |
| Approach: | They propose to use randomized smoothing to add noise to the input and then make predictions based on these denoised versions. |
| Outcome: | The proposed method surpasses existing methods in both empirical and certified robustness in defending against adversarial perturbations for both downstream tasks and human alignments (i.e., jailbreak attacks). |
Copied to clipboard
| Challenge: | Audio-aware large language models (ALLMs) can understand textual and non-textual information in the audio input. |
| Approach: | They use audio-aware large language models (ALLMs) to evaluate the speaking styles of SLMs on two tasks: voice style instruction following and role-playing. |
| Outcome: | The proposed models can understand the textual and non-textual information in the audio input and can be used as a judge to assess the speaking styles of SLMs. |
Copied to clipboard
| Challenge: | Existing evaluations of large language models fail to reflect fine-grained capabilities . existing benchmarks are manually curated or domain-generic, limiting scalability and alignment with real use cases. |
| Approach: | They propose a framework that allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs. |
| Outcome: | The proposed framework reveals fine-grained differences in scientific capabilities that standard benchmarks overlook . it allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific capabilities in LLMs. |
Copied to clipboard
| Challenge: | Existing molecule-language models obscure the hierarchical organization of chemical semantics . Existing models rely on linear or uniform encodings, causing structural distortion . |
| Approach: | They propose a framework that integrates intrinsic molecular topology into large language models. |
| Outcome: | The proposed framework improves on cross-modal retrieval, captioning, and property prediction benchmarks. |
Copied to clipboard
| Challenge: | Existing solutions to fine-tune large language models for domain-specific tasks are ineffective in addressing privacy concerns. |
| Approach: | They propose a privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation. |
| Outcome: | The proposed framework fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation. |
Copied to clipboard
| Challenge: | Existing methods to encourage diversity among multi-head attention are limited. |
| Approach: | They propose a disagreement regularization term to encourage diversity among attention heads . they validated their approach on EnglishGerman and ChineseEnglish translation tasks . |
| Outcome: | The proposed approach improves translation performance across language pairs on English-German and Chinese-English translation tasks. |
Copied to clipboard
| Challenge: | Text summarization is a key natural language generation task, but the high cost of inaccurate summaries raises concerns about the reliability of uncertainty estimation on text summarisation (UE-TS) evaluation methods. |
| Approach: | They propose a UE-TS benchmark that evaluates the uncertainty estimation capabilities of two large language models and one pre-trained language model on three datasets. |
| Outcome: | The proposed benchmark evaluates the uncertainty estimation capabilities of two large language models and one pre-trained language model on three datasets, with human-annotation analysis incorporated where applicable. |
Copied to clipboard
| Challenge: | Existing studies have improved the performance of Large language models on well-defined mathematical benchmarks, but they often overlook ill-defined problems. |
| Approach: | They develop a large-scale benchmark that contains over 5,000 ill-defined mathematical problems. |
| Outcome: | The proposed framework improves the accuracy of identifying unsolvable problems by at least 12% across different LLMs, thus achieving stronger robust mathematical reasoning ability. |
Copied to clipboard
| Challenge: | Experimental results show that consistency preference for lexical chains reduces lexical translation inconsistency . Lexical translation consistency is a common discourse phenomenon . |
| Approach: | They propose a consistency-aware model which captures consistency context . they then define consistency-tailored latent variables which guide translation of corresponding sentences . |
| Outcome: | The proposed model significantly improves translation performance in ChineseEnglish and FrenchEnglish translation tasks. |
Copied to clipboard
| Challenge: | Multimodal in-context learning (ICL) is a key mechanism for harnessing the capabilities of large vision–language models. |
| Approach: | They propose a transformer-based model with task-aware attention that dynamically configures ICL sequences. |
| Outcome: | Experiments on five LVLMs and nine datasets show that TACO surpasses baselines across diverse ICL tasks. |
Copied to clipboard
| Challenge: | Existing generative language models (LMs) can generate new or reusable theorems, but their ability to generate new theorels is under-explored. |
| Approach: | They propose to use Metamath library to generate new theorems that can be saved as reusable knowledge for future theoretical proving. |
| Outcome: | The proposed benchmark evaluates whether an agent can generate valuable (and possibly brand new) theorems that are applicable for downstream theoretic proving as reusable knowledge. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been widely used in planning but lack interpretability and control. |
| Approach: | They propose to augment widely used planning benchmarks with manually annotated, fine-grained, and rich natural language constraints spanning four formally defined categories. |
| Outcome: | The proposed model outperforms existing models in 4 state-of-the-art reasoning LLMs, 4 formal languages, and 4 datasets. |
Copied to clipboard
| Challenge: | Existing methods focus on disentangling speakers and content, while others focus on preserving the source's prosody. |
| Approach: | They propose a rhythm-controllable and efficient zero-shot voice conversion model that transforms the source speaker’s timbre into an unseen one while retaining speech content. |
| Outcome: | The proposed model adapts the linguistic content duration to the desired speaking style, facilitating the transfer of the target speaker’s rhythm. |
Copied to clipboard
| Challenge: | Existing research has proposed external memory modules for Large Language Models (LLMs) to overcome the limitations of finite input length and obtain contextual memory beyond the current input. |
| Approach: | They propose a two-stage alternating co-optimization reinforcement learning method that optimizes evidence retrieval and utilization using semantic feedback and rewards. |
| Outcome: | The proposed method outperforms baselines on lexical overlap and semantic similarity metrics, confirming the co-optimization in memory retrieval and memory utilization. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) enhances the question answering abilities of large language models (LLMs) however, adapting general-purpose RAG systems to specialized fields poses unique challenges due to distribution shifts and limited access to domain-specific data. |
| Approach: | They propose a method that equips large language models with joint capabilities of question answering and question generation for domain adaptation. |
| Outcome: | Experiments on 11 datasets across three different domains verify the efficacy of SimRAG over baselines by 1.2%–8.6%. |
Copied to clipboard
| Challenge: | Chinese idioms are hard to understand by children and non-native speakers due to their non-compositionality and metaphorical meaning. |
| Approach: | They propose a task to rephrase idiom-containing sentences to non-idiomatic ones under the premise of preserving the original sentence’s meaning. |
| Outcome: | The proposed method has better performance than baselines based on the established dataset. |
Copied to clipboard
| Challenge: | Large language models with instruction-following capabilities have revolutionized the field of artificial intelligence. |
| Approach: | They propose an annotation-free framework for empowering large language models with instruction-following capabilities. |
| Outcome: | The proposed framework generates multi-turn multimodal instruction-response conversations from a language model. |
Copied to clipboard
| Challenge: | Extensive experiments on a high-quality, real-world PCQA dataset validate its accuracy and preference. |
| Approach: | They propose a multi-perspective preference alignment for programming-community question answering to generate user-centric responses. |
| Outcome: | Experiments on a high-quality, real-world PCQA dataset validate the proposed model's accuracy and preference. |
Copied to clipboard
| Challenge: | Existing general-domain visual language models lack ability of music notation understanding . Symbolic music is represented in two distinct forms: auditory music and symbolic music . |
| Approach: | They propose to train a multimodal music notation model using a large-scale dataset . they use cross-modal alignment to train the model for music notations analysis . |
| Outcome: | The proposed model improves on music understanding by training with a multimodal music notation model. |
Copied to clipboard
| Challenge: | Existing video benchmarks often resemble image-based questions with scans of only a few key frames, without deep temporal reasoning. |
| Approach: | They propose a video benchmark to assess whether large vision-language models can genuinely think with videos rather than perform superficial frame-level analysis. |
| Outcome: | The proposed benchmark consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. |
Copied to clipboard
| Challenge: | Existing methods for retrieving medical textual knowledge Graphs struggle to perform well, a study finds . existing methods struggle to provide accurate answers to complex questions, he says . |
| Approach: | They synthesize user queries integrating diverse topological structures, relational information, and complex textual descriptions. |
| Outcome: | a new dataset for medical textual knowledge graphs shows that existing methods struggle to perform well . main bottlenecks lie in the scarcity of existing medical TKGs and the limited expressiveness of their topological structures . |
Copied to clipboard
| Challenge: | Existing evaluations of multimodal large language models rely on limited case studies . however, they lack the ability to generate accurate edits according to the instructions . |
| Approach: | They propose a benchmark for chart editing that includes 1,405 edit instructions applied to 233 real-world charts. |
| Outcome: | The proposed benchmark includes 1,405 diverse editing instructions applied to 233 real-world charts. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved remarkable success across various natural language processing tasks, but they still face challenges in performing fundamental NLP tasks, such as syntactic parsing. |
| Approach: | They propose a method that leverages grammar rules from existing treebanks to guide LLMs in correcting previous errors. |
| Outcome: | The proposed method significantly improves performance on in-domain and cross-domain datasets. |
Copied to clipboard
| Challenge: | a recent study has focused on simple settings, but their reliability in complex tasks remains understudied. |
| Approach: | They propose to use large language models as judges to evaluate reliability in complex tasks . they use a challenge benchmark to expose and quantify Auxiliary Information Induced Biases . |
| Outcome: | The proposed benchmark exposes and quantifies Auxiliary Information Induced Biases across 12 basic and 3 advanced scenarios. |
Copied to clipboard
| Challenge: | Long, multi-round, multirole interaction trajectories lead to severe information dilution and context window overload, triggering context collapse which destabilizes reasoning. |
| Approach: | They propose a multi-agent framework that compresses and reorganizes multi-round consensus. |
| Outcome: | The proposed framework outperforms baselines across text-based and multimodal tasks while demonstrating superior diagnostic performance and stability in complex clinical scenarios. |
Copied to clipboard
| Challenge: | Multimodal large language models have demonstrated promising results in a variety of tasks that combine vision and language. |
| Approach: | They propose a benchmark to assess the ability of models to use contextual information in free-form text to enhance visual comprehension. |
| Outcome: | The proposed model fails to extract and utilize contextual information to improve understanding of images. |
Copied to clipboard
| Challenge: | Large-scale pre-trained language models (PLMs) have made extraordinary progress in most NLP tasks, but they fail to achieve state-of-the-art (SOTA) performance. |
| Approach: | They propose a Guassian HMM variant for unsupervised POS tagging that incorporates contexualized word representations into the decoder. |
| Outcome: | The proposed model outperforms state-of-the-art models on Penn Treebank and multilingual Universal Dependencies treebank v2.0. |
Copied to clipboard
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Copied to clipboard
| Challenge: | Existing approaches to automating ML are time-consuming and difficult to understand for human developers. |
| Approach: | They propose a framework that leverages large language models to develop ML solutions for novel tasks. |
| Outcome: | The proposed framework bridges the gap between machine intelligence and human knowledge by exploiting state-of-the-art large language models. |
Copied to clipboard
| Challenge: | Existing methods to translate natural language descriptions into visualization queries focus on spoken languages, not sign languages. |
| Approach: | They propose a sign language interface that enables the DHH community to engage more fully with data analysis. |
| Outcome: | The proposed interface can be used by the deaf and hard-of-hearing community. |
Copied to clipboard
| Challenge: | Membership inference attacks aim to determine whether a specific example was used to train a given language model. |
| Approach: | They propose a membership inference approach that iteratively refines prefix effectiveness and membership scores using an expectation-maximization strategy without requiring labeled non-member examples. |
| Outcome: | The proposed approach outperforms baselines under systematically varied distributional overlap and difficulty. |
Copied to clipboard
| Challenge: | Existing research on learning with noisy labels dates back to the 1980s, but it is still vibrant today. |
| Approach: | They propose a novel DNN model called NetAb to deal with noisy labels during training and train the networks using their respective loss functions in mutual reinforcement. |
| Outcome: | The proposed model can fit training data with noisy labels and predict clean labels. |
Copied to clipboard
| Challenge: | Large language model-based multi-agent systems (MAS) are increasingly used to extend agentic problem solving via role specialization and collaboration. |
| Approach: | They propose a graph-centric framework for orchestrating large language model-based multi-agent systems . they compile a user's natural-language intent into an editable workflow specification and then into an executable graph . |
| Outcome: | The proposed framework compiles natural-language intent into an executable graph and then compile and executes it at runtime. |
Copied to clipboard
| Challenge: | Symbols are used in abstract reasoning, chemical property prediction, and tabular question-answering. |
| Approach: | They propose a method that converts symbols to language-based representations to improve their accuracy. |
| Outcome: | The proposed method improves the accuracy of symbols in language-based models. |
Copied to clipboard
| Challenge: | Recent studies have revealed the vulnerability of pre-trained language models to adversarial attacks. |
| Approach: | They propose a novel approach to repair adversarial examples using an adversarial detector. |
| Outcome: | The proposed approach is effective in various adversarial attack scenarios. |
Copied to clipboard
| Challenge: | Experimental results show that M-TRACE effectively reduces time-anchor drift . external knowledge may be inaccurate while internal knowledge can become outdated . |
| Approach: | They propose a multi-agent reasoning framework for temporal knowledge conflicts . they propose 'TimeConfQA' which guides conflict-aware final reasoning . |
| Outcome: | Experimental results show that M-TRACE reduces time-anchor drift and improves performance on complex temporal question answering tasks. |
Copied to clipboard
| Challenge: | Existing approaches to enhance robustness of deep neural networks focus on perturbation . weak robustness is a problem for many types of adversarial attacks, authors say . |
| Approach: | They propose a lightweight framework for enhancing robustness by perturbing parameters of a model and diversifying adversarial example distributions among different models. |
| Outcome: | The proposed method can improve robustness against adversarial attacks while maintaining accuracy on clean data. |
Copied to clipboard
| Challenge: | Recent work has focused on evaluating large audio models (LAMs) that directly accept audio inputs. |
| Approach: | They propose an interactive approach to evaluate large audio models and collect 7,500 LAM interactions from 484 participants. |
| Outcome: | The proposed model is based on a set of user-generated audio interfaces with 7,500 interactions from 484 participants. |
Copied to clipboard
| Challenge: | Large language models struggle with input errors, often failing to interpret user intent or altering the original question’s structure (over-correction). |
| Approach: | They propose a framework that uses reinforcement learning to address misinterpretation and over-correction by integrating external knowledge with the input. |
| Outcome: | The proposed framework unlocks the full potential of LLMs for the question correction task. |
Copied to clipboard
| Challenge: | Existing methods for multimodal content detection fail to capture cross-modal semantic inconsistencies and ignore inherent noise in multimodal features. |
| Approach: | They propose a multimodal rumor detection method based on a frequency domain spectral selection method and entropy-guided uncertainty fusion method to capture cross-modal semantic inconsistencies. |
| Outcome: | The proposed method outperforms state-of-the-art methods in multimodal rumor detection . it shows stronger detection capability and robustness on multiple datasets . |
Copied to clipboard
| Challenge: | Existing methods for listwise reranking exhibit intrinsic position bias . existing methods are constrained by an inherent trade-off between efficiency and flexibility . |
| Approach: | They propose a training-free framework that mechanically decouples positional bias from ranking decisions. |
| Outcome: | a training-free framework decouples position bias from ranking decisions . evaluations show it outperforms training-based methods and outperformed expensive methods . |
Copied to clipboard
| Challenge: | Recent advances struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world. |
| Approach: | They propose to train different MMT models to support translations between different languages. |
| Outcome: | The proposed model is able to handle the above issues by providing a shared semantic space for multiple languages. |
Copied to clipboard
| Challenge: | Existing approaches to answer selection are limited in domains with limited labeled data. |
| Approach: | They propose a Knowledge-aware Attentive Network framework for cross-domain answer selection that uses the knowledge base as a bridge to enable knowledge transfer from the source domain to the target domain. |
| Outcome: | The proposed model outperforms strong competitors by a noticeable margin in cross-domain answer selection. |
Copied to clipboard
| Challenge: | Existing pre-training frameworks for text-to-SQL parsing have shown inherent differences in distributions between tables and plain text. |
| Approach: | They propose a framework to improve context-dependent Text-to-SQL parsing by leveraging Linking information. |
| Outcome: | The proposed framework achieves state-of-the-art performance on two leading downstream benchmarks. |
Copied to clipboard
| Challenge: | Traditional Function Calling (FC) approaches operate statelessly, requiring multiple exploratory calls to build environmental awareness before execution, leading to inefficiency and limited error recovery. |
| Approach: | They propose a state-based function call approach that maintains explicit system state awareness and implements direct state transitions to achieve target conditions. |
| Outcome: | The proposed approach outperforms traditional function calling approaches, achieving superior execution accuracy and reduced latency. |
Copied to clipboard
| Challenge: | Existing evaluations of LLMs in finance are text-only, monolingual, and largely saturated by current models. |
| Approach: | They propose a multilingual and multimodal benchmark for evaluating LLMs in real financial contexts. |
| Outcome: | The first expert-annotated multilingual and multimodal benchmark is released . it evaluates 21 leading LLMs and shows they perform better in multilingual settings . |
Copied to clipboard
| Challenge: | Autoregressive speech token generation models suffer from hallucinations and undesired vocalizations that do not conform to conditioning inputs. |
| Approach: | They propose an encoder-decoder transformer model that improves contextual adherence of speech token generation LLMs through preference alignment and classifier-free guidance. |
| Outcome: | The proposed model outperforms previous LLM-based models on intelligibility, speaker similarity and naturalness. |
Copied to clipboard
| Challenge: | Existing methods to evaluate consistency capacity of open-domain chatbots are costly and low-efficient. |
| Approach: | They propose an efficient framework for evaluating consistency of open-domain chatbots . they use human judges to interact with chatbot, which is costly and low-efficient . |
| Outcome: | The proposed framework can assess the consistency capacity of chatbots and achieve a high ranking correlation with the human evaluation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) excel in various tasks but remain vulnerable to jailbreak attacks, where adversaries manipulate prompts to generate harmful outputs. |
| Approach: | They propose a Reverse Embedded Defense Attack mechanism that disguises the attack intention as the "defense" intention against harmful content. |
| Outcome: | The proposed method outperforms existing methods on open-source and closed-source models and enables successful jailbreak in one iteration. |
Copied to clipboard
| Challenge: | Existing visual token pruning methods leverage simple metrics derived from human experience, such as attention or similarity, to rank and select tokens within a highly entangled feature space. |
| Approach: | They propose a novel visual token pruning method that uses a concept-driven paradigm to quantify the Marginal Semantic Gain of each token's contribution to uncovered concepts. |
| Outcome: | The proposed method outperforms state-of-the-art methods in a concept-driven model while maintaining semantic completeness. |
Copied to clipboard
| Challenge: | Current parallel corpora are not publicly accessible but trained models are more readily available. |
| Approach: | They propose a method to take advantage of existing translation models to improve one model of interest. |
| Outcome: | The proposed method improves on Chinese-English and German-English datasets and is robust to malicious models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face a fundamental challenge with delayed disambiguation: how can a model update the meaning of an early, ambiguous token when clarifying context only appears later in the sequence? |
| Approach: | They propose a method to defer semantic re-evaluation to subsequent tokens in a process they call "Deferred Semantic Drift" they demonstrate this mechanism in metaphor comprehension and provide causal validation by steering model outputs towards literal or metaphorical meanings via targeted activation interventions. |
| Outcome: | The proposed model can update the meaning of an ambiguous word when clarifying context arrives only after it has been processed. |
Copied to clipboard
| Challenge: | Backdoor attacks are a serious threat to the safety of reusing deep neural networks (DNNs). |
| Approach: | They propose an efficient online defense mechanism based on robustness-aware perturbations to distinguish poisoned and clean samples to defend against backdoor attacks on natural language processing models. |
| Outcome: | The proposed method achieves better defending performance and lower computational costs than existing defense methods. |
Copied to clipboard
| Challenge: | Existing paradigms rely on unreliable prompting or rigid constrained decoding strategies to achieve aesthetic unity. |
| Approach: | They propose a framework to embed external constraints into the model’s intrinsic intuition and use it to generate open-ended creative texts. |
| Outcome: | The proposed framework surpasses baselines in both strict constraint adherence and literary aesthetics. |
Copied to clipboard
| Challenge: | Recent years, neural machine translation systems are often developed with large-scale parallel data extracted from the Web. |
| Approach: | They propose a memory-augmented adapter to steer pretrained neural machine translation models in a pluggable manner by combining model representations and retrieved results. |
| Outcome: | The proposed method outperforms several representative pluggable baselines on style- and domain-specific experiments. |
Copied to clipboard
| Challenge: | Existing models employ a fixed gating network where each token is computed by the same number of experts. |
| Approach: | They propose a flexible training strategy that allows tokens to be processed by a variable number of experts based on expert probability distribution. |
| Outcome: | The proposed model reduces training time and inference quality while maintaining sparsity while maintaining inference accuracy. |
Copied to clipboard
| Challenge: | Existing methods to protect privacy of sensitive data are differential privacy (DP) and DP is used to protect users from privacy leakage. |
| Approach: | They propose an LDP-based Dynamic Text sanitization for privacy-preserving LLM inference that dynamically constructs semantic-aware adjacency lists of sensitive tokens to sample non-sensitive tokens for perturbation. |
| Outcome: | The proposed model excels on three datasets. |
Copied to clipboard
| Challenge: | Existing studies on word class flexibility have been fraught with difficulties in quantifying it accurately and at scale. |
| Approach: | They propose a method to quantify word class flexibility in 37 languages using contextualized word embeddings. |
| Outcome: | The proposed method builds on recent work in contextualized word embeddings to quantify semantic shift between word classes and uncovers shared tendencies in class flexibility across languages. |
Copied to clipboard
| Challenge: | Existing models for natural language processing (NLP) have been challenging to scale attention to longer inputs. |
| Approach: | They propose an extended Transformer construction architecture that scales attention to longer inputs by combining global-local attention with relative position encodings and a "Contrastive Predictive Coding" objective. |
| Outcome: | The proposed architecture scales attention to longer inputs and encodes structured inputs. |
Copied to clipboard
| Challenge: | Modern language models rely on Reinforcement Learning from Human Feedback (RLHF) to encourage safe behaviors, but they remain vulnerable to adversarial attacks due to three key limitations: (1) the inefficiency and high cost of human annotation; (2) the vast diversity of potential adversarials; and (3) the risk of feedback bias and reward hacking. |
| Approach: | They propose an iterative adversarial training method that incorporates three key innovations to address these challenges. |
| Outcome: | Experiments on Mistral-7B-Instruct-v0.3 show that the proposed method significantly enhances robustness and reduces harmful outputs from 5.88% to 0.43%. |
Copied to clipboard
| Challenge: | Pronouns are often dropped in conversational genres as their referents can be easily understood from context. |
| Approach: | They propose an end-to-end neural network model to recover dropped pronouns in conversational data. |
| Outcome: | The proposed model improves on three different conversational genres. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) are powerful but still suffer from inefficient and off-target reasoning. |
| Approach: | They propose a training-free framework that automatically optimizes Large Reasoning Models' reasoning by generating think-prefixes that evolve driven by a taxonomy of reasoning behaviors. |
| Outcome: | The proposed framework significantly improves accuracy-length trade-off for efficient reasoning, drastically improves safety and improves instruction following. |
Copied to clipboard
| Challenge: | Existing LLMs are primarily used for simple text-related tasks, but LLM-based agents can undertake more complex tasks that require planning and interaction with the physical world and humans. |
| Approach: | They propose an Agent-Constitution-based agent framework with a particular focus on improving the LLM-based agents' safety. |
| Outcome: | The proposed framework can enhance an LLM agent’s safety across multiple domains by identifying and mitigating potential dangers during the planning process. |
Copied to clipboard
| Challenge: | Existing Class Activation Mapping methods are ill-suited for interpreting non-autoregressive behaviors of diffusion-based architectures. |
| Approach: | They propose to use a method to generate parallel activation maps by probing intermediate representations in the transformer backbone to capture latent features and their class-specific gradients. |
| Outcome: | Experiments show that Diffusion-CAM significantly outperforms SoTA methods in localization accuracy and visual fidelity. |
Copied to clipboard
| Challenge: | A critical bottleneck is the lack of ground-truth human data to link personality traits to emotional shifts. |
| Approach: | They propose a large-scale dataset to capture reader-based emotional variations across news, social media, and life narratives. |
| Outcome: | The proposed model captures reader-based emotional variations across news, social media, and life narratives. |
Copied to clipboard
| Challenge: | Existing approaches to optimize tool-use policies are monolithic and prone to entangling behaviors. |
| Approach: | They propose a framework that decomposes agent’stool-use policy into four modules and improves them via three mechanisms. |
| Outcome: | The proposed framework outperforms strong baselines on bothGPT-4.1 and Qwen3-8B while maintaining superior efficiency and transferability. |
Copied to clipboard
| Challenge: | HITL-ML approaches are too low-level and far-removed from human’s conceptual models. |
| Approach: | They propose a prototype HITL-ML system that exposes the machine-learned model through high-level, explainable linguistic expressions formed of predicates representing semantic structure of text. |
| Outcome: | The proposed system exposes the machine-learned model through high-level, explainable linguistic expressions formed of predicates representing semantic structure of text. |
Copied to clipboard
| Challenge: | Existing approaches to domain adaptation only use reliable pseudo instances, i.e., pseudo instances with high prediction confidence, to retrain the model. |
| Approach: | They propose a domain adversarial learning enhanced self-training framework that uses meta-learning to estimate the importance of each pseudo instance and a meta constructor to construct the meta-validation set. |
| Outcome: | The proposed framework reduces label noise and preserves hard examples while maintaining accuracy. |
Copied to clipboard
| Challenge: | Existing attempts to unify large language models are limited to open-domain QA with fixed retrieval settings. |
| Approach: | They propose a general reinforcement learning framework that dynamically coordinates retrieval and reasoning. |
| Outcome: | The proposed framework outperforms existing paradigms on open-domain QA, MMLU-Pro, medical, and mathematical reasoning tasks. |
Copied to clipboard
| Challenge: | Existing approaches to generate captions using image captioning are based on multi-head attention (MHA) |
| Approach: | They propose to transform scene graphs into more descriptive captions by using multi-head attention to build a Graph Neural Network (GNN) . they construct a Mixture-of-Expert (MOE)-based decoder where each expert is built on MHA for discriminating the graph embeddings to generate different kinds of words. |
| Outcome: | The proposed framework can generate captions from multiple visual features and objects . it is based on a mixture-of-expert (MOE)-based decoder based upon MHA . |
Copied to clipboard
| Challenge: | Alignment of large language models (LLM) is a process that ensures the model’s responses to user prompts align with human intentions and social values. |
| Approach: | They propose an alignment method based on a two-agent game consisting of an adversarial agent and a defensive agent. |
| Outcome: | The proposed method improves on a two-agent game with an adversarial agent and a defensive agent. |
Copied to clipboard
| Challenge: | Existing embedding-based methods rely on triples in the KG, which is vulnerable to specious relation patterns and long-tail entities. |
| Approach: | They propose a context-enriched framework for KGC that uses a large language model to generate potential answers for each query triple. |
| Outcome: | The proposed framework improves on FB15k237 and WN18RR datasets. |
Copied to clipboard
| Challenge: | Distant supervision uses triple facts to label corpus for relation extraction, leading to wrong labeling and long-tail problems. |
| Approach: | They propose a model to enrich distantly-supervised sentences with entity types by injecting context-free and -related backgrounds into sentences to alleviate sentence-level wrong labeling. |
| Outcome: | The proposed model achieves state-of-the-art on benchmarks and in overall and long-tail performance. |
Copied to clipboard
| Challenge: | Existing work on variableal autoencoders and waterstein autoencoding models has shown significant progress in open-domain response generation. |
| Approach: | They propose to embed user-level and utterance-level information into two multimodal distributions and combine them into a mixed distribution. |
| Outcome: | The proposed model outperforms state-of-the-art models on a large-scale real-world dataset. |
Copied to clipboard
| Challenge: | Existing models for classical Chinese poetry generation only allow users to use keywords to interfere with the meaning of generated poems. |
| Approach: | They propose a model to generate classical Chinese poems from vernacular . their model uses unsupervised machine translation to generate Chinese poems . human evaluation shows it can generate high-quality poems comparable to amateur poems - authors . |
| Outcome: | The proposed model improves the perplexity and BLEU of the proposed model compared with typical models and human evaluation shows it generates high-quality poems comparable to amateur poems. |
Copied to clipboard
| Challenge: | Misaligned large language models can magnify harm by exploiting them to undermine safety . et al., 2022b; Bai e.t., 2023): misalignment, realignment and model-specific resistance are important . |
| Approach: | They evaluate four methods to identify a mechanism asymmetry between attack and defense . they find that ORPO is most effective for misalignment, but DPO excels in realignment . |
| Outcome: | The proposed methods show a mechanism asymmetry between attack and defense . the proposed methods excel in realignment, but at the expense of model utility . |
Copied to clipboard
| Challenge: | Cloud-hosted Large Language Models (LLMs) offer unmatched reasoning capabilities and dynamic knowledge, yet submitting raw queries to these services can expose sensitive user intent. |
| Approach: | They propose a framework that formulates the trade-off between knowledge utility and privacy as a strategic game. |
| Outcome: | The proposed framework reduces intent leakage while maintaining high-fidelity answer quality. |
Copied to clipboard
| Challenge: | Global scientific publications are growing annually by about 4%-5% (Pinedo et al., 2024). |
| Approach: | They introduce an AI-assisted platform that answers diverse questions from researchers using Retrieval-Augmented Generation (RAG) they develop various tools to understand queries, search from the scientific literature, filter retrieved information, provide accurate and comprehensive answers, and self-refine answers. |
| Outcome: | OpenResearcher is built on Retrieval-Augmented Generation (RAG) to integrate Large Language Models (LLMs) with up-to-date, domain-specific knowledge. |
Copied to clipboard
| Challenge: | Existing methods for learning representations of structured knowledge are limited to the minority of people with technical skills. |
| Approach: | They propose a large pretraining dataset and strategy for learning representations of text, tables, and SQL code that leverages the entire context of the problem. |
| Outcome: | The proposed model improves on two SQL tasks and shows a 1.7 and 2.2 percentage point improvement over existing methods. |
Copied to clipboard
| Challenge: | Existing inference optimizations for coarse-grained Mixture-of-Experts models implicitly assume a fixed activation budget, which is poorly understood. |
| Approach: | They propose a training-free policy that adapts token-level activation using router confidence and entropy while remaining within the model’s original budget. |
| Outcome: | The proposed skipping policy can provide substantial throughput gains, but optimal static schedules vary significantly across models and routing mechanisms. |
Copied to clipboard
| Challenge: | Existing multimodal benchmarks overlook linguistic and visual ambiguities, authors say . ambiguity resolution between modalities is lacking in multimodal large language models . |
| Approach: | They propose a benchmark to evaluate multimodal ambiguity resolution across multilingual and cross-modal scenarios. |
| Outcome: | a new benchmark evaluates multimodal ambiguity resolution across multilingual and cross-modal scenarios . the benchmark shows that MLLMs can resolve ambiguities in image-text alignment . however, existing benchmarks often overlook linguistic and visual ambiguties . |
Copied to clipboard
| Challenge: | Existing studies focus on evaluating large language models in close-ended QA tasks, but many clinical decisions involve answering open-ended questions without pre-set options. |
| Approach: | They construct a benchmark to better understand large language models in the clinic . they use existing datasets to evaluate LLMs in clinical situations . |
| Outcome: | The proposed model outperforms human experts in multiple medical tasks. |
Copied to clipboard
| Challenge: | Existing data selection methods suffer from severe domain specificity . existing methods for general instruction-following fail on reasoning tasks . |
| Approach: | They propose a framework that operationalizes contrastive entropy as a domain-adaptive selection criterion through warmup calibration, bi-directional NLL filtering, and entropic-based ranking. |
| Outcome: | Experiments show that InstructDiff outperforms baseline training on reasoning tasks while using only 10% of the data. |
Copied to clipboard
| Challenge: | Existing models for multilingual biomedical training are monolingual, resulting in limited cross-lingual capability. |
| Approach: | They propose a model that transforms a multilingual biomedical corpus into a biomedically domain using a knowledge-anchored approach. |
| Outcome: | The proposed model outperforms monolingual and multilingual models in cross-lingual scenarios. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on single task, simple evaluation metrics, and readily available ground truth (GT) DataSciBench is built on curated, natural, and challenging prompts with complex evaluation criteria and uncertain GT. |
| Approach: | They propose a benchmark for evaluating Large Language Models in data science that integrates LLM-based self-consistency and human verification to ensure accuracy. |
| Outcome: | The proposed framework outperforms open-source models in all metrics and offers rigorous insights into LLM strengths and weaknesses. |
Copied to clipboard
| Challenge: | Existing methods to evaluate factual consistency of text depend on limited data . e.g., generated text can contain factual inconsistencies that are irrelevant to context . |
| Approach: | They propose a new holistic metric that measures factual inconsistencies . they use 4.7M training examples from 7 well-established tasks . |
| Outcome: | The proposed metric outperforms existing metrics on 22 datasets and matches or outperFORMs them. |
Copied to clipboard
| Challenge: | Existing approaches to handle wrong labeling and long-tail relations are labor-intensive and scarce training data. |
| Approach: | They propose a neural network to handle wrong labeling and long-tail relations by collaborating relation-augmented attention. |
| Outcome: | The proposed neural network improves the state-of-the-art on the NYT dataset . |
Copied to clipboard
| Challenge: | Multi-hop question answering (QA) is a central challenge in natural language processing . early mistakes can cause errors and undermine the final result, authors say . |
| Approach: | They propose a reversible multi-agent reasoning framework that backtracks to earlier valid states when conflicts arise. |
| Outcome: | Empirical evaluation shows that the framework improves on forward-only benchmarks by 6% . the approach enables agents to backtrack to valid states when conflicts arise . |
Copied to clipboard
| Challenge: | a new cross-domain entity linking problem exists between FDA-approved medical devices and patents . a recent study compared the semantics of FDA documents with patents, resulting in minimal overlap . |
| Approach: | They propose a framework that links FDA-approved medical devices to their patents . they propose 'MedDevKG' framework that integrates a domain-specific ontology . |
| Outcome: | The proposed framework outperforms existing methods in lower-bound recalls and noise reductions. |
Copied to clipboard
| Challenge: | Existing studies have demonstrated that direct preference optimization (DPO) can be effective in generalizing large language models, but its effectiveness in video domain remains limited. |
| Approach: | They propose a framework that utilizes detailed video captions as a proxy of video content to enable language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. |
| Outcome: | The proposed framework shows that it can be used to align language models with video content and improves performance on open-ended video QA tasks. |
Copied to clipboard
| Challenge: | Existing studies show that standard splits produce low reproducible and unreliable conclusions . reproducibility of empirical experimental conclusions is a problem in NLP domain . |
| Approach: | They propose to transform the reproducibility of a model comparison into a probabilistic function . they propose to use a regularized corpus splitting strategy to estimate the model's performance . |
| Outcome: | The proposed estimator achieves a high SNR and significantly increases reproducibility. |
Copied to clipboard
| Challenge: | Concatenating large language models are adapted to context-aware neural machine translation in a concatenated way . a recent paradigm shift has been witnessed in discourse-related challenges such as zero pronoun translation . |
| Approach: | They propose an alternative adaptation approach to make large language models discriminately model and utilize inter- and intra-sentence contexts. |
| Outcome: | The proposed approach outperforms concatenation mode and improves performance in discourse modeling. |
Copied to clipboard
| Challenge: | Existing methods for table-to-text generation use encoder-decoder framework, but lack of large parallel data is a problem for many domains. |
| Approach: | They propose a model to separate table-to-text generation into two stages: key fact prediction and surface realization. |
| Outcome: | The proposed model achieves 27.34 BLEU score with only 1,000 parallel data, while the baseline model only achieves 9.71 BLUE score. |
Copied to clipboard
| Challenge: | Existing methods for generating test cases with limited training data are not reliable and may be counterproductive. |
| Approach: | They propose a method that splits code snippets into smaller, granular blocks, creating more diverse DPO pairs from the same test cases. |
| Outcome: | The proposed approach shows significant improvements in code generation tasks on benchmark datasets such as HumanEval (+), MBPP (+), and APPS. |
Copied to clipboard
| Challenge: | Existing discourse segmenters rely on complicated hand-crafted features and are not practical in actual use. |
| Approach: | They propose an end-to-end neural segmenter based on BiLSTM-CRF framework that can segment texts fast and accurately using a large corpus. |
| Outcome: | The proposed model is significantly faster than previous methods while achieving state-of-the-art performance on the RST-DT corpus. |
Copied to clipboard
| Challenge: | despite the rapid development of Large Language Models, there is no dedicated benchmark for evaluating LLMs in Chinese K-12 education. |
| Approach: | They propose to develop a benchmark specifically tailored for Chinese K-12 education. |
| Outcome: | EVAL is the first evaluation benchmark specifically tailored for Chinese K-12 education. |
Copied to clipboard
| Challenge: | Language model pretraining is the cornerstone of universal language models (LMs), creating generalpurpose representations to excel across a variety of downstream tasks. |
| Approach: | They propose to use multi-granular tokens to sample large-scale language models for domain-specific use cases. |
| Outcome: | The proposed model outperforms random sampled samples on eight benchmarks with 1% of the data and performs on par with the full RefinedWeb data. |
Copied to clipboard
| Challenge: | Existing methods to identify semantic relations between entities are time-consuming and labor-intensive. |
| Approach: | They propose a relation-aware prototype learning method for document-level relation extraction (FSDLRE) they propose RAPL, which judiciously leverages relation descriptions and real NOTA instances as guidance . |
| Outcome: | The proposed method outperforms state-of-the-art approaches by 2.61% F1 . it generates task-specific NOTA prototypes and refines relation prototypes . |
Copied to clipboard
| Challenge: | Existing studies use LoRA to fine-tune existing LLMs, but this is limited by the data and training gap between them and embedding models. |
| Approach: | They propose a new 1.4B-parameter LLM trained from scratch and fine-tuned as a text embedder that integrates embeddings across different languages. |
| Outcome: | The proposed model improves performance on the Massive Text Embedding Benchmark (MTEB) and Chinese MTEB (May 19, 2025). |
Copied to clipboard
| Challenge: | Existing methods for image-goal navigation fail to extract informative visual cues, leading agents to wander around. |
| Approach: | They propose a framework that decomposes image-goal navigation into high-level planning and low-level execution. |
| Outcome: | The proposed method is superior to existing methods in both simulation and real-world environments. |
Copied to clipboard
| Challenge: | Existing continual learning methods use data replay, parameter isolation and regularization to mitigate catastrophic forgetting. |
| Approach: | They propose a parameter-efficient continual learning framework that updates parameters offline and then trains using an online regularization method. |
| Outcome: | The proposed framework reduces catastrophic forgetting and saves the model with the changed parameters instead of all parameters. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have witnessed remarkable advancements in recent years, prompting the exploration of tool learning. |
| Approach: | They propose a virtual API server and stable evaluation system to assess the stability of large-scale real-time APIs. |
| Outcome: | The proposed benchmarks demonstrate the stability of the proposed system and its caching system. |
Copied to clipboard
| Challenge: | Empirical studies on three different benchmark conversation datasets demonstrate the effectiveness of the proposed model over several strong baselines. |
| Approach: | They propose an addressee-aware module to automatically learn whether the participant keeps the historical emotional state or is affected by others in the next upcoming turn. |
| Outcome: | The proposed model can predict the participant's emotion in the next upcoming turn without knowing the participant’s response yet. |
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models (MLLMs) focus on visual abilities, but audio is essential for video understanding. |
| Approach: | They propose an audio-centric video understanding benchmark to evaluate video comprehension capabilities of multimodal LLMs with a particular focus on auditory information. |
| Outcome: | The proposed video understanding benchmarks evaluate video comprehension capabilities of multimodal models with a particular focus on auditory information. |
Copied to clipboard
| Challenge: | Existing methods for MLLMs struggle with fine-grained temporal reasoning . despite advances in video understanding, current methods struggle with time-sensitive tasks . |
| Approach: | They propose a time-stamp-aware multi-segment grounding method that enhances temporal understanding by introducing timestamps. |
| Outcome: | The proposed method outperforms existing methods on time-sensitive tasks and generalizes well across diverse temporal understanding scenarios. |
Copied to clipboard
| Challenge: | FineState-Bench evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. |
| Approach: | They propose a benchmark that evaluates whether an agent can correctly ground an instruction to the intended UI control and reach the exact target state. |
| Outcome: | The proposed benchmark evaluates whether an agent can ground an instruction to the intended UI control and reach the exact target state. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are evolving rapidly on code generation tasks. |
| Approach: | They propose to automate the vulnerability code benchmark creation with iterative auto validation. |
| Outcome: | The proposed benchmark covers 232 CWE categories across C/C++, Java, and Python languages. |
Copied to clipboard
| Challenge: | Spreadsheets are among the most widely used data formats in real-world applications . existing large language models treat tables as plain text, overlooking layout cues and visual semantics. |
| Approach: | They propose a two-stage multi-agent framework for spreadsheet understanding that adopts a step-by-step reading and reasoning paradigm. |
| Outcome: | Extensive experiments on two spreadsheet datasets show the proposed framework outperforms existing methods on Spreadsheet Bench. |
Copied to clipboard
| Challenge: | Existing experiential memory approaches rely on task-level memory, but this lacks the situational alignment required for complex multi-step decision-making. |
| Approach: | They propose a new fine-grained memory paradigm that aligns memory retrieval with the current state instead of storing and reusing globally shared experiences. |
| Outcome: | Experiments on complex decision-making benchmarks show that the proposed state-aware memory outperforms existing experiential memory approaches and significantly improves task-solving efficiency. |
Copied to clipboard
| Challenge: | Existing methods for controlling the generation of pre-trained language models infuse domain bias into the generation process, making it difficult to generate out-of-domain texts. |
| Approach: | They propose a retrieval-augmented generation framework that uses retrieval to generate fluent sentences with high attribute relevance. |
| Outcome: | The proposed method can generate fluent sentences with high attribute relevance while keeping domain bias out of the model. |
Copied to clipboard
| Challenge: | relying on large language models for information has raised concerns about reliability and accuracy of outputs. |
| Approach: | They propose a hallucination taxonomy with 11 categories for various NLG tasks and propose HAllucination Detection models which integrate hallucinism detection, span-level identification, and correction into a single inference process. |
| Outcome: | The proposed models outperform baselines on HaluEval, FactCHD, and FaithBench, confirming their robustness and versatility. |
Copied to clipboard
| Challenge: | Existing methods for dialogue state tracking ignore the slot imbalance problem and treat all slots indiscriminately, which limits the learning of hard slots. |
| Approach: | They propose to employ a contextual hierarchical attention network to enhance the DST by learning contextual representations. |
| Outcome: | The proposed approach achieves 52.68% and 58.55% joint accuracy on multiWOZ 2.0 and MultiWOZ 2.1 datasets and significantly improves performance (+1.24% and +5.98%) |
Copied to clipboard
| Challenge: | Existing approaches to classify human affect and subjective information from multiple data sources are limited by the lack of high-level feature associations. |
| Approach: | They propose a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data. |
| Outcome: | The proposed model outperforms state-of-the-art approaches on published datasets and visualizes and interprets synchronized attention over modalities. |
Copied to clipboard
| Challenge: | Existing open domain response generation models are limited to paired data, but are less explored in real-world applications. |
| Approach: | They propose to train a neural response generation model with unpaired data and paired data as prior. |
| Outcome: | The proposed model outperforms state-of-the-art models in both automatic and human evaluation when only a few pairs are available. |
Copied to clipboard
| Challenge: | In 2015 alone, about 100 manuscripts describing randomized controlled trials for medical interventions were published every day. |
| Approach: | They propose a corpus of 5,000 medical articles annotated with demarcations of text spans that describe the Patient population enrolled, the Interventions studied and to what they were Compared, and the Outcomes measured. |
| Outcome: | The proposed corpus includes 5,000 medical articles describing clinical randomized controlled trials. |
Copied to clipboard
| Challenge: | Existing benchmarks focused on simplified or isolated aspects of coding, ignoring the full spectrum of programming challenges. |
| Approach: | They propose a case study that examines the performance of large language models across the entire software development lifecycle with four programming languages, multiple domains, and carefully designed and verified metrics for each task. |
| Outcome: | The proposed model performs across the entire software development lifecycle, including design, environment setup, implementation, acceptance testing, and unit testing. |
Copied to clipboard
| Challenge: | Existing methods for training effective AI agents often resort to synthetic data generation. |
| Approach: | They propose a plug-and-play framework for data quality control in tool-use scenarios . they construct a tool-verify dataset and release a benchmark to assess its performance . |
| Outcome: | The proposed framework surpasses Qwen2.5-72B-Instruct on Tool-V-Bench and the previous APIGen-MT dataset. |
Copied to clipboard
| Challenge: | Existing methods for grounding large language models suffer from inefficient querying . Existing approaches that rely on physical verification or self-reflection suffer from excessive querying. |
| Approach: | They propose a framework that introduces Reinforced Advantage feedback for efficient self-refinement of plans. |
| Outcome: | The proposed framework surpasses baselines in success rate and significantly decreases interaction steps of agents and query rounds of LLMs. |
Copied to clipboard
| Challenge: | Experimental results show document-level translation repair improves translation consistency but still suffers from lexical translation inconsistency due to the lack of inter-sentence context. |
| Approach: | They propose a document-level translation repair model to model translation inconsistency via automatic post-editing. |
| Outcome: | The proposed model improves translation quality and lexical consistency on document-level translation datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for conversational machine reading comprehension are inconsistent with real scenarios. |
| Approach: | They propose to use a Chinese CMRC benchmark to evaluate model's generalization ability towards diverse domains by using zero-shot/few-shot settings. |
| Outcome: | The proposed benchmarks are based on 831 hot-topic driven conversations with 4,742 turns and cover 33 domains. |
Copied to clipboard
| Challenge: | Existing approaches to explain NLP models have two key challenges: spurious correlation and degeneration. |
| Approach: | They propose a rationalization framework using a generator and a predictor to construct a self-explaining NLP model with spurious correlation and degeneration as key challenges. |
| Outcome: | The proposed method improves the F1 score by 20.9% compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are gaining popularity as scalable tools for mental health support . however, nearly half of individuals do not receive timely support due to limited selfawareness or reluctance to seek help. |
| Approach: | They propose a proactive emotional support framework that leverages principles of active listening to uncover implicit user needs. |
| Outcome: | The proposed model elicits implicit emotional needs and delivers empathetic support compared to baselines . |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, but they introduce safety and reliability risks, such as CoT-hijacking and prompt-induced inefficiencies. |
| Approach: | They propose a unified benchmark to assess the trustworthiness of Large Reasoning Models. |
| Outcome: | The proposed benchmark evaluates truthfulness, safety and efficiency on 26 models. |
Copied to clipboard
| Challenge: | Neural Processing Units (NPUs) are critical for AI infrastructure, but their development remains a bottleneck due to vendor-specific Domain-Specific Languages (DSLs). |
| Approach: | They propose a framework for NPU kernel development that bridges the gap in hardware-specific coding . compiler success on complex Level-2 kernels improves from 0% to 95.5%, they say . |
| Outcome: | The proposed framework bridges the gap in hardware-specific coding, showing a near-zero success rate on complex kernels. |
Copied to clipboard
| Challenge: | Using multiple sequence alignments (MSA) to extract evolutionary knowledge is limited. |
| Approach: | They propose to use multiple sequence alignments to augment protein representations . they propose to employ Retrieved Sequence Augmentation to enhance protein representation learning . |
| Outcome: | The proposed method surpasses MSA Transformer by 5% in structural and property prediction tasks while being 373 times faster. |
Copied to clipboard
| Challenge: | Existing long-context benchmarks do not accurately evaluate large language models’ comprehension and reasoning abilities in extended texts. |
| Approach: | They propose a new evaluation benchmark that adopts a multiple-choice question format and uses a multi-choke question format to assess the comprehension and reasoning skills of large language models. |
| Outcome: | The proposed benchmark provides a rapid, precise, and unbiased appraisal of the long-context comprehension skills of large language models. |
Copied to clipboard
| Challenge: | Experimental results show that ChatCRS improves language quality and informativeness by 17% and proactivity by 27%. |
| Approach: | They propose a framework to decompose the CRS task into several sub-tasks . they propose 'knowledge retrieval agent' and 'goal-planning agent' |
| Outcome: | The proposed framework improves language quality and informativeness by 17% and proactivity by 27% on two multi-goal CRS datasets. |
Copied to clipboard
| Challenge: | Existing general-domain benchmarks do not capture complexity of real-world judicial cognition and decision-making. |
| Approach: | They propose a benchmark specifically designed to evaluate LLM Agents in the legal domain. |
| Outcome: | The proposed benchmark includes 17 corpora from real-world legal scenarios and provides 37 tools for interacting with external knowledge. |
Copied to clipboard
| Challenge: | Recent advances in large reasoning models have broadened the capabilities of medical artificial intelligence. |
| Approach: | They propose a reasoning framework for complex medical inference that reformulates medical reasoning as a parallelizable directed acyclic graph process based on Petri Net theory. |
| Outcome: | The proposed reasoning framework improves strong general-purpose LLMs by up to 8.9%. |
Copied to clipboard
| Challenge: | Existing methods to extract product features from unstructured text still suffer from problems . e-commerce platforms are focusing on multi-scale values, which can be confusing . |
| Approach: | They propose a pre-training technique to automatically obtain attribute value pairs from product descriptions to aid e-commerce. |
| Outcome: | The proposed method improves on the existing token-level masking strategy and achieves state-of-the-art on four benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for grammatical error correction (GEC) are mainly divided into detection-based and end-to-end generative models. |
| Approach: | They propose an end-to-end framework which Leverages Error Type (LET) information in the generation process to introduce more convincing error type information. |
| Outcome: | The proposed framework outperforms existing methods on various datasets by a clear margin. |
Copied to clipboard
| Challenge: | Existing methods for conversational KBQA assume the independence of utterances and model them in isolation. |
| Approach: | They propose a History Semantic Graph Enhanced KBQA model that models long-range semantic dependencies in conversation history while maintaining low computational cost. |
| Outcome: | The proposed model outperforms baselines on a widely used question type dataset. |
Copied to clipboard
| Challenge: | a survey of older adults shows that many LLMs mishandle elderly-specific contextual risks. |
| Approach: | They propose a framework to assess elderly-specific contextual risks in LLM interactions . they use a taxonomy to identify 50 fine-grained risk types across mental well-being, financial, medical, toxicity, and privacy domains . |
| Outcome: | a new framework assesses elderly-specific contextual risks in LLM interactions . it achieves 96.2% and 90.9% unsafe-prompt detection accuracy, respectively . |
Copied to clipboard
| Challenge: | Existing approaches to building generalizable verifiable data are task-specific and lack a principled, universal evaluator of verifikatability. |
| Approach: | They propose a task-agnostic, strategy-guided, executably-checkable data synthesis framework that synthesizes problems, diverse candidate solutions and verification artifacts from a single source. |
| Outcome: | The proposed framework synthesizes problems, candidates, and verification artifacts from human-annotated and strategy-induced checks and iteratively discovers strategies. |
Copied to clipboard
| Challenge: | Parody is an emerging phenomenon on social media, where individuals imitate a role or position opposite to their own . limited available data and deficient diversity in current datasets hinder study of parody . |
| Approach: | They build a dataset of parody users and annotated comments from both English and Chinese corpora to test parody detection and comment sentiment analysis. |
| Outcome: | The proposed datasets provide richer contextual information, which is lacking in existing datasets. |
Copied to clipboard
| Challenge: | Existing studies on spatial intelligence from the perspective of visual-spatial intelligence have not explored whether visual intelligence alone is sufficient to endow models with spatial intelligence. |
| Approach: | They propose to use a linguistic perspective to investigate spatial intelligence from a theoretical perspective. |
| Outcome: | The proposed model performs poorly on the proposed dataset while human can easily achieve 100% accuracy. |
Copied to clipboard
| Challenge: | Existing studies on the effectiveness of the Retentive Networks have not yet been conducted. |
| Approach: | They propose a retention mechanism that integrates the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models. |
| Outcome: | The proposed retention mechanism combines the inductive bias of recurrent neural networks with the parallelizable training advantages of attention-based models. |
Copied to clipboard
| Challenge: | Existing prompt engineering techniques are limited to producing single flow instructions, struggling with handling diverse patterns. |
| Approach: | They propose an automatic prompt optimization method that iteratively develops a multi-branched prompt using failure cases as feedback. |
| Outcome: | The proposed method achieves the best results across five tasks and demonstrates significant optimization efficiency due to adoption of a minimal search strategy. |
Copied to clipboard
| Challenge: | tens of thousands of ancient characters must be deciphered by experts to interpret unearthed documents. |
| Approach: | They propose a diachronic Chinese knowledge base to help researchers discover glyph similar characters by measuring glyph similarities between ancient Chinese characters. |
| Outcome: | The proposed method shows strong correlations between the scores obtained from the method and from human experts. |
Copied to clipboard
| Challenge: | Recent advances in multimodal recommenders lack explicit reasoning and self-awareness of uncertainty. |
| Approach: | They propose a reasoning-augmented multimodal agent structured around a three-stage explicit reasoning pipeline. |
| Outcome: | The proposed agent improves ranking metrics and performance on four standard recommendation tasks across five real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods for supervised fine-tuning focus on unit test feedback to construct preference pairs. |
| Approach: | They propose a preference alignment framework that mimics human iterative debugging to refine Code LLMs. |
| Outcome: | Experiments show that Preference Learning improves on BigCodeBench and BigCodeBind tasks. |
Copied to clipboard
| Challenge: | Existing datasets focus on relation extraction between two entities in one sentence, and some focus on cross-sentence relationships. |
| Approach: | They propose to use a Chinese multi-party dialogue dataset for automatic extraction of dialogue-based character relationships. |
| Outcome: | The proposed dataset extracts relationships between 140 entities on the CRECIL corpus and another existing relation extraction corpus. |
Copied to clipboard
| Challenge: | Self-supervised representation learning (SSL) uses proxy supervised learning tasks to obtain training data from unlabeled corpora. |
| Approach: | They propose to survey the latest SSL techniques, tools, datasets, and performance achievement in speech processing to scale up current machine learning technologies. |
| Outcome: | The proposed tutorial is highly relevant to the special theme of ACL about language diversity. |
Copied to clipboard
| Challenge: | Existing research focused on model-specific adversarial methods, but real-world applications demand a more generalizable approach to audio adversarials. |
| Approach: | They propose a Chat-Audio Attacks benchmark to evaluate LALMs' robustness . they propose standard evaluation, GPT-4o-based evaluation and human evaluation . |
| Outcome: | The proposed benchmark aims to explore the robustness of six state-of-the-art LALMs with voice interaction capabilities. |
Copied to clipboard
| Challenge: | BrainLoc is a lightweight object detection model guided by fMRI signals. |
| Approach: | They propose a brain-based object detection model guided by fMRI signals . they employ a multi-modal alignment strategy that enhances fmr feature extraction . |
| Outcome: | The proposed model improves fMRI-based object detection accuracy and convenience. |
Copied to clipboard
| Challenge: | Syntactic parsing aims to reveal how sentences are syntactically structured. |
| Approach: | They propose to produce compatible constituency and dependency trees simultaneously for input sentences . they adopt a much more efficient decoding algorithm and explore joint modeling at training phase . |
| Outcome: | The proposed model significantly improves matching ratio of whole trees compared to separate models . the proposed model adopts a much more efficient decoding algorithm . |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating long-context language models employ irrelevant noise texts to artificially extend the length of test cases, diverging from the real-world scenarios of long-constituency applications. |
| Approach: | They propose a long-context benchmark, Loong, aligning with realistic scenarios through extended multi-document question answering (QA) . |
| Outcome: | The proposed model can scale up the context window of large language models to perform in-depth analysis of multiple long documents. |
Copied to clipboard
| Challenge: | Existing studies on automatic summary evaluation metrics focus on lexical similarity and require a reference summary which is expensive to obtain. |
| Approach: | They propose to use a weakly supervised summary evaluation approach without the presence of reference summaries to transform existing summarization datasets into corrupted reference summarizers. |
| Outcome: | The proposed method outperforms baselines and shows that it improves linguistic quality over all metrics. |
Copied to clipboard
| Challenge: | Existing studies treat named entity recognition as a sequential labeling problem. |
| Approach: | They propose a span selection framework for nested named entity recognition . they propose nesting entities with different input categories would be separately extracted . |
| Outcome: | The proposed framework outperforms competing models on four benchmark datasets. |
Copied to clipboard
| Challenge: | Large language models suffer from repetitive text generation, a phenomenon we refer to as the ”Repeat Curse”. |
| Approach: | They propose a method to induce and analyze the Repeat Curse in large language models by using mechanistic interpretability. |
| Outcome: | The proposed method induces and analyzes the Repeat Curse in large language models using mechanistic interpretability. |
Copied to clipboard
| Challenge: | Existing studies on visual deep semantics focus primarily on superficial description of images, revealing a notable deficiency in the systematic investigation of the inherent deep semantic. |
| Approach: | They propose a benchmark to assess Large Multimodal Models’ (LMMs) capacities of visual deep semantics. |
| Outcome: | The proposed benchmark demonstrates a substantial gap between the deep semantic comprehension capabilities of existing LMMs and humans. |
Copied to clipboard
| Challenge: | Recent studies on review helpfulness prediction require labeled samples for each domain/category of interest. |
| Approach: | They propose a convolutional neural network based model which leverages word-level and character-based representations to transfer knowledge between domains. |
| Outcome: | The proposed model outperforms the state-of-the-art on the Amazon product review dataset. |
Copied to clipboard
| Challenge: | Existing state-of-the-art event coreference resolution systems rely on spurious and spurious associations in the input mention pair text. |
| Approach: | They propose a rationale-centric counterfactual data augmentation method that leverages the debiasing capability of counterfact data haussed by LLM-in-the-loop to mitigate spurious association while emphasizing causation. |
| Outcome: | The proposed method achieves state-of-the-art on three popular cross-document benchmarks and demonstrates robustness in out-of domain scenarios. |
Copied to clipboard
| Challenge: | Existing text augmentation methods generate instances with shifted feature spaces, which leads to a drop in performance on large datasets. |
| Approach: | They propose a hybrid instance-filtering framework that generates instances with shifted feature spaces, which leads to a drop in performance on augmented data. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three classification tasks and nine public datasets. |
Copied to clipboard
| Challenge: | Language models exhibit increasingly consciousness-like behaviors, requiring a baseline to evaluate their cognitive abilities. |
| Approach: | They propose a benchmark to assess the cognitive abilities of language models (LMs) they compare 18 state-of-the-art LMs to human models in metacognition, self-awareness, social awareness and situational awareness . |
| Outcome: | Evaluating 18 state-of-the-art LMs, they find they consistently surpass baselines . but most models fall short in metacognition and self-awareness, the study finds . |
Copied to clipboard
| Challenge: | Existing approaches to solve the data imbalance problem are limited in extremely imbalanced data. |
| Approach: | They propose a hybrid approach which adapts general networks for head categories and few-shot techniques for tail categories. |
| Outcome: | The proposed approach improves the performance of Single networks with diverse loss objectives on tail or entire categories. |
Copied to clipboard
| Challenge: | Despite recent advances in speech-to-text translation, the impact of the emotion content has been overlooked. |
| Approach: | They propose to use generative error correction (GER) to generate the translation based on the decoded N-best hypotheses and combine emotion and sentiment labels into the LLM finetuning process to enable the model to consider the emotion content. |
| Outcome: | The proposed model can translate speech in English-Chinese using GER and emotion and sentiment labels. |
Copied to clipboard
| Challenge: | a new approach to training large language models (LLMs) overlooks task-specific characteristics in tool use, leading to performance bottlenecks. |
| Approach: | They propose a task-feature-based framework that mitigates the effects of suboptimal training data . they use a dataset to train large-scale LLMs and a reward mechanism tailored to error categories . |
| Outcome: | The proposed framework matches or surpasses open- and closed-source LLMs in tool-use performance using only 1,217 training data points. |
Copied to clipboard
| Challenge: | Existing methods for distantly supervised relation extraction suffer from noisy labeling problem, which can severely degrade its performance. |
| Approach: | They propose a framework for distantly supervised relation extraction that leverages text corpus and knowledge graph and a cooperative module involving their mutual learning. |
| Outcome: | The proposed method reduces the noisy labels and achieves substantial improvement over the state-of-the-art methods. |
Copied to clipboard
| Challenge: | Current methods require large amount of bilingual training data, which is challenging and sometimes impossible task. |
| Approach: | They propose a method to modify the style of inputs by modifying the source side of BT data. |
| Outcome: | The proposed method significantly improves translation quality against popular BT benchmarks on high-resource and low-resourced language pairs. |
Copied to clipboard
| Challenge: | a new evaluation platform for large language models and text-driven AIGCs is available for free. |
| Approach: | They propose an evaluation platform for side-by-side comparisons of large language models and text-driven AIGC systems. |
| Outcome: | a new evaluation platform for large language models and text-driven AIGC systems is available for free . the platform is more focused on the Chinese language and more models developed by Chinese institutes . |
Copied to clipboard
| Challenge: | Existing multilingual table benchmarks suffer from geolinguistic imbalance - overrepresenting certain languages and lacking sufficient scale for rigorous cross-lingual analysis. |
| Approach: | They propose a framework for massively multilingual table question answering that includes tables expanded to 97 languages from Chinese and English sources. |
| Outcome: | Experiments on state-of-the-art LLMs show that synthetically generated training data significantly boosts performance, especially for low-resource languages. |
Copied to clipboard
| Challenge: | Existing studies on short text classification focus on long texts and achieve unsatisfactory performance due to the sparsity and limited labeled data. |
| Approach: | They propose a heterogeneous graph neural network based method for semi-supervised short text classification that leverages the full advantage of few labeled data and large unlabeled data through information propagation along the graph. |
| Outcome: | The proposed method outperforms state-of-the-art methods across six benchmark datasets significantly. |
Copied to clipboard
| Challenge: | Existing models often rely on specific words to predict offensive content, compromising model fairness and potentially exacerbates biases against vulnerable and minority groups. |
| Approach: | They propose a bias self-awareness and data self-iteration framework to help models identify and mitigate biases by integrating multiple natural language processing techniques. |
| Outcome: | The proposed framework reduces false positive rate of models in in-distribution and out-of-difference tests, enhances model accuracy and fairness, and shows promising performance improvements on larger datasets. |
Copied to clipboard
| Challenge: | Existing large language model evaluation benchmarks focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. |
| Approach: | They propose a comprehensive benchmark covering 29 languages, built on an English benchmark. |
| Outcome: | The MMLU-ProX is a comprehensive benchmark covering 29 languages, built on an English benchmark. |
Copied to clipboard
| Challenge: | Existing benchmarks to evaluate LLMs' capabilities are inadequate for assessing their musical capabilities. |
| Approach: | They propose to use a large-scale music benchmark specifically designed to evaluate the music-related capabilities of large language models (LLMs). |
| Outcome: | The proposed framework evaluates 16 large language models in the domain of music. |
Copied to clipboard
| Challenge: | Existing multimodal neural machine translation models focus on bilingual translation, but experimental results show that they outperform the text-only baselines and multilingual multimodal methods by a large margin. |
| Approach: | They propose a framework to leverage the multimodal prompt to guide the Multimodal Multilingual Neural Machine Translation (m3P) this framework aligns the representations of different languages with the same meaning and generates the conditional vision-language memory for translation. |
| Outcome: | The proposed framework outperforms previous text-only baselines and multilingual multimodal methods by a large margin. |
Copied to clipboard
| Challenge: | Existing work has treated procedures as shallow structures without modeling the parent-child relation. |
| Approach: | They propose to construct an open-domain hierarchical knowledge-base (KB) of procedures based on wikiHow . they link steps in an article to other articles with similar goals, recursively building the KB . |
| Outcome: | The proposed method significantly outperforms baselines according to automatic evaluation, human judgment, and application to downstream tasks such as instructional video retrieval. |
Copied to clipboard
| Challenge: | Competitive programming has become a key task for training and evaluating large language models . but test cases of competitive programming problems are often difficult to obtain . |
| Approach: | They propose an LLM-based agent system that creates high-quality test cases for competitive programming problems. |
| Outcome: | The proposed system improves code tests on a CodeContests dataset with pass/fail labels. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have brought significant changes to various domains, especially through autonomous agents. |
| Approach: | They propose a framework that lets agents learn shortcuts from their past tasks and use them for future task execution. |
| Outcome: | The proposed framework enables agents to tackle unseen software-developing tasks more effectively. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive performance, but they keep repeating similar mistakes due to their inability to capture relationships among samples. |
| Approach: | They propose a tuning-free rule accumulation framework that guides LLMs in improving their performance by learning from previous mistakes. |
| Outcome: | The proposed framework improves over baselines by a large margin over previous frameworks. |
Copied to clipboard
| Challenge: | Model merging is an effective technique for composing the capabilities of a multilingual model and a reasoning model. |
| Approach: | They propose a model merging framework that modulates the contribution of each source model. |
| Outcome: | Experiments show that the proposed model merging framework outperforms strong baselines on multilingual reasoning benchmarks across 21 different languages. |
Copied to clipboard
| Challenge: | Existing studies on LLMs focused on supervised fine-tuning but their effectiveness has been limited. |
| Approach: | They propose a paradigm consisting of three stages: Secondary Pre-training using extensive monolingual data, Continual Pre- training with interlinear text format documents, and Leveraging source-language consistent instruction for supervised fine-tuning. |
| Outcome: | The proposed approach surpasses previous work and achieves superior performance compared to models such as NLLB-54B(CITATION) and GPT3.5-text-davinci-003. |
Copied to clipboard
| Challenge: | Considering the vast size and wide-ranging sources of LLMs’ training data, it could explicitly or implicitly include test data. |
| Approach: | They propose a Contamination Detection via output Distribution (CDD) which detects data contamination only by identifying the peakedness of LLM's output distribution. |
| Outcome: | The proposed method improves performance by 21.8%-30.2% on humanEval and TED: trustworthy evaluation via output distribution. |
Copied to clipboard
| Challenge: | Deciphering oracle bone scripts using AI technology is not an overnight task due to the evolution of written language over millennia. |
| Approach: | They propose a framework that utilizes Large Multi-modal Models (LMMs) for interpreting Oracle Bone Script (OBS). |
| Outcome: | The proposed framework provides quantitative analyses and superior deciphering capability. |
Copied to clipboard
| Challenge: | Existing Chinese preference datasets suffer from limited scale, restricted domain coverage, and insufficiently rigorous data validation. |
| Approach: | They propose an LLM-based data annotation pipeline with no human intervention to annotate Chinese preference datasets. |
| Outcome: | The proposed pipeline outperforms existing Chinese preference datasets on AlignBench and Chinese Reward Benchmark. |
Copied to clipboard
| Challenge: | Existing backdoor attacks are not stealthy to system deployers or users. |
| Approach: | They propose a novel backdoor attack method based on negative data augmentation and modifying word embeddings that is much stealthier while maintaining pretty good attacking performance. |
| Outcome: | The proposed method is much stealthier while maintaining pretty good attacking performance. |
Copied to clipboard
| Challenge: | Experimental results show that UniCoder with the universal code significantly outperforms the previous prompting methods by a large margin. |
| Approach: | They introduce the universal code (UniCode) as the intermediate representation of algorithm steps using conventions of programming languages. |
| Outcome: | The proposed model outperforms previous prompting methods by a large margin . the proposed model is based on a dataset of natural-language questions and code solutions . |
Copied to clipboard
| Challenge: | Existing retrieval-augmented approaches focus on ignoring the structural information of the Knowledge Base (KB) and the question. |
| Approach: | They propose a structure-aware subgraph retrieval stage that ranks candidate subgraphs by aligning them with the question’s structure, along with semantic relevance. |
| Outcome: | Experiments on GrailQA, WebQSP, and GraphQuestions show that the proposed framework achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Empirical evaluations on ML models show substantial reductions in unsafe generations and improved robustness against jailbreak attacks. |
| Approach: | They propose a resource-efficient pruning framework that directly identifies unsafe behaviors while preserving model utility. |
| Outcome: | The proposed framework reduces unsafe generations and improves robustness against jailbreak attacks with minimal utility loss. |
Copied to clipboard
| Challenge: | Existing evaluations focus on problem-solving from examiner perspective, overlooking a dual perspective of examiner regarding error identification and correction. |
| Approach: | They propose to use an annotated dataset to evaluate large language models from the examiner perspective and to use diverse prompts to evaluate eleven representative LLMs. |
| Outcome: | The proposed model outperforms all models while LLaMA-2-7B has comparable abilities to closed-source models GPT-3.5 and Gemini Pro. |
Copied to clipboard
| Challenge: | Automatic Chinese poetry generation is one of the first attempts towards computer writing. |
| Approach: | They propose a model which requires no supervised style labeling to generate stylistic poems . they incorporate mutual information, a concept in information theory, into modeling . |
| Outcome: | The proposed model generates stylistic poems without losing fluency and coherency . it is based on mutual information, a concept in information theory . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on correctness, overlooking optimality . large language models excel at math, coding, logic and puzzles . |
| Approach: | They propose a framework for training and evaluating Large Language Models on NP-hard optimization problems through quality-aware RLVR. |
| Outcome: | The proposed framework outperforms existing benchmarks on math, coding, logic and puzzles. |
Copied to clipboard
| Challenge: | Current Event Extraction methods focus on high-resource scenarios, which requires large amount of annotated data. |
| Approach: | They propose a demonstration-based learning paradigm for EE to fully use annotated data . they propose EE as a natural language generation task guided by schema-based prompts . |
| Outcome: | The proposed model outperforms current methods in low-resource scenarios. |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucinations are confounded by epistemic uncertainty and cannot distinguish genuine uncertainty from fabricated content. |
| Approach: | They propose a model-agnostic metric that captures epistemic boundary deviations by measuring answer-level stability across multiple stochastic forward passes. |
| Outcome: | The proposed metric outperforms strong uncertainty-only baselines and can be used to detect hallucinations on open-domain question answering, dialogue generation, and code completion. |
Copied to clipboard
| Challenge: | Rapid advances in multimodal large language models have revolutionized cross-modality understanding. |
| Approach: | They propose a method that uses whitening transformations to adjust MLLM representation spaces . they propose ML models that are dominated by textual semantics and visual semantics . |
| Outcome: | The proposed approach improves zero-shot multimodal retrieval performance without fine-tuning efforts. |
Copied to clipboard
| Challenge: | OpenAI's GPT-4 has demonstrated remarkable multimodal capabilities, but specific mechanics of GPT4 remain unknown. |
| Approach: | They propose a data collection methodology that synchronously synthesizes images and dialogues for visual instruction tuning. |
| Outcome: | The proposed method improves on ten commonly assessed models and provides greater flexibility compared to existing methods. |
Copied to clipboard
| Challenge: | Current Chain-of-thought Distillation methods hinder CoT reasoning performance . student models are separately distilled from specific reasoning tasks . parameter update of student models severely harms CoT ability on unseen reasoning tasks. |
| Approach: | They propose a method which distills Chain-of-thought reasoning ability of large language models to much smaller student models. |
| Outcome: | The proposed method improves the reasoning ability of large language models on 14 datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are powerful at question-answering but prone to hallucinations due to limited domain-specific or up-to-date knowledge. |
| Approach: | They propose a framework for IDentifying RAG properties in LLM services that integrates LLMs with retrieval systems and adds an external retriever and knowledge database to mitigate hallucinations. |
| Outcome: | The proposed framework detects RAG-enhanced LLMs with 99.97% accuracy with partial or no optional knowledge and nearly 100% when the LLM and database are known. |
Copied to clipboard
| Challenge: | Recent efforts to train speech-only LLMs have led to models “forging” speech information from text-only models. |
| Approach: | They propose a paradigm for training Speech Large Language Models without instruction data by using the response of a text-only LLM to transcripts as self-supervision. |
| Outcome: | The proposed model generalizes to Spoken Question Answering, Classification, and Translation and achieves a 72% win rate compared with state-of-the-art models like Qwen 2 Audio . |
Copied to clipboard
| Challenge: | Discourse relations are annotated on scientific articles. |
| Approach: | They propose a domain-specific discourse treebank annotated on scientific articles . they use dependency trees to represent discourse structure, which is flexible and simplified . |
| Outcome: | The proposed treebank is a benchmark for evaluating discourse dependency parsers. |
Copied to clipboard
| Challenge: | Introspection-driven approach equips LLM agents with introspection, enhancing consistency and adaptability in solving complex tasks. |
| Approach: | They propose a zero-shot approach that equips LLM agents with introspection, enhancing consistency and adaptability in solving complex tasks. |
| Outcome: | The proposed approach improves performance and efficiency by reducing the number of trials and plan revisions by 45%. |
Copied to clipboard
| Challenge: | Large Language Models generate inconsistent and sometimes contradictory outputs when presented with a prompt that has equivalent semantics but is expressed differently from the original prompt. |
| Approach: | They propose to refine a Large Language Model (LLM) with prompt-output pairs with equivalent semantics to achieve semantic consistency. |
| Outcome: | The proposed method improves the semantic consistency and task performance of LLMs. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) based sparse architectures are prone to overfitting on low-resource language translation. |
| Approach: | They propose a modularized MNMT framework that flexibly assembles dense and MoE-based sparse modules to achieve the best of both worlds. |
| Outcome: | The proposed framework outperforms existing models on low-resource language translation and zero-shot translation on benchmark datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promise as automated evaluators for assessing the quality of answers generated by AI systems. |
| Approach: | They propose an alignment-based system that calibrates position bias in a lightweight yet effective manner by taking into account both length and semantics and combining them into a single prompt. |
| Outcome: | Extensive experiments with six LLMs on 11,520 answer pairs show that PORTIA significantly improves consistency and consistency rates with humans. |
Copied to clipboard
| Challenge: | Goal-oriented script planning is used by humans to plan for typical activities . however, this capability remains underexplored due to several challenges . |
| Approach: | They propose a framework that enables product-enriched scripts by associating products with each step based on the semantic similarity between the actions and their purchase intentions. |
| Outcome: | The proposed framework can generate product-enriched scripts from 2.4 million scripts . human annotations are conducted to provide gold labels for a sampled subset . |
Copied to clipboard
| Challenge: | Existing methods to perform implicit knowledge transfer from machine translation to ST model are difficult because of the task complexity and data scarcity. |
| Approach: | They recommend a method which conducts explicit knowledge transfer from MT to ST model by fine and coarse granularity contrastive learning. |
| Outcome: | The proposed method improves the performance of the end-to-end speech translation model on all 8 languages. |
Copied to clipboard
| Challenge: | Existing work on document-level ASR error correction ignores contextual information . however, there are limited studies on incorporating contextual information into AEC . |
| Approach: | They propose a context-aware method that retrieves contextual information from a datastore . they use two English and two Chinese datasets to model document-level AEC . |
| Outcome: | The proposed model can utilize contextual information to improve document-level AEC . the data store containing contextual information provides even better results . |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable reasoning capabilities, but they still face challenges in knowledge-intensive multi-hop reasoning. |
| Approach: | They propose a method that uses self-critique feedback to guide iterative reasoning by enabling iteration and self-evaluation of its intermediate reasoning steps. |
| Outcome: | The proposed method surpasses the previous SOTA by 8.6% on three multi-hop reasoning datasets. |
Copied to clipboard
| Challenge: | et al., 2024) show large language models exhibit internal chain-of-thought, meaning they decompose and execute composite tasks layer-by-layer. |
| Approach: | They propose a method to decode hidden states using LogitLens . they also propose 'chain of thought' to decompose and execute composite tasks . |
| Outcome: | The proposed method decodes hidden states and shows consistent execution pattern . it opens avenues for fine-grained, instruction-level activation steering. |
Copied to clipboard
| Challenge: | Current approaches to decoding language from the human brain rely on unimodal representations, neglecting the brain’s inherently multimodal processing. |
| Approach: | They propose a framework that leverages Multimodal Large Language Models to align brain signals with a shared semantic space encompassing text, images, and audio. |
| Outcome: | The proposed framework achieves an 8.48% improvement on the most commonly used benchmark on fMRI datasets with textual, visual, and auditory stimuli. |
Copied to clipboard
| Challenge: | Emotion plays a crucial role in human conversation. |
| Approach: | They present a MELD-ST dataset for the emotion-aware speech translation task . they show that fine-tuning with emotion labels can enhance translation performance . |
| Outcome: | The proposed dataset shows that fine tuning with emotion labels can improve translation performance in some settings. |
Copied to clipboard
| Challenge: | stance detection studies focus on evaluating stances within individual instances, hindering progress of conversational stance analysis. |
| Approach: | They propose a multi-turn conversation stance detection dataset that encompasses multiple targets for conversational stance detector. |
| Outcome: | The proposed dataset encompasses multiple targets for conversational stance detection. |
Copied to clipboard
| Challenge: | Existing methods for detecting factual inconsistencies in abstractive summarization are lacking in factual consistency detection. |
| Approach: | They propose to use real model-generated summaries with human annotations to detect factual inconsistencies. |
| Outcome: | The proposed model outperforms the SOTA on CoGenSumm, FactCC, Frank, and SummEval datasets. |
Copied to clipboard
| Challenge: | Existing methods for detecting LLM-Generated text suffer from distribution misalignment and limited interpretability. |
| Approach: | They propose a statistical framework utilizing supervised subspace learning to extract compact features and construct conditional semantic distributions based on syntactic structures. |
| Outcome: | The proposed framework is superior in cross-domain, cross-model, and adversarial scenarios. |
Copied to clipboard
| Challenge: | Existing knowledge poisoning attacks against RAG systems require multiple poisoned documents or can only function effectively on simplistic queries. |
| Approach: | They propose a more realistic knowledge poisoning attack that poisons only a single document while remaining effective for complex multi-hop questions involving complex relationships between multiple elements. |
| Outcome: | The proposed attack achieves success by poisoning only a single document while remaining effective for complex multi-hop questions involving complex relationships between multiple elements. |
Copied to clipboard
| Challenge: | Existing methods for storytelling lack coherence and consistency, compromising the overall storytelling experience. |
| Approach: | They propose a novel approach that improves the coherence and consistency of automatically generated stories by managing plot nodes and enabling dynamic interactions between different parts of the story. |
| Outcome: | The proposed approach outperforms existing methods in 84.33% of the trials. |
Copied to clipboard
| Challenge: | Prior studies have focused on the role of well-chosen examples in in-context learning . |
| Approach: | They propose to use multiple translational factors for in-context example selection by using monotone submodular function maximization. |
| Outcome: | The proposed approach outperforms random selection and robust single-factor baselines across various NLP tasks. |
Copied to clipboard
| Challenge: | Existing methods for text watermarking rely on arbitrary vocabulary partitioning during decoding, which compromises the availability of suitable tokens and significantly degrades the quality of responses. |
| Approach: | They propose a method that leverages linguistic prior knowledge of lexical redundancies in LLM vocabularies to seamlessly integrate watermarks. |
| Outcome: | The proposed approach preserves the expressive power of large language models while preserving watermark detectability. |
Copied to clipboard
| Challenge: | Existing methods for modifying parameters are unsystematic and rely on empirical experience. |
| Approach: | They propose a controllable alignment prompting for unlearning framework that decouples unlearning into a learnable prompt optimization process via reinforcement learning. |
| Outcome: | The proposed framework achieves precise, controllable unlearning without updating model parameters. |
Copied to clipboard
| Challenge: | Existing benchmarks for extracting structured procedural knowledge from unstructured business documents are limited by simplistic schemas and shallow logical dependencies. |
| Approach: | They propose a framework for extracting structured procedural knowledge from unstructured business documents . they propose BREX, a carefully curated benchmark comprising 409 real-world business documents and 2,855 expert-annotated rules . |
| Outcome: | The proposed framework outperforms standard prompts in rule extraction and execution. |
Copied to clipboard
| Challenge: | under the pandemic of COVID-19, people experiencing COVI D19-related symptoms have a pressing need to consult doctors. |
| Approach: | They develop a medical dialog system that can provide COVID19-related consultations . they use two dialog datasets containing conversations between doctors and patients . |
| Outcome: | The proposed system can provide COVID19-related consultations, but is too small compared with general-domain dialog datasets. |
Copied to clipboard
| Challenge: | Existing methods to remove undesirable data from Large Language Models suffer from cumulative catastrophic utility loss under continuous unlearning requests. |
| Approach: | They propose a method that leverages the rotational salience weight of RCU to quantify and control the unlearning degree in the continuous unlearning process. |
| Outcome: | The proposed method achieves SOTA performance without a retained dataset. |
Copied to clipboard
| Challenge: | Contemporary large language models (LLMs) are pre-trained on multilingual corpora, but their performance lags behind in most languages compared to a few resource-rich languages. |
| Approach: | They propose a method that leverages the internal capabilities of large language models on resource-rich languages to enhance multilingual performance. |
| Outcome: | The proposed method improves multilingual performance while minimizing impact on original performance in resource-rich languages. |
Copied to clipboard
| Challenge: | Recent work has shown that reinforcement learning with simple rule-based reward functions (RLVR) can induce emergent reasoning behaviors and yield gains in challenging domains such as math problem solving. |
| Approach: | They propose a rollout-alignment-quantization-aware RL which aligns training-side forward with the quantized rollout to minimize mismatch. |
| Outcome: | The proposed approach outperforms quantized-rollout training by +5.5 on Qwen3-30B-A3B MoE for math problems while maintaining low-bit throughput. |
Copied to clipboard
| Challenge: | In language model training, it is difficult to obtain the right data mixtures for various tasks as the relationship between data and tasks is difficult. |
| Approach: | They propose to identify checkpoint models based on their respective capabilities and leverage them as data mixers by using their aggregated first-order influence approximation over source data. |
| Outcome: | The proposed framework shows significant improvements on eight reasoning benchmarks, with accuracy increases of up to 1.93%. |
Copied to clipboard
| Challenge: | Pre-trained models such as BERT have achieved success in learning sequence representations, but they tend to learn representations that are covariant with the noise of pre-training. |
| Approach: | They propose to train self-trained models to learn noise invariant sequence representations . they encourage consistency between original sequence and corrupted version via unsupervised instance-wise training signals. |
| Outcome: | The proposed model improves on 11 natural language understanding and cross-modal tasks and achieves 0.6% gain on GLUE benchmarks and 0.8% increment on NLVR2 . |
Copied to clipboard
| Challenge: | Existing safety defenses for large language models fail to explicitly repel harmful patterns . Optimal transport (SOT) allows for safe fine-tuning without sacrificing safety . |
| Approach: | They propose a framework that reframes safe fine-tuning from instance-level filtering challenge to distribution-level alignment task grounded in Optimal Transport. |
| Outcome: | a new framework improves safety of large language models while maintaining competitive performance . the proposed framework reduces the risk of errors and improves model performance compared to baselines . |
Copied to clipboard
| Challenge: | Recent studies have used dependency trees to extract relation between aspects and contexts, but there is a potential mismatch between the dependency tree and sentiment classification as a semantic task. |
| Approach: | They propose to replace the syntactic dependency tree with a semantic structure to capture the relation between an aspect and a context. |
| Outcome: | The proposed model improves ABSA on four public datasets with 1.13% improvement over baselines. |
Copied to clipboard
| Challenge: | Generative dialogue systems tend to produce generic and boring responses, causing boring conversations . a novel commonsense knowledge-aware dialogue generation model is proposed to solve this problem . |
| Approach: | They propose to retrieve and introduce knowledge facts from knowledge graphs to reduce boring conversations . they use a Felicitous Fact mechanism to help the model focus on context-relevant knowledge facts . |
| Outcome: | The proposed model outperforms the state-of-the-art approach in most experiments. |
Copied to clipboard
| Challenge: | Existing studies aim to integrate multiple sub-tasks into a unified ABSA model but suffer from major disadvantages . |
| Approach: | They propose a multi-task learning approach to make use of sub-tasks for a unified ABSA. |
| Outcome: | The proposed model can work well when some sub-tasks are absent, and the interactive relations among subtasks not adequate. |
Copied to clipboard
| Challenge: | Existing large language models rely on one-shot output without explicit verification, resulting in rough, incomplete, and potentially unsafe treatment plans. |
| Approach: | They propose an agentic framework that replaces one-shot generation with an iterative generate-judge-refine pipeline. |
| Outcome: | The proposed framework achieves state-of-the-art results on HealthBench, leading in Accuracy and Completeness. |
Copied to clipboard
| Challenge: | Anomaly detection (AD) is an important machine learning task, but its effectiveness in detecting harmful content, phishing attempts, and spam reviews is limited. |
| Approach: | They introduce NLP-ADBench, the most comprehensive NLP anomaly detection benchmark to date . it includes eight curated datasets and 19 state-of-the-art algorithms . |
| Outcome: | The NLP-ADBench benchmark includes 19 state-of-the-art methods and 8 curated datasets . no single model dominates across all datasets, indicating need for automated model selection . |
Copied to clipboard
| Challenge: | Multimodal large language models are increasingly deployed in open-ended, real-world environments where inputs are messy, underspecified, and not always trustworthy. |
| Approach: | They evaluate multimodal large language models in real-world environments where inputs are messy, underspecified, and not always trustworthy. |
| Outcome: | The proposed models fail to detect hidden issues even when they possess the necessary perceptual and reasoning skills. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are inherently dual-use and can be leveraged for both beneficial and harmful purposes. |
| Approach: | They propose a retention-prioritized gradient synthesis framework that decouples task-specific gradient extraction from conflict-aware combination. |
| Outcome: | The proposed method achieves tighter alignment on WMDP Bio and RWKU benchmarks. |
Copied to clipboard
| Challenge: | Existing models for word sense disambiguation lack images or senses in textual and visual datasets. |
| Approach: | They propose a unified image-text WSD model that uses image-sense complementarity to generate visual representations for word senses and a disambiguation-oriented image-sensor dataset to provide implicit textual representations. |
| Outcome: | The proposed model achieves 2.53% F1-score increase over state-of-the-art models on Textual-WSD and 2.22% HR@1 improvement on Visual-WSS. |
Copied to clipboard
| Challenge: | Existing solutions for text-to-image synthesis are sensitive on textual prompts, posing a challenge for novice users. |
| Approach: | They propose a dialogue-based TIS prompt generation model that emphasizes user experience for novice users. |
| Outcome: | The proposed model emphasizes user experience for novice users . it improves user-centricity score while maintaining a competitive quality of synthesized images. |
Copied to clipboard
| Challenge: | Existing video captioning models fail to capture nuanced semantics of videos . existing models generate coarse descriptions of human motions, resulting in poor quality . |
| Approach: | They construct a fine-grained human motion video captioning dataset named BoFiT and a model that generates fine-grain descriptions of human motions via prompting. |
| Outcome: | The proposed model outperforms existing models on comprehensive metrics. |
Copied to clipboard
| Challenge: | Existing LLMs break down on long-horizon tasks due to unbounded context growth and accumulated errors. |
| Approach: | They propose a framework that externalizes persistent state into a file-centric state abstraction and keeps the agent’s reasoning context strictly bounded regardless of task duration. |
| Outcome: | Experiments on DeepResearch and an 80-paper literature review show that the proposed framework maintains higher long-horizon coverage than baseline models without task-specific fine-tuning. |
Copied to clipboard
| Challenge: | Existing benchmarks assess basic Theory of Mind abilities but neglect temporal evolution of mental states in real-world social contexts. |
| Approach: | They propose a benchmark specifically designed to evaluate Large Language Models' ability to understand and track the temporal progression of mental states across interconnected scenarios. |
| Outcome: | The proposed benchmarks underperform humans by 44.7% and show that they can model the dynamic nature of human mental states better than existing models. |
Copied to clipboard
| Challenge: | Recent research has shown that large language models (LLMs) can enhance translation quality through self-refinement. |
| Approach: | They propose to extend translation refinement from sentence-level to document-level by using document-to-document (Doc2Doc) translations. |
| Outcome: | The proposed method improves translation quality across ten translation tasks with LLaMA-3-8B-Instruct and Mistral-Nemo-Instru. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have a high inference latency stemming from autoregressive decoding. |
| Approach: | They propose a novel decoding paradigm that drafts multiple tokens and verifies them in parallel . they aim to provide a catalyst for further research on Speculative Decoding . |
| Outcome: | The proposed method drafts multiple tokens and verifies them in parallel . it can be used to accelerate inference in large language models. |
Copied to clipboard
| Challenge: | Existing studies have suggested that the composition of the pretraining corpus exerts a significant impact upon the performance of LLMs. |
| Approach: | They analyze the impact of 48 datasets from 5 major categories of pretraining data of Large Language Models and measure their impacts on LLMs using benchmarks about nine major categories. |
| Outcome: | The proposed analysis provides insights into the organization of data to support more efficient pretraining of Large Language Models. |
Copied to clipboard
| Challenge: | Existing studies investigate ways to refuse to answer unknown questions . Large Language Models (LLMs) display a significant level of overconfidence when answering questions that they are aware of. |
| Approach: | They propose a self-alignment method to utilize Large Language Models to enhance its response-ability to unknown questions. |
| Outcome: | The proposed method is superior to baseline methods on four types of unknown questions. |
Copied to clipboard
| Challenge: | Existing data on suicidal ideation in private conversations are limited . a new dataset of 1,200 test cases is presented to address this gap . |
| Approach: | They propose a dataset of 1,200 test cases simulating implicit suicidal ideation in private contexts. |
| Outcome: | The proposed dataset includes 1,200 test cases simulating implicit suicidal ideation in dialogue scenarios. |
Copied to clipboard
| Challenge: | Existing Large Language Model (LLM) enabled agents lack flexibility to respond to users’ varying needs and preferences. |
| Approach: | They propose a test-time user-preference alignment strategy that optimizes the persona prompt, ensuring real-time preference alignment through textual loss feedback between simulated and ground-truth responses. |
| Outcome: | The proposed framework outperforms baseline methods in real-time and in real applications. |
Copied to clipboard
| Challenge: | Existing approaches to retrieval-agmented generation fail to generalize effectively in black-box scenarios. |
| Approach: | They propose a framework that leverages the semantic importance of words to dynamically adjust retrieval thresholds and filter information. |
| Outcome: | The proposed framework achieves the highest score on four long-form, knowledge-intensive generation datasets. |
Copied to clipboard
| Challenge: | Existing agent benchmarks neglect hierarchical rule application in real-world domains . a critical gap persists in numerous real-life professional domains where decision-making is governed by expert-written rules. |
| Approach: | They propose a benchmark requiring agents to assign a unique 10-digit Harmonized System (HS) Code to products by aligning their fuzzy attributes with strict tariff classification rules. |
| Outcome: | The proposed benchmarks lack hierarchical rule application capability in real-world domains . the proposed benchmark is based on e-commerce and is open-source . |
Copied to clipboard
| Challenge: | Existing approaches to learn multi-modal tasks are based on chain-of-thought . however, human thought processes are non-linear and employ dynamic adjustment and updating mechanisms. |
| Approach: | They propose a chain-of-thought technique that adjusts the length of the chain to improve the performance of generated prompts. |
| Outcome: | The proposed model improves multi-modal representation learning in visual, visual, and audio-visual tasks and also has good domain generalization performance due to better reasoning. |
Copied to clipboard
| Challenge: | Existing research has focused on evaluating detection methods for specific domains or language models. |
| Approach: | They build a testbed to detect texts from diverse human writings and LLMs using different detection methods. |
| Outcome: | Empirical results show that the top performing detector can identify 84.12% out-of-domain texts generated by a new LLM, indicating the feasibility for application scenarios. |
Copied to clipboard
| Challenge: | Existing studies have not explored aspect sentiment coherency, including its implications in adversarial defense. |
| Approach: | They propose a local sentiment aggregation paradigm that models aspect sentiment coherency . they demonstrate the capability of LSA in adversarial defense . |
| Outcome: | The proposed model outperforms existing models and achieves state-of-the-art sentiment classification performance. |
Copied to clipboard
| Challenge: | Low-rank adaptation methods for large language models have limitations in preserving world knowledge and limiting updates to preserve world knowledge. |
| Approach: | They propose a Fisher-optimized adaptive low Rank and Singular-VectorSelection framework for knowledge-preserving fine-tuning that allows efficient and task-sensitive updates. |
| Outcome: | The proposed framework outperforms existing methods for knowledge-preserving fine-tuning. |
Copied to clipboard
| Challenge: | Existing methods for text clustering use static pseudo-oracles, i.e., unidirectionally querying them for similarity assessment or data augmentation. |
| Approach: | They propose a training framework that enables bidirectional refinement between LLMs and embedding models by using task-aware prompts to guide the LLM in generating interpretations for the input texts. |
| Outcome: | Experiments on 14 benchmark datasets across 5 tasks demonstrate the effectiveness of the proposed training framework. |
Copied to clipboard
| Challenge: | Multi-modal information retrieval (MMIR) is a rapidly evolving field . current benchmarks for image-text pairings overlook the scientific domain . |
| Approach: | They develop a scientific domain-specific MMIR benchmark to evaluate image-text pairings using open-access research paper corpora. |
| Outcome: | The proposed benchmarks are based on 530K image-text pairs extracted from scientific documents with detailed captions. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable success across diverse tasks such as instruction following, code generation, and medical diagnosis. |
| Approach: | They propose a supervised fine-tuning-based auxiliary loss for Q-value estimations during supervised refinement. |
| Outcome: | The proposed method outperforms beam search on GSM8K, MATH, and GAOKAO on reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for named entity recognition ignore visual context bias . NER is a key component of many information extraction tasks . |
| Approach: | They propose to use a multimodal interaction module to generate word-aware visual representations and leverage purely text-based entity span detection as an auxiliary module to guide the final predictions. |
| Outcome: | The proposed approach achieves state-of-the-art on two benchmark datasets. |
Copied to clipboard
| Challenge: | Currently, LLMs learn in a data-driven schema while the instructions about complex tasks are both scarce and hard to collect or construct. |
| Approach: | They employ a gradient-based method to dissect the process that the Supervised Fine-tuning Process (SFT) adapts LLMs to downstream tasks via the perspective of attention patterns. |
| Outcome: | The proposed method dissects the process that the SFT process adapts LLMs to downstream tasks via the perspective of attention patterns. |
Copied to clipboard
| Challenge: | Using large-scale annotation data, large language models can generate noise, errors and biases, leading to unexpected behaviours. |
| Approach: | They propose a dataset to promote safety alignment in large language models . they separate helpfulness and harmlessness annotations for question-answering pairs . |
| Outcome: | The proposed dataset provides 44.6k prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels, with answers generated by Llama-family models. |
Copied to clipboard
| Challenge: | Existing large language models favor high-resource languages, such as English, at the expense of low-resourced and regional languages. |
| Approach: | They propose a series of language models that specifically focuses on Southeast Asian languages. |
| Outcome: | SeaLLM models outperform ChatGPT-3.5 in non-Latin languages by large margins . linguistic disparity impedes access to state-of-the-art AI technologies for non-English-speaking populations . |
Copied to clipboard
| Challenge: | Existing work on Aspect-based sentiment analysis ignores the rich label semantics of ABSA. |
| Approach: | They propose to tackle various ABSA tasks in a unified generative framework . they propose to use annotation-style and extraction-style modeling to enable training . |
| Outcome: | The proposed framework achieves state-of-the-art on four ABSA tasks across multiple benchmark datasets. |
Copied to clipboard
| Challenge: | Backdoor attacks are a new threat to neural natural language processing models due to the fragility and lack of interpretability of NLP models. |
| Approach: | They propose a method to perform backdoor attacks without an external trigger . they propose to use clean-labeled examples to generate poisoned clean-labelled examples . |
| Outcome: | The proposed strategy is effective and hard to defend due to its triggerless nature. |
Copied to clipboard
| Challenge: | Recent advances in training optimization for Transformer-based large language models lack systematic optimization of weight patterns during training. |
| Approach: | They propose a Weight Scaling method that rescales weights while preserving model outputs to improve model training efficiency and model quality. |
| Outcome: | The proposed method significantly improves convergence quality and loss reduction in LLMs with Grouped Query Attention architectures and LoRA fine-tuning tasks. |
Copied to clipboard
| Challenge: | Long-context modeling capabilities are important for large language models (LLMs) however, training LLMs with long context windows is insufficient since some samples do not exhibit strong semantic dependencies across long contexts. |
| Approach: | They propose a data mining framework ProLong that assigns each training sample with a long dependency score and ranks and filters them according to their results. |
| Outcome: | The proposed framework can rank and filter training samples that exhibit more powerful long-context modeling abilities. |
Copied to clipboard
| Challenge: | Existing tools for examining and fixing missing captions are lacking in mobile UIs. |
| Approach: | They propose a task for automatically generating language descriptions for UI elements from multimodal input including both the image and structural representations of user interfaces. |
| Outcome: | The proposed task can generate captions from image and structural representations of UI elements. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit severe hallucinations, which undermine reliability of automated scientific document understanding systems. |
| Approach: | They propose a framework for mitigating scientific measurement hallucinations through enhanced reasoning and targeted optimization. |
| Outcome: | The proposed framework significantly reduces hallucination rates and improves overall accuracy on the MeasEval benchmark. |
Copied to clipboard
| Challenge: | Using linguistic content and vocal characteristics for multimodal deep learning is difficult for computers to interpret human meaning . |
| Approach: | They propose a deep multimodal network with feature attention and modality attention to classify utterance-level speech data. |
| Outcome: | The proposed system achieves state-of-the-art or competitive results on three published multimodal datasets. |
Copied to clipboard
| Challenge: | Existing methods for debiasing may generate incorrect or nonsensical predictions but leave aside individual commonsense facts, resulting in modified knowledge that elicits unreasonable or undesired predictions. |
| Approach: | They propose a framework that identifies encoding locations of biases within language models and then applies the Fairness-Stamp (FAST) they also propose 'BiaScope' to evaluate the retention of commonsense knowledge and generalization across paraphrased social biase. |
| Outcome: | The proposed framework surpasses state-of-the-art baselines with superior debiasing performance while not compromising the overall model capability for knowledge retention and prediction. |
Copied to clipboard
| Challenge: | Social event detection relies on labeled data, but annotation is costly and labor-intensive. |
| Approach: | They propose a plug-and-play dual augmentation framework that combines explicit text-based and implicit feature-space augmentation to enhance data diversity and model robustness. |
| Outcome: | The proposed framework outperforms the best baseline model by 17.67% on the Twitter2012 dataset and 15.57% on the twitter2018 dataset in terms of the average F1 score. |
Copied to clipboard
| Challenge: | Existing LLM planning benchmarks emphasize local, step-level reasoning rather than global constrained optimization. |
| Approach: | They propose a benchmark for practical long-horizon agent planning that uses local constrained reasoning and global constrained optimization. |
| Outcome: | The proposed benchmarks show that even frontier agentic LLMs struggle with these problems. |
Copied to clipboard
| Challenge: | Recent work on Chain-of-Thought prompting imposes substantial computational overhead . lack of supervision obscures the analyzability of the latent reasoning chain. |
| Approach: | They propose a framework to render latent reasoning chain into images, making latent rationale explicit and traceable. |
| Outcome: | The proposed framework achieves 3-4 token compression and substantial inference acceleration compared to explicit CoT prompting. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have exhibited remarkable fluency across tasks, but their unethical applications are unclear. |
| Approach: | They propose a grammar error-free black-box attack that exploits LLM embeddings at the word-level while preserving original text quality. |
| Outcome: | The proposed attack compromises all detectors across domains and is transferable across source models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have excellent performance in many tasks, but they still face challenges in document translation. |
| Approach: | They propose a method that leverages the capabilities of Large Language Models to optimize document translation using only monolingual data. |
| Outcome: | The proposed method improves translation quality and improves contextual consistency in document translation using only monolingual data. |
Copied to clipboard
| Challenge: | Existing multilingual neural machine translation models perform poorly on language pairs with no parallel corpus. |
| Approach: | They propose a two-stage approach that encourages original models to acquire language-agnostic multilingual representations from new data and preserves the model architecture without introducing parameters. |
| Outcome: | The proposed approach improves performance in translation directions where existing models are weak and mitigates degeneration in the well-performing translation directions, offering flexibility in the real-world scenario. |
Copied to clipboard
| Challenge: | EASYTOOL combines tools from diverse tool documentation into a single tool instruction. |
| Approach: | They propose a framework that transforms tool documentation into a unified tool instruction. |
| Outcome: | EASYTOOL combines extensive tool documentation into a concise tool instruction . it reduces token consumption and improves performance of LLM-based agents . |
Copied to clipboard
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Copied to clipboard
| Challenge: | Existing methods to extract aspects and sentiments are limited due to lack of annotated sequence data. |
| Approach: | They propose a Selective Adversarial Learning method to align latent correlation vectors . they propose tagging a set of aspect boundary tags and sentiment tags to create a joint label space . |
| Outcome: | The proposed method can learn weights for words to achieve fine-grained adaptation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can solve complex tasks through iterative information retrieval. |
| Approach: | They propose a turn-level stage-aware policy optimization approach to solve this problem . they introduce a first-occurrence latent reward mechanism to allocate partial rewards . |
| Outcome: | Experiments show that TSPO outperforms state-of-the-art models on Qwen2.5-3B and 7B models. |
Copied to clipboard
| Challenge: | Existing studies fail to distinguish different classification errors with a standard cross-entropy classification loss and ignore the numbers in the fact description for predicting the term of penalty. |
| Approach: | They propose to extract crime amounts from fact description and use them to learn distinguishable representations to exploit the numbers in the fact description for predicting the term of penalty. |
| Outcome: | The proposed method achieves state-of-the-art results on real-world datasets and ablation studies demonstrate the effectiveness of each component. |
Copied to clipboard
| Challenge: | In order to perform downstream tasks, Large Language Models (LLMs) need continual adaptation without catastrophic forgetting. |
| Approach: | They propose a new paradigm that allows for continual adaptation without catastrophic forgetting . they propose to replay previous data based on task similarity with instructions . |
| Outcome: | The proposed method improves performance over 16 tasks with different training orders. |
Copied to clipboard
| Challenge: | TableLLM is a robust large language model capable of handling tabular data manipulation tasks. |
| Approach: | They propose a distant supervision method for training which includes a reasoning process extension strategy and a cross-way validation strategy. |
| Outcome: | The proposed model has 8 billion parameters and is capable of handling tabular data tasks. |
Copied to clipboard
| Challenge: | Existing topic models suffer from poor performance when applied to short text contents due to the limited length of a single topic. |
| Approach: | They propose a neural short text topic model that augments reconstruction labels with k-nearest documents to complement relevant but unobserved words. |
| Outcome: | The proposed model outperforms the state-of-the-art models on multiple public short-text datasets and can derive high-quality topics and document representations. |
Copied to clipboard
| Challenge: | Existing methods to transfer text style focus on sentence-level data, limiting performance . current LLMs struggle to generate public speaking texts that align with human preferences . |
| Approach: | They propose a task to transform official texts into public-speaking styles by analyzing real-world data. |
| Outcome: | The proposed task aims to transform public speaking texts into public-speaking styles . the proposed framework analyzes characteristics and identifies problems of stylized texts . |
Copied to clipboard
| Challenge: | AEGIS examines whether current models can effectively audit AI-generated images in academic papers. |
| Approach: | They propose a holistic benchmark for forensic analysis of AI-Generated academic ImageS that reveals limitations in academic image forensics. |
| Outcome: | AEGIS compared with existing benchmarks on seven academic categories and features key advances in forensic analysis. |
Copied to clipboard
| Challenge: | Existing methods for relation classification suffer from the scarcity of manually annotated data. |
| Approach: | They propose a novel relation classification model that incorporates query representation into the encoding of novel prototypes and utilizes iteratively to achieve more interaction. |
| Outcome: | The proposed model outperforms the state-of-the-art model on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on the repair generation capability of LLMs, lacking fine-grained evaluation of reflection. |
| Approach: | They propose a benchmark with oracle reflections and a dual-task protocol to decouple evaluation of reflection from repair. |
| Outcome: | The proposed benchmarks show that underperforming reflection capabilities remain a bottleneck for code repair. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) are predominantly trained on consistent visual-textual inputs, leaving open the question of whether they can handle semantic mismatches in layout-rich content. |
| Approach: | They propose to use multimodal inconsistency reasoning to assess MLLMs' ability to reason about semantic mismatches in webpages, presentation slides, and posters. |
| Outcome: | The proposed model outperforms open-source models in detecting inconsistencies in webpages, presentation slides, and posters while remaining vulnerable to inconsistent errors. |
Copied to clipboard
| Challenge: | Existing approaches to Multimodal Relation Extraction focus on single image scenarios . current approaches focus on text paired with a single image, ignoring valuable insights provided by remaining images. |
| Approach: | They propose a human-annotated dataset that includes multi-image and single-image instances for relation extraction. |
| Outcome: | The proposed model significantly improves relation extraction in multi-image scenarios. |
Copied to clipboard
| Challenge: | Existing symbolic music generation models represent musical notes as a sequence of attribute tokens with fixed unidirectional dependencies. |
| Approach: | They propose a symbolic music generation framework that adopts a autoregressive and a discrete diffusion architectures for note attributes. |
| Outcome: | The proposed framework improves state-of-the-art models across objective and subjective metrics. |
Copied to clipboard
| Challenge: | a new open-domain question answering system integrates best practices from IR with a BERT-based reader to identify answers from a large corpus of Wikipedia articles. |
| Approach: | They propose an end-to-end question answering system that integrates BERT with an IR reader. |
| Outcome: | The proposed system improves on a standard benchmark test collection. |
Copied to clipboard
| Challenge: | Extensive experiments on six real-world datasets show our approach outperforms the best baselines by 7.33% in NDCG@10, 4.65% in Recall@10 and 8.42% in MRR. |
| Approach: | They propose a framework for mapping sequential item texts to sequential item IDs that incorporates multi-query input and item linear projection to model conditional probability distribution of items. |
| Outcome: | The proposed framework outperforms baseline models on six real-world datasets by 7.33% and 4.65% respectively. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have achieved great success in various NLP tasks, but the vast model parameters pose challenges in downstream fine-tuning. |
| Approach: | They propose a task-agnostic prompting strategy that analyzes each dialogue utterance before task execution to enhance LLMs' comprehension in multi-turn dialogues. |
| Outcome: | The proposed strategy outperforms other zero-shot prompts and matches or exceeds efficacy of few-shot ones. |
Copied to clipboard
| Challenge: | Existing approaches for constituency parsing are expensive and lack high-quality labeled data. |
| Approach: | They propose a data augmentation method via lightweight large language model (LLM) generation and tree hybridization to generate a large number of structurally diverse instances. |
| Outcome: | The proposed method achieves significant improvements on five target domains with a lightweight LLM generation cost. |
Copied to clipboard
| Challenge: | Existing approaches to finding effective predictive signals from financial data are limited by their complexity and low signal-to-noise ratio. |
| Approach: | They propose a framework that combines code-level alpha representation with LLM-driven reasoning and evolutionary search. |
| Outcome: | The proposed framework combines code-level alpha representation with LLM-driven reasoning and evolutionary search. |
Copied to clipboard
| Challenge: | Speech Large Language Models (SpeechLLMs) have emerged as dominant speech processing approaches. |
| Approach: | They compare self-supervised learning-based discrete and continuous features . they compare performance across six spoken language understanding-related tasks . |
| Outcome: | The proposed models outperform discrete tokens and continuous features in six spoken language understanding-related tasks. |
Copied to clipboard
| Challenge: | Prompt-based learning can tackle zero-shot and few-shot NLP tasks . authors propose a method that makes use of pre-trained language models . |
| Approach: | They propose to map NLP tasks into natural language prompts, which are then filled by pre-trained language models. |
| Outcome: | The proposed method outperforms standard prompt-based methods in few-shot settings. |
Copied to clipboard
| Challenge: | Existing approaches to event detection require annotated triggers and event types in training data. |
| Approach: | They propose a framework that encodes the representation of a sentence based on target event types. |
| Outcome: | The proposed framework achieves competitive performances compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing reference-based metrics are limited by their reliance on human input. |
| Approach: | They propose to adapt some reference-based metrics to assess system summary against human-written references. |
| Outcome: | The proposed model outperforms reference-based metrics on two datasets and is comparable to reference-free metrics. |
Copied to clipboard
| Challenge: | Existing annotation tools do not consider post-annotation quality analysis due to inter-annotator disagreement. |
| Approach: | They propose a lightweight but efficient open-source tool for text span annotation that can be used for collaborative user annotation and administrator evaluation and analysis. |
| Outcome: | The proposed system reduces the annotation time by half compared with existing tools and the time can be compressed by 16.47% through intelligent recommendation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are capable of generating human-like text, but the potential for freely customisable characters remains underexplored. |
| Approach: | They propose a framework which employs Large Language Models to create freely customisable characters through personalised characteristic feature injection. |
| Outcome: | The proposed framework provides valuable insights for developing more accurate and customisable human simulacra. |
Copied to clipboard
| Challenge: | Existing studies link hallucination to data or representation biases, but their causal origins remain unclear. |
| Approach: | They propose a causal framework to analyze and mitigate hallucination in vision-language models by using counterfactual analysis to estimate the Natural Direct Effect (NDE) of each modality and their interaction. |
| Outcome: | The proposed framework significantly reduces hallucination while preserving task performance while retaining reliability. |
Copied to clipboard
| Challenge: | Existing automatic question generation methods focus on encoding passage and answer to generate question. |
| Approach: | They propose an automatic question generation approach which integrates question generation with its dual problem, question answering, into a unified primal-dual framework. |
| Outcome: | The proposed approach outperforms existing methods on SQuAD and HotpotQA benchmarks. |
Copied to clipboard
| Challenge: | Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence. |
| Approach: | They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included . |
| Outcome: | The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) demonstrate impressive reasoning ability and the maintenance of world knowledge in natural language tasks. |
| Approach: | They propose a framework that enables LLMs to ask relevant questions to uncover more details in the image, along with filters for refining the generated information. |
| Outcome: | The proposed framework boosts the performance of baseline methods by 2.15% on OK-VQA and achieves consistent improvements across different LLMs. |
Copied to clipboard
| Challenge: | Existing models for language analysis are inadequate for specialized domains like psychology. |
| Approach: | They have enriched a Chinese social media database with psychological lexicons to enhance its applicability to psychological text analysis. |
| Outcome: | The proposed model performed better on six public datasets and provided relevant predictions given the masked sentences. |
Copied to clipboard
| Challenge: | Existing low-rank adaptations have limited expressiveness, a tendency to overfit, and sensitivity to hyperparameter settings. |
| Approach: | They propose a technique to enhance LoRA’s expressiveness and generalization capabilities while preserving its training efficiency. |
| Outcome: | The proposed method outperforms baselines, mitigates overfitting, enhances model stability, and improves OOD robustness. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) have transformed image captioning . existing evaluations lack standardized criteria and a standardized evaluation framework . |
| Approach: | They propose a leaderboard for evaluating detailed captions that addresses three main gaps in existing evaluations: lack of standardized criteria, bias-aware assessments, and user preference considerations. |
| Outcome: | The proposed model evaluates caption quality, descriptiveness, risks, and societal biases while tailoring criteria to user preferences. |
Copied to clipboard
| Challenge: | Existing methods for dialog understanding only consider self-augmented dialogs as positive samples and treat all other dialogs like negative ones. |
| Approach: | They propose a tree-structured pre-trained conversation model which learns dialog representations from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised contrastive pre-training. |
| Outcome: | The proposed model can achieve state-of-the-art results on the DialoGLUE benchmark. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. |
| Approach: | They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences. |
| Outcome: | The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences. |
Copied to clipboard
| Challenge: | Existing methods for grammatical error correction (GEC) have been developed. |
| Approach: | They propose a method which integrates the detection labels from a Seq2Edit model to construct a template as the input. |
| Outcome: | The proposed method can perform human-in-the-loop error correction tasks. |
Copied to clipboard
| Challenge: | Existing methods for generating multiple translations for source and target languages neglect the one-to-many mapping between the source and the target languages. |
| Approach: | They propose a method to generate different translations for the input sentence by linearly interpolating it with different sentence pairs sampled from the training corpus during decoding. |
| Outcome: | Experiments on WMT’16 en-ro, WMT'14 en de, and WMT ‘17 zh-en show that the proposed method outperforms all previous diverse machine translation methods. |
Copied to clipboard
| Challenge: | Recent studies show that shallow semantic role labeling (SRL) performance drops under out-of-domain setting. |
| Approach: | They propose to annotate a multi-domain Chinese predicate-argument dataset using a frame-free annotation methodology and strict double annotation for improving data quality. |
| Outcome: | The proposed dataset is compared with a dataset from six different domains. |
Copied to clipboard
| Challenge: | Word sense disambiguation (WSD) is a fundamental task in natural language processing. |
| Approach: | They propose a training framework for generative WSD with chain-of-thought (CoT) and preference optimization. |
| Outcome: | The proposed framework achieves significant performance gains on rare and unseen settings and exhibits strong generalization in standard evaluation settings. |
Copied to clipboard
| Challenge: | Low-rank adaptation (LoRA) is an efficient approach for adapting large language models (LLMs) but many of the weights in these matrices are redundant, leading to inefficiencies in parameter utilization. |
| Approach: | They propose a low-rank adaptation approach that fine-tunes two low-ranked matrices and adapts them through a dense low-Rank matrix, improving parameter utilization and adaptation efficiency. |
| Outcome: | The proposed approach achieves 83.8% accuracy with only 0.01% of trainable parameters compared to LoRA's 80.8% with 0.70% of trainability parameters on LLaMA3-8B. |
Copied to clipboard
| Challenge: | Existing methods for enhancing cross-lingual transfer are limited by parallel resources and lack linguistic and domain coverage. |
| Approach: | They propose a cross-lingual in-context pre-training approach that leverages semantically related bilingual Wikipedia documents to enhance cross-linguistic transfer. |
| Outcome: | The proposed approach improves multilingual performance on three models across six target languages. |
Copied to clipboard
| Challenge: | Despite recent progress, most prior work studies confidence in single-turn question answering. |
| Approach: | They propose a logit-based probe that measures confidence in multi-turn dialogues . they propose 'infoECE' and a "hinter-guesser" paradigm for generating controlled evaluations based on data . |
| Outcome: | The proposed framework is grounded in calibration and monotonicity of confidence as more information becomes available. |
Copied to clipboard
| Challenge: | NER is one of the most important natural language processing tasks. |
| Approach: | They propose to annotate sentences from human-computer interaction, social media, and e-commerce using two rounds of annotation. |
| Outcome: | The proposed system performs the best on all the data sets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) lack structural information and semantic context to infer missing entities . large language models often lack structural signals to infuse missing entities into knowledge graphs . |
| Approach: | a modular framework integrates structural information and semantic context into a frozen LLM backbone for link prediction. |
| Outcome: | a new framework integrates KG-derived structural information and semantic context to infer missing entities. |
Copied to clipboard
| Challenge: | Prior work focused on typographic and pixel-level perturbations, leaving the study of SCO unexplored. |
| Approach: | They propose a framework that exploits MLLMs' diagrammatic reasoning capabilities to bypass safety guardrails. |
| Outcome: | The proposed framework exploits the model's reasoning capabilities to bypass safety guardrails. |
Copied to clipboard
| Challenge: | Prior studies modeled multimodal UI grounding in one round, but such an interaction is inherently iterative. |
| Approach: | They propose a task where a user and an agent collaborate on an interface screen . they use a dataset of 77,820 sequences of human user-agent interaction on mobile interfaces . |
| Outcome: | The proposed task improves the absolute task completion by 18% over the entire test set and 31% over the challenging split. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) are built across modalities and the misalignment between two modality can result in "hallucination" . developing LMMs faces challenges such as a lack of data and a limited number of data sets. |
| Approach: | They propose a new algorithm that augments the reward model with additional factual information such as image captions and ground-truth multi-choice options. |
| Outcome: | The proposed approach improves on the LLaVA-Bench dataset with the 96% performance level of the text-only GPT-4 and an improvement of 60% on MMHAL-BENCH over other baselines. |
Copied to clipboard
| Challenge: | Existing knowledge graphs lack robustness and incompleteness to provide link prediction. |
| Approach: | They propose to capture prior schema-level interactions related to relations by leveraging entity type information and introduce schema-guided negatives to bolster the efficiency of normal contrastive representation learning. |
| Outcome: | The proposed method achieves state-of-the-art performance on multiple established metrics across multiple datasets for link prediction. |
Copied to clipboard
| Challenge: | Existing methods for event causality identification (ECI) do not consider event causal label information and interaction information between event pairs. |
| Approach: | They propose a framework to enrich the representation of event pairs by introducing the event causal label information and the interaction information between event pairs. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Abstractive summarization models suffer from factual inconsistency problem . post-editing methods focus on replacing suspicious entities, failing to modify incorrect content hidden in sentence structures. |
| Approach: | They propose to use sentence pruning operation to correct possible errors . they propose to apply sentence pruning operations to the syntactic dependency tree . |
| Outcome: | The proposed method improves factual consistency on the FRANK dataset compared with baselines . it is model-independent and can serve as the final step in ensuring factual consistentness. |
Copied to clipboard
| Challenge: | Existing methods focusing on this task usually concatenate the concatened concepts words as the inputs of a pre-trained language model (PLM) however, in pre-training, the input is often corrupted sentences with correct word order. |
| Approach: | They propose a two-stage framework to improve the ability of pre-trained language models to deal with masked sentences with incorrect word order and a special token to make the input distribution more similar to the one used in pre-training. |
| Outcome: | The proposed method is able to generate a sentence containing all given concepts and correctly describe the relations between concepts. |
Copied to clipboard
| Challenge: | Existing word-level dependency parsing methods in Chinese lack explicit word boundaries due to the lack of word boundaries. |
| Approach: | They propose to model latent internal structures within Chinese words by constrained Eisner algorithm . they propose to guarantee a single root for intra-word structures and establish inter-word dependencies . |
| Outcome: | The proposed model outperforms existing models on Chinese treebanks and shows that it can predict plausible intra-word structures. |
Copied to clipboard
| Challenge: | Recent work finds that realizing who holds the initiative can help select knowledge . however, there is a strong semantic transition between two rounds, probably leading to initiative misjudgment . |
| Approach: | They propose a topic-shift Aware Knowledge sElector(TAKE) model which locates relevant parts from dialogue history to improve knowledge selection. |
| Outcome: | The proposed model outperforms baseline models on the WoW. |
Copied to clipboard
| Challenge: | Existing methods to improve pre-trained language models for many-class classification suffer from verbalizer ambiguity . a significant disparity exists between the pre-training and fine-tuning stages of the model . |
| Approach: | They propose a method to tune pre-trained language models to a broad spectrum of tasks . they use an instance-dependent soft prefix to complement language verbalizers in many-class classification . |
| Outcome: | The proposed method outperforms baselines on many-class datasets. |
Copied to clipboard
| Challenge: | Named Entity Recognition and Entity Linking are challenging for voice assistants . utterances are relatively short, so there is not much context to help disambiguate . |
| Approach: | They propose a Named Entity Understanding system that combines NER and EL in a joint reranking module. |
| Outcome: | The proposed framework improves NER accuracy by up to 3.13% and EL accuracy by 3.6% in F1 score . it also leads to better accuracies in other natural language understanding tasks . |
Copied to clipboard
| Challenge: | Existing tools for word segmentation are based on different linguistic theories or target different scenarios. |
| Approach: | They propose a probabilistic toolkit for multi-grained word segmentation in Chinese . they adopt semi-Markov CRF for single-grain word segmenting (SWS) . |
| Outcome: | The proposed approach can produce marginal probabilities of words during inference and significantly improve performance in the cross-domain scenario. |
Copied to clipboard
| Challenge: | Existing studies rely on deep graph neural networks (GNNs) to capture rich structural information, but they lack the structural information needed for QA. |
| Approach: | They propose a framework which captures structural information from KBs and models long-distance node relations from two perspectives. |
| Outcome: | The proposed framework models long-distance node relations from two perspectives . it is based on two widely used multi-hop KBQA datasets . |
Copied to clipboard
| Challenge: | Existing methods to extract event data are laborious to create and limited in size. |
| Approach: | They propose an event extraction model to overcome the roles overlap problem by separating the argument prediction in terms of roles. |
| Outcome: | The proposed method surpasses existing methods on the ACE2005 dataset and improves on the previous methods. |
Copied to clipboard
| Challenge: | Existing table benchmarks lack the capacity to adequately assess the practical application of table reasoning in industrial applications. |
| Approach: | They propose a bilingual table-to-report task and a table-based benchmark to assess the quality of table reasoning. |
| Outcome: | The proposed task is based on a bilingual benchmark with 457 industrial tables and evaluation criteria to measure the quality of report generation. |
Copied to clipboard
| Challenge: | Semantic phrases (SP) are lexical combinations whose meanings or usages may not be fully derived from their individual components. |
| Approach: | They propose to consolidate existing multiword expression resources into a unified testbed to assess language models in semantic phrase processing tasks. |
| Outcome: | The evaluation suite covers idiomatic expressions, noun compounds, and verbal constructions. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have redefined the role of AI in software engineering . current benchmarks focus on localized code generation, but neglect dynamic, full-process requirements of real-world engineering. |
| Approach: | They propose a benchmark to evaluate agentic backend coding within a realistic, executable workflow. |
| Outcome: | The ABC-Bench benchmark evaluates agentic backend coding within a realistic, executable workflow. |
Copied to clipboard
| Challenge: | Experimental results on ten datasets across seven domains demonstrate the effectiveness of PeerDA. |
| Approach: | They propose a new approach which uses span pairs with the PR relation as the augmentation data for training. |
| Outcome: | The proposed approach achieves state-of-the-art results on ten datasets across seven domains. |
Copied to clipboard
| Challenge: | Existing methods for procedural planning over-rely on visual inputs and lack structured semantic information. |
| Approach: | They propose a vision–language framework for multimodal procedural planning that exploits implicit spatial relations and deep semantics encoded in object attributes. |
| Outcome: | The proposed framework outperforms existing methods in terms of execution success rate, LCS, and planning correctness. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown superior performance without task-specific fine-tuning due to the computational costs. |
| Approach: | They propose a method which lets LLMs refer to the questions they have previously encountered and adaptively call for external resources when dealing with new questions. |
| Outcome: | The proposed method outperforms chain-of-thought based and fully retrieval-based methods on multiple datasets and outperformed chain- of-though, chatGPT and InstructGPT. |
Copied to clipboard
| Challenge: | Existing methods rely on external tool documentation during reasoning, leading to tool mastery difficulty, tool size constraints, and inference inefficiency. |
| Approach: | They propose a tool-internalized reasoning framework for unified reasoning and tool usage that integrates external tools into Large Language Models (LLMs) to address these issues, they propose 'tool-internet-based' reasoning. |
| Outcome: | The proposed method achieves superior performance across in-domain and out-of-domain settings, highlighting its effectiveness and efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) evolve into agentic systems capable of autonomous tool invocation and complex reasoning. |
| Approach: | They propose a trajectory-level preference benchmark to evaluate judges' ability to distinguish preferred versus distractor agent trajectories in tool-integrated environments. |
| Outcome: | The proposed benchmark evaluates how well judges distinguish preferred versus distractor agent trajectories in complex tool-using scenarios. |
Copied to clipboard
| Challenge: | Existing automated evaluation systems of chatbots rely on static chat scripts as ground truth, which is hard to obtain. |
| Approach: | They propose an interactive chatbot evaluation framework that allows chatbots to compete with each other like in a sports tournament. |
| Outcome: | The proposed framework can rank chatbots independently from their model architectures and domains . existing evaluation systems rely on static chat scripts as ground truth . |
Copied to clipboard
| Challenge: | a recent study has focused on how to recognize punchlines from dialogues, but has neglected character information. |
| Approach: | They propose a character-fusion conversational humor recognition model that uses character information to recognize punchlines from dialogue. |
| Outcome: | The proposed model improves performance on Chinese sitcoms corpus and punchline identification. |
Copied to clipboard
| Challenge: | Lightweight Vision-Language Models (VLMs) are indispensable for resource-constrained applications. |
| Approach: | They propose a framework that retrieves context from a memory bank to enhance alignment . they propose EMI-based approach to align vision and language models . |
| Outcome: | The proposed framework reduces training loss, accelerates convergence, and enhances task performance with negligible computational overhead. |
Copied to clipboard
| Challenge: | Existing knowledge enhancement techniques for pre-trained language models (PLMs) introduce noisy entity representations. |
| Approach: | They propose a knowledge enhancement filter that integrates external knowledge bases to enhance PLMs' ability to capture entity knowledge. |
| Outcome: | The proposed method achieves the highest F1-score and accuracy while reducing the computational cost by 1.7-2.5x. |
Copied to clipboard
| Challenge: | a study of slow reasoning models for multimodal reasoning finds that they are more prone to fabricating plausible yet false details when confronted with incomplete or misleading visual inputs. |
| Approach: | They conduct the first systematic study of the inverse scaling law in slow-thinking paradigms for multimodal reasoning. |
| Outcome: | The findings suggest that slower reasoning models are more prone to fabricating false details . the study analyzed 5,000-sample hierarchical prompt dataset by 50 participants . |
Copied to clipboard
| Challenge: | Existing methods to extract product attribute value require multiple extractions to obtain all corresponding values. |
| Approach: | They propose an Efficient product Attribute Value Extraction approach using lightweight sparse-layer interaction. |
| Outcome: | The proposed method achieves significant efficiency gains with neutral or marginal loss in performance when the context is long and number of attributes is large. |
Copied to clipboard
| Challenge: | Existing methods overlook the challenge of effectively transforming structure information from NL to SQL. |
| Approach: | They propose a text-to-SQL framework that unites content and structure pipes to bridge the gap between NL and SQL. |
| Outcome: | The proposed framework bridges the gap between natural language questions and SQL by combining content and structure pipes. |
Copied to clipboard
| Challenge: | Current approaches use a numeric ID or text piece as the identifier, but these identifieres cannot cover a passage’s content well. |
| Approach: | They propose a new type of identifier that is generated based on the content of a passage and could integrate contextualized information that text pieces lack. |
| Outcome: | The proposed approach performs the best in generative retrieval on three public datasets. |
Copied to clipboard
| Challenge: | Existing personalization methods rely on static personality modeling to achieve optimal performance. |
| Approach: | They propose a training-free framework for advanced situational personality steering that incorporates situation-dependent behavior patterns within LLM personalities through analysis of persona neurons. |
| Outcome: | The proposed framework surpasses baselines on PersonalityBench and SPBench, demonstrating generalization and robustness to complex, unseen situations and different models architecture. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs)-based agents have fundamentally reshaped artificial intelligence . however, the inherent statelessness of LLMs hinders their ability to maintain logical consistency across complex, multi-step tasks . |
| Approach: | They propose a framework for LLM agent memory mechanisms that formalizes the development process into three stages: storage, reflection, and experience. |
| Outcome: | The proposed framework breaks the development process into three stages . it analyzes the need for long-range consistency, challenges in dynamic environments, and the ultimate goal of continual learning. |
Copied to clipboard
| Challenge: | Existing studies focus on detecting known and previously undefined categories of user intent . skewed and long-tailed distributions often encountered in open-world scenarios . |
| Approach: | They propose to use imbalanced new intent discovery task to identify familiar and novel intent categories within long-tailed distributions. |
| Outcome: | The proposed model outperforms the existing benchmark on three datasets to simulate the real-world long-tail distributions. |
Copied to clipboard
| Challenge: | Chinese spelling check (CSC) is a fundamental NLP task that detects and corrects spelling errors in Chinese texts. |
| Approach: | They propose an auxiliary task of Chinese pronunciation prediction to improve CSC . they propose adaptive weighting schemes and a delicate correction strategy . |
| Outcome: | The proposed auxiliary task improves Chinese pronunciation prediction on three benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for zero-shot stance detection are labor-intensive to train for each new target. |
| Approach: | They propose a generative data augmentation approach to generate training samples containing unseen and seen targets and map them into the same embedding space with contrastive learning. |
| Outcome: | The proposed model achieves state-of-the-art on most topics in the task of zero-shot stance detection. |
Copied to clipboard
| Challenge: | Xu and Peng, 2025) . . SPUR is a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images. |
| Approach: | They propose to use 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images to evaluate the visual perception of multimodal large language models (MLLMs) . they also propose to utilize cross-panel relation understanding to evaluate MLLM’s ability to decipher intricate cross-panel relations. |
| Outcome: | The proposed model is based on 4,264 question-answering pairs derived from 1,084 expert-curated images. |
Copied to clipboard
| Challenge: | Existing methods for ambiguous queries struggle to retrieve high-quality documents . DFAMS outperforms advanced FR methods by 14.37% in knowledge classification accuracy . |
| Approach: | They propose a framework that leverages dynamic information flow to identify latent query intents and construct semantically aligned knowledge partitions for accurate retrieval across heterogeneous sources. |
| Outcome: | The proposed framework outperforms existing methods in classification accuracy and retrieval recall tests. |
Copied to clipboard
| Challenge: | Large language models produce content lacking pedagogical depth when asked to generate lessons . |
| Approach: | They propose a framework that allows teachers to select content according to pedagogical intent and sequence topics so foundations precede applications. |
| Outcome: | The framework achieves 67.8% win rate in human evaluation and 79.6% in LLM-based evaluation against eight baselines. |
Copied to clipboard
| Challenge: | Large language models have advantages over neural machine translation systems, but they suffer from high computational costs and significant latency. |
| Approach: | They propose a scheduling policy that optimizes translation result while ensuring fast speed and as little LLM usage as possible. |
| Outcome: | The proposed model achieves optimal translation performance with less LLM usage on multilingual test sets. |
Copied to clipboard
| Challenge: | Existing methods for generating stories with complex plots rely on detailed prompts, which inadvertently limit the creative potential of the generated stories. |
| Approach: | They propose a retrieval-auGmented stoRy generation framework with a fOrest of eVidEnce to enhance stories’ complexity. |
| Outcome: | The proposed framework enables generating more diverse plotlines from human-written stories. |
Copied to clipboard
| Challenge: | Root cause analysis (RCA) in Micro-services architectures with escalating complexity is challenging due to fault propagation and circular dependencies among nodes. |
| Approach: | They propose a framework where multiple agents follow Agent Workflow and collaborate in blockchain-inspired voting to ensure the reliability of root cause analysis. |
| Outcome: | The proposed framework reduces the number of steps and standardizes task processing through Agent Workflow. |
Copied to clipboard
| Challenge: | EvoRoute is a self-evolving model routing paradigm that transcends static, pre-defined model assignments. |
| Approach: | They propose a model routing paradigm that transcends static, pre-defined model assignments. |
| Outcome: | Experiments on GAIA and BrowseComp+ show that EvoRoute reduces execution cost and latency by over 70%. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual natural language inference lack annotated parallel corpora. |
| Approach: | They propose a new prompt learning framework with the Multilingual Verbalizer for XNLI that uses a multilingual verbalizer to align the representations of original and augmented multilingual questions into a unified semantic space with consistency regularization. |
| Outcome: | The proposed framework outperforms existing methods under few-shot and full-shot cross-lingual transfer settings. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have remarkable capabilities in understanding complex tasks, but they can only handle graph partitioning tasks that require global perception abilities. |
| Approach: | They propose a pipeline for coarsening, reasoning, and refining to enable LLMs to perform graph partitioning on small-scale graphs. |
| Outcome: | The proposed pipeline can handle graph partitioning tasks on small graphs with coarsening, reasoning, and refining. |
Copied to clipboard
| Challenge: | Existing systems for automatic poetry generation are model-oriented, resulting in poor user participation. |
| Approach: | They propose a human-machine collaborative Chinese classical poetry generation system called Jiuge . Jiuge allows users to revise unsatisfied parts of a generated poem draft repeatedly . |
| Outcome: | The proposed system allows users to revise unsatisfied parts of a generated poem draft repeatedly. |
Copied to clipboard
| Challenge: | Pretrained language models are often finetuned for downstream tasks, which has been shown to improve performance over non-pretrained models. |
| Approach: | They propose a genetic algorithm to automatically search for the best prompt for few-shot learning with pretrained language models by gradient-free algorithm. |
| Outcome: | Experiments on diverse datasets show that the proposed method outperforms manual prompts by 2.6 points. |
Copied to clipboard
| Challenge: | Existing language models have limited sensitivity to temporal information and inadequate temporal reasoning capabilities. |
| Approach: | They propose a framework that enhances temporal awareness and reasoning . they propose to use Temporal Information-Aware Embedding and Granular Contrastive Reinforcement Learning . |
| Outcome: | The proposed framework outperforms existing LLMs on time-sensitive question answering tasks. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed model improves naturalness and prosody diversity with clear margins. |
| Approach: | They propose a cross-utterance conditional VAE to estimate posterior probability distribution of latent prosody features for each phoneme by conditioning on acoustic features, speaker information, and text features from past and future sentences. |
| Outcome: | The proposed model improves naturalness and prosody diversity with clear margins. |
Copied to clipboard
| Challenge: | Recent advances in vision-language learning have significantly advanced Human-Computer Interactions (HCI). |
| Approach: | They propose a method to align the semantic spaces between speech and text by incorporating two modules to align semantic spaces. |
| Outcome: | The proposed method outperforms state-of-the-art approaches on AVOS benchmarks. |
Copied to clipboard
| Challenge: | 3D visual grounding aims to localize the desired objects in a 3D point cloud by a free-form language description. |
| Approach: | They propose a relation-aware framework which captures relative spatial relationships between objects and enhances object attributes. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three benchmarks . it captures relative spatial relationships between objects and enhances object attributes . |
Copied to clipboard
| Challenge: | Existing studies focus on dummy tokens but fail to leverage the inherent sentence-level structure of natural language. |
| Approach: | They propose a method that inserts delimiters at sentence boundaries to enhance large language models' capabilities. |
| Outcome: | The proposed method improves performance on 7B LLMs to 600B Deepseek-V3 with 7.7% gains on GSM8k and 12.5% on DROP. |
Copied to clipboard
| Challenge: | Large language models struggle to meet user’s needs when required to generate responses of a specific length due to their inherent difficulty in accurately perceiving numerical constraints. |
| Approach: | They propose a Target Length Generation Task and propose RULER, a model-agnostic approach that controls generated length for large language models. |
| Outcome: | The proposed model-agnostic approach improves instruction-following ability of large language models under length-constrained instructions and can generate appropriate MLT when length constraints are not explicitly provided. |
Copied to clipboard
| Challenge: | Recent works require a model to learn from trace examples of a task via supervised learning or few/many-shot prompting. |
| Approach: | They propose a model that iteratively learns from its mistakes via self-reflection and structured thought management. |
| Outcome: | The proposed model outperforms previous models on easy tasks with more efficient reasoning and self-reflection. |
Copied to clipboard
| Challenge: | Large language models excel at processing unstructured data, but integrating time series data with text remains a challenge. |
| Approach: | They propose a self-supervised multimodal framework that uses prompt-guided learning to unify heterogeneous data types. |
| Outcome: | The proposed framework outperforms state-of-the-art approaches on disease diagnosis tasks using real-world datasets. |
Copied to clipboard
| Challenge: | Modern NLP workflows require different models for generation and embedding tasks. |
| Approach: | They propose a method that transforms an LLM into a Uni-Directional Masked Auto-Encoder. |
| Outcome: | The proposed method achieves state-of-the-art under unsupervised conditions with merely 100 training steps. |
Copied to clipboard
| Challenge: | Existing work generates long videos segment by segment sequentially, which is inefficient. |
| Approach: | They propose a Diffusion over Difference architecture for eXtremely Long video generation. |
| Outcome: | The proposed architecture reduces the average inference time from 7.55min to 26s (94.26%) and generates high-quality long videos with both global and local coherence. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models can perform well in general tasks, but their effectiveness and limitations in domainspecific tasks remain unclear. |
| Approach: | They examine the proficiency of Large Language Models (LLMs) in generating succinct survey articles specific to the niche field of NLP in computer science. |
| Outcome: | The LLMs perform better in generating succinct survey articles specific to the niche field of NLP in computer science, compared to human-authored surveys, but they exhibit bias in evaluation. |
Copied to clipboard
| Challenge: | Experimental results show that our method outperforms full-model tuning in few-shot abstractive summarization tasks. |
| Approach: | They propose a soft prompts architecture with prompt pre-training and prompt fine-tuning paradigm to support few-shot abstractive summarization. |
| Outcome: | The proposed model outperforms Prompt Tuning and Profix-Tuning on CNN/DailyMail and XSum datasets and outperfies Profix Tuning by a large margin. |
Copied to clipboard
| Challenge: | Existing benchmarks for table reasoning are incomplete due to the complexity of the tables and user questions in real-world applications. |
| Approach: | They propose a Multi-scale spreadsheet benchmark with Meta operations for Table reasoning that incorporates two key features and a new criterion with six categories of meta operations for measuring the difficulty of each question. |
| Outcome: | The proposed model outperforms Claude-3.5-Sonnet with 77.4% accuracy on the existing benchmarks. |
Copied to clipboard
| Challenge: | a new problem of grounding natural language instructions to mobile UI actions is emerging . we use a Transformer to extract action phrase tuples from long-range natural language instruction . |
| Approach: | They propose a dataset that pairs English instructions with actions performed by people on a mobile UI emulator. |
| Outcome: | The proposed model achieves 70.59% accuracy on predicting complete ground-truth action sequences in PixelHelp. |
Copied to clipboard
| Challenge: | Large language models have significantly advanced Multilingual Machine Translation (MMT) yet scaling to many languages while maintaining robust performance across directions remains challenging. |
| Approach: | They propose a strategy to reduce the number of translations in one direction . they propose auxiliary parallel sentences to promote cross-lingual transfer . |
| Outcome: | The proposed model performs on par with or better than substantially larger baselines. |
Copied to clipboard
| Challenge: | MM-Verifier and MM Reasoner are a powerful multimodal reasoning model . large language models (LLMs) have demonstrated exceptional performance across tasks spanning myriad domains. |
| Approach: | They propose a method which combines tree search and verification to generate high-quality chain-of-thought data. |
| Outcome: | The proposed method outperforms all larger models on the MathCheck, MathVista, and MathVerse benchmarks. |
Copied to clipboard
| Challenge: | Contextualised word embeddings are powerful tool to detect contextual synonyms, but most of the current SOTA methods are supervised and underexploit the potential of the context. |
| Approach: | They propose a self-supervised approach which detects contextual synonyms of concepts being trained on the data created by shallow matching. |
| Outcome: | The proposed approach outperforms the previous SOTA with gains of up to 4.5 and 4.0 absolute points after fine-tuning with as little as 20% of the labelled data. |
Copied to clipboard
| Challenge: | Advanced GUI agents suffer from prohibitive deployment costs on resource-constrained devices. |
| Approach: | They propose a lightweight GUI agent with GUI-specific knowledge and task scalability . LAMO-3B supports monolithic execution and MAS-style orchestration . |
| Outcome: | The proposed GUI agent LAMO-3B supports monolithic execution and MAS-style orchestration. |
Copied to clipboard
| Challenge: | a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios is a major challenge for end-to end spoken dialogue models. |
| Approach: | They propose to provide an extensive evaluation framework for end-to-end spoken dialogue models (SDMs) that includes both cognitive dimensions and paralinguistic cues . |
| Outcome: | The proposed benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model’s abilities in U**nderstanding, **R**easoning, and **O**ral conversation. |
Copied to clipboard
| Challenge: | Multi-step reasoning remains challenging for language models with limited capacity . et al., 2025) demonstrate remarkable reasoning capabilities across diverse tasks . |
| Approach: | They propose a module-aware reasoning distillation framework that explicitly targets key Transformer components for effective reasoning transfer. |
| Outcome: | The proposed framework targets key components for effective reasoning transfer . it adopts an offline distillation setting, where a strong teacher model provides reasoning trajectories in advance . |
Copied to clipboard
| Challenge: | Annotated data plays a critical role in training models and evaluating their performance. |
| Approach: | They propose a paradigm for Human-LLM co-annotation of unstructured texts at scale that utilizes uncertainty to estimate LLMs’ annotation capability. |
| Outcome: | The proposed model outperforms existing models on many text-annotation tasks with up to 21% performance improvement over random baseline. |
Copied to clipboard
| Challenge: | Existing methods to update a multilingual model with new language pairs are expensive and time-consuming. |
| Approach: | They propose an entropy-based vocabulary substitution method that walks through new language pairs for incremental learning while remaining the size of the original vocabulary. |
| Outcome: | The proposed method achieves better performance and saves excess overhead in a multilingual machine translation task. |
Copied to clipboard
| Challenge: | Large language models (LLMs) acquire substantial world knowledge during pretraining, which is further shaped by post-training techniques such as supervised fine-tuning (SFT). |
| Approach: | They evaluate closed-book question answering (CBQA) performance across five LLMs from the LLaMA-2 and LLama-3 families and examine the impact of supervised fine-tuning on model knowledge. |
| Outcome: | The proposed model performance is 14% worse than models fine-tuned on 1,920 samples and 12% worse on 240 samples. |
Copied to clipboard
| Challenge: | RAG implementations face challenges in addressing retrieved noise and redundant content . current RAG methods lack the ability to exploit fine-grained inter-document relationships . |
| Approach: | They propose a retrieval-augmented generation framework that exploits latent inter-document relationships while removing irrelevant information and redundant content. |
| Outcome: | The proposed framework achieves consistent performance improvements on knowledge-QA and hallucination-Detection datasets. |
Copied to clipboard
| Challenge: | Existing code debugging benchmarks focus on the Code Repair stage of the code generation process. |
| Approach: | They propose a framework to evaluate the debugging abilities of large language models by emulating the human debug process. |
| Outcome: | The proposed framework outperforms human-curated and GPT-4-generated training data, enabling 7B-scale LLMs to achieve comparable debugging performance to GPT-3.5. |
Copied to clipboard
| Challenge: | Entity typing fails to assign an entity to the types beyond the predefined type set. |
| Approach: | They propose a generative entity typing paradigm that assigns types to entities . traditional classification-based approaches fail to assign entities to the types beyond the predefined set . they employ curriculum learning to train the model on heterogeneous data . |
| Outcome: | The proposed model outperforms the state-of-the-art model on heterogeneous training data. |
Copied to clipboard
| Challenge: | Large language models (LLMs) may exhibit undesirable behaviors due to the inevitable biases and harmful content present in training. |
| Approach: | They propose to investigate the elasticity of large language models by examining their performance. |
| Outcome: | The proposed model performance declines rapidly before reverting to the pre-training distribution, the authors show . the proposed model weight and code are available at pku-lm-res ist-alignment.github.io. |
Copied to clipboard
| Challenge: | Evaluation and Management (E/M) coding is performed by physicians and trained human coders who review clinical encounter notes and electronic health record data to assign appropriate codes. |
| Approach: | They propose a framework that automates evaluation and management coding tasks using the Current Procedural Terminology (CPT) taxonomy. |
| Outcome: | The proposed framework achieves an increase in coding accuracy of more than 36% over a commercial CPT E/M coding system and almost 5% over our strongest single-prompt baseline. |
Copied to clipboard
| Challenge: | Existing large language models exhibit unidirectional behavior when processing bidirectional relationships . authors propose a solution to alleviate the reversal curse in Diffusion LLMs . |
| Approach: | They propose a model that addresses the "reversal curse" of bidirectional behavior in large language models . they propose 'entity-aware training' and balanced data construction to alleviate asymmetry and missing relations . |
| Outcome: | The proposed model alleviates the "reversal curse" in Diffusion LLMs . the proposed model employs whole-entity masking to mitigate entity fragmentation . |
Copied to clipboard
| Challenge: | Existing methods for hallucination detection are coarse-grained and lack long-range consistency checks. |
| Approach: | They propose a benchmark for long-form hallucination detection that incorporates diverse entity types and intricate factual dependencies spanning extended contexts. |
| Outcome: | The proposed framework outperforms baselines and robustly integrates fact-centric hyper-relational knowledge graphs. |
Copied to clipboard
| Challenge: | a new generation of (M)LLMs is enabling the creation of superintelligent AI assistants . OS Agents can complete tasks autonomously and have the potential to significantly enhance the lives of billions of users worldwide. |
| Approach: | They propose to build OS Agents that operate within operating systems' GUIs and GUIs . they examine evaluation metrics and benchmarks to identify promising directions . |
| Outcome: | The proposed agents are based on operating systems (OS) and operating systems frameworks. |
Copied to clipboard
| Challenge: | Existing studies have revealed the robustness degra-dation caused by data distillation. |
| Approach: | They propose a framework to evaluate and quantify model distillation . they aim to identify identity cognition contradictions and analyse multi-granularity response similarities across models to measure the extent of homogenization. |
| Outcome: | The proposed framework addresses two key aspects: (1) Identifying identity cognition contradictions to assess discrepancies in how models perceive and represent identity-related information; (2) Analyzing multi-granularity response similarities across models to measure the extent of homogenization. |
Copied to clipboard
| Challenge: | a new framework for complex reasoning with LLMs is developed to improve reasoning proof accuracy and interpretability. |
| Approach: | They propose to use LLMs to generate search logs that can be interpreted into human-readable reasoning proofs. |
| Outcome: | The proposed framework improves reasoning accuracy but lacks interpretability due to black-box nature of the solvers. |
Copied to clipboard
| Challenge: | a recent study shows that agent research practices are far from standard, rigorous . lack of a standard evaluation protocol makes previous works not reproducible, authors say . |
| Approach: | They conduct an empirical study on the GAIA benchmark to investigate agent design choices . they find that lack of a standard evaluation protocol makes previous works not reproducible . |
| Outcome: | The proposed framework achieves state-of-the-art performance among open-source projects. |
Copied to clipboard
| Challenge: | Weakly supervised vision-and-language pre-training (WVLP) uses only local descriptions of images as cross-modal anchors to construct weakly-aligned image-text pairs for pre- training. |
| Approach: | They propose to take a small number of aligned image-text pairs as anchors and represent each unaligned image and text by its similarities to these anchors. |
| Outcome: | The proposed model reduces the cost of pre-training while maintaining decent performance on downstream tasks. |
Copied to clipboard
| Challenge: | Existing embedding models support only 512 input tokens, hindering their application in scenarios requiring long inputs. |
| Approach: | They evaluate the performance of existing embedding models by using a new benchmark and a training-free context window extension strategy. |
| Outcome: | The proposed model extends the input window of existing models by several folds. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) has been gaining popularity in high-resource translation tasks, but struggles in low-ressource and morphologically-rich scenarios. |
| Approach: | They propose a multi-task neural model that jointly learns to perform bi-directional translation and agglutinative language stemming. |
| Outcome: | The proposed model can significantly improve translation performance on agglutinative languages by using a small amount of monolingual data. |
Copied to clipboard
| Challenge: | Recent studies have proposed tool learning, which augments LLMs with external tools. |
| Approach: | They propose an adaptive and hierarchy-aware reranking method to refine retrieval results by truncating the retrieval result related to seen and unseen tools at different positions. |
| Outcome: | The proposed method improves retrieval results, leading to better execution results generated by the LLM. |
Copied to clipboard
| Challenge: | Existing methods to learn speech representations for end-to-end speech-totext translation (ST) neglect the representation discrepancy across modalities. |
| Approach: | They propose a method to calibrate the representation discrepancy between modalities by mixing up the representation sequences of different modality inputs. |
| Outcome: | The proposed method alleviates the cross-modal representation discrepancy and improves on a strong baseline on eight translation directions. |
Copied to clipboard
| Challenge: | Group Relative Policy Optimization (GRPO) uses a coarse-grained credit assignment mechanism that propagates group-level rewards uniformly to to every token in a sequence, neglecting the varying contribution of individual reasoning steps. |
| Approach: | They introduce Outcome-grounded Advantage Reshaping (OAR) which redistributes advantages based on how much each token influences the model’s final answer. |
| Outcome: | Empirical results show that OAR-G outperforms GRPO on a high-fidelity attribution signal and suppresses low-impact tokens while preserving the advantage mass. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive performance in various natural language processing tasks. |
| Approach: | They propose a benchmark for the evaluation of large language models in the IP domain . they also propose supervised multilingual large language model called MoZi . |
| Outcome: | The proposed model outperforms four well-known LLMs on the MoZIP benchmark . the most powerful ChatGPT does not reach the passing level . |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models are constrained to datasets where each sample is manually injected with only one type of bias. |
| Approach: | They propose a multi-bias benchmark where each sample contains multiple types of biases. |
| Outcome: | The proposed benchmark shows that existing LLMs and debiasing methods perform poorly on this benchmark, highlighting the challenge of eliminating compounded biases. |
Copied to clipboard
| Challenge: | Existing studies of stance detection focus on learning stance information about specific targets from context, but in real-world scenarios, we usually have a certain understanding of a target when we express our stance on it. |
| Approach: | They propose to take the background knowledge of the target into account for better stance detection by categorizing it into episodic and discourse knowledge categories and a heuristic retrieval algorithm based on the topic to retrieve the Wikipedia documents relevant to the sample. |
| Outcome: | The proposed framework achieves state-of-the-art on four benchmark datasets showing that the proposed framework is able to detect stances in-target and zero-shot scenarios. |
Copied to clipboard
| Challenge: | Word sense disambiguation (WSD) is a fundamental yet challenging task in natural language processing. |
| Approach: | a novel multi-agent Debate framework for adversarial word Sense disambiguation is proposed . the framework simulates a real-world debate environment where multiple agents engage in discussions about ambiguous words in the context of adversarials. |
| Outcome: | The proposed framework integrates with existing LLMs and improves models in Chinese language . it shows that it can be used to improve models in the Chinese language and improve performance . |
Copied to clipboard
| Challenge: | Knowledge graph completion (KGC) is a critical task to predict missing facts among entities. |
| Approach: | They propose a knowledge-constrained generative re-ranking method based on generative large language models for KGC that can predict missing facts among entities. |
| Outcome: | The proposed method achieves state-of-the-art performance on four datasets and 9.0% and 11.1% compared to the previous methods. |
Copied to clipboard
| Challenge: | Existing studies on document-level neural machine translation (NMT) assume that all repeated source words should be translated consistently. |
| Approach: | They construct a test set of 310 bilingual news articles to evaluate lexical translation consistency. |
| Outcome: | The proposed test sets show that translation consistency is consistent across multiple languages. |
Copied to clipboard
| Challenge: | Existing approaches treat length as an incidental output property rather than a statistically regular phenomenon worthy of rigorous modeling. |
| Approach: | They propose a statistical framework for modeling and controlling large language model response lengths using extreme value theory and cross-validation on Qwen and DeepSeek architectures. |
| Outcome: | The proposed model improves tail fit and generalizability while maintaining generalizzability. |
Copied to clipboard
| Challenge: | Existing unlearning methods suffer from a geometric mismatch, causing catastrophic forgetting or unsafe substitution. |
| Approach: | They propose a framework for surgical semantic pruning within the Lorentz manifold. |
| Outcome: | Experiments on MLLMU-Bench show that LOTUS significantly outperforms baselines while maintaining general utility. |
Copied to clipboard
| Challenge: | Existing methods for generating responses following a desired style are lacking of parallel data for training. |
| Approach: | They propose a KL loss and a style classifier to fine-tune response generation . they show that their model can significantly outperform state-of-the-art methods . |
| Outcome: | The proposed model outperforms state-of-the-art models in style consistency and contextual coherence with two public datasets. |
Copied to clipboard
| Challenge: | Existing VQA models suffer from language bias that indicates a spurious correlation between textual questions and answers. |
| Approach: | They propose a model agnostic dual-debiasing framework that models two types of language bias by separate branches under counterfactual inference framework. |
| Outcome: | The proposed framework significantly reduces language bias and achieves state-of-the-art performance on the benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for embedding knowledge graphs are difficult due to complicated query structures and incomplete graph data. |
| Approach: | They propose a probabilistic embedding model for encoding entities and queries to answer different types of FOL queries on KGs. |
| Outcome: | The proposed model outperforms state-of-the-art models on public benchmarks on three large logical query datasets. |
Copied to clipboard
| Challenge: | Knowledge Distillation (KD) is a predominant approach for BERT compression. |
| Approach: | They propose a weight-inherited distillation method which directly transfers knowledge from the teacher to a compact student model by inheriting the weights. |
| Outcome: | The proposed method outperforms state-of-the-art KD-based methods on GLUE and SQUAD benchmarks. |
Copied to clipboard
| Challenge: | Multi-turn, long-horizon tasks require dozens of sequential model calls per episode. |
| Approach: | They propose a cost-aware multi-turn LLM routing tool which encodes interaction history and candidate models into joint history–model embeddings and learns an outcome estimator from logged trajectories to predict turn-level model utility. |
| Outcome: | The proposed model reduces cost and performance by 58.7% on ScienceWorld and on Humanity’s Last Exam (HLE) and even reduces costs for held-out tasks. |
Copied to clipboard
| Challenge: | Existing methods for movie dubbing break phonemes in scripts, resulting in incomplete phoneme pronunciation and poor identity stability. |
| Approach: | They propose a method that switches dubbing learning from frame level to phoneme level . it uses a multimodal style adaptor to learn pronunciation style from audio . |
| Outcome: | The proposed method improves on two benchmarks, V2C and Grid, and is available on github. |
Copied to clipboard
| Challenge: | Existing tasks to generate question-answer pairs from visual images are under-explored. |
| Approach: | They propose a task that targets question-answer pair generation from visual images. |
| Outcome: | The proposed model can generate diverse or consistent QAPs on two benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for fraud detection rely on transcribed text, lacking acoustic cues . a proposed framework for audio-based slow-thinking fraud detection eliminates transcription errors . |
| Approach: | They propose a framework for audio-based slow-thinking fraud detection that eliminates transcription errors and rewards slow-thought reasoning by capturing fine-grained audio details. |
| Outcome: | The proposed method improves accuracy, inference efficiency, and real-time processing capabilities. |
Copied to clipboard
| Challenge: | Maximum likelihood estimation (MLE) is used to train models, but during testing, the model is conditioned on previously generated tokens, resulting in exposure bias. |
| Approach: | They propose to use optimal transport to match the sequences generated in MLE and test modes to reduce exposure bias. |
| Outcome: | The proposed method is validated on machine translation, text summarization, and text generation tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on single agentic capability, failing to capture long-horizon real-world scenarios. |
| Approach: | They propose a benchmark that evaluates 6 agentic capabilities across 32 real-world scenarios. |
| Outcome: | Experiments show that closed-source models outperform open-source model (48.4% vs 32.1%) integrating models with advanced scaffolds to form autonomous agents is a paradigm shift. |
Copied to clipboard
| Challenge: | e-commerce companies often have the option of escalating complaints by filing grievances with a government authority . this is detrimental to an ecommerce company, but this problem is challenging to solve by integrating recurrent neural networks with manually-engineered features. |
| Approach: | They propose a model that integrates recurrent neural networks with manually-engineered features to identify cases where the customer expresses such an intent. |
| Outcome: | The proposed model outperforms baseline models and provides better recall and triage for specialized agents. |
Copied to clipboard
| Challenge: | Existing methods for creating video content are limited by high costs and slow update cycles. |
| Approach: | They propose a paradigm shifting educators from manual creators to high-level directors who focus on pedagogical intents while agents handle execution. |
| Outcome: | The proposed framework reduces production costs to 0.3% of traditional course videos and provides a robust solution for scalable education. |
Copied to clipboard
| Challenge: | Existing approaches to argument summarization rely on single-pass generation, offering limited support for factual correction or structural refinement. |
| Approach: | They propose a large language diffusion framework that iteratively improves argument summarization by sufficiency-guided remasking and regeneration. |
| Outcome: | Empirical results show that Arg-LLaDA surpasses state-of-the-art baselines in 7 out of 10 evaluation metrics. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can solve complex multi-step math reasoning problems, but their internal implementation is limited. |
| Approach: | They propose to use a "C**ausal **E**ffect **D**riven **F**ine-tuning method" to improve LLMs' reasoning ability. |
| Outcome: | The proposed method improves the model's reasoning ability by enhancing key components that are used to execute mixed arithmetic calculations. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to perceive the visual world with a deep concept structure cognition. |
| Approach: | They propose a concept-level benchmark to assess MLLMs’ hierarchical concept understanding and reasoning abilities. |
| Outcome: | The proposed model outperforms state-of-the-art models in concept structure reasoning evaluation. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable capabilities in code generation tasks, but their effectiveness relies on supervised training with extensive labeled data and computational resources. |
| Approach: | They propose an unsupervised method that leverages Internal Probing of Large language models for Code generation without any external corpus, even unlabeled code snippets. |
| Outcome: | The proposed method can achieve competitive performance compared to supervised approaches while reducing the dependency on labeled data and computational resources. |
Copied to clipboard
| Challenge: | Existing knowledge-enhanced methods use a single-source homogeneous knowledge base with limited knowledge coverage. |
| Approach: | They propose a multi-source heterogeneous knowledge-enhanced dialogue generation model that leverages multiple knowledge sources to improve knowledge coverage. |
| Outcome: | The proposed model outperforms existing knowledge-enhanced models on a Chinese dataset and shows that it can leverage multiple heterogeneous knowledge sources to improve knowledge coverage. |
Copied to clipboard
| Challenge: | Current compression techniques entail structural pruning and a recovery phase that leverages the Low-Rank Adaptation algorithm. |
| Approach: | They propose a hierarchical rank allocation method that enables efficient fine-tuning of pruned LLMs according to layerwise specific recovery requirements. |
| Outcome: | The proposed algorithm outperforms state-of-the-art methods across pruning settings and LLM architectures with improvements ranging from 0.7% to 5.5%. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated remarkable success across a wide range of tasks, however, they still face challenges in reasoning tasks that require understanding and inferring relationships between distinct pieces of information within text sequences. |
| Approach: | They propose to construct explicit graphs from context and leverage them to enhance LLM reasoning performance on reasoning tasks. |
| Outcome: | Extensive experiments show that the proposed method improves both logical reasoning and multi-hop question answering tasks. |
Copied to clipboard
| Challenge: | Existing multimodal dialogue systems are limited by the scale and quality of available datasets or the coarse concept of visual knowledge. |
| Approach: | They propose to explicitly split visual knowledge into finer granularity and turn-level . they propose a framework to add visual representation into vanilla dialogue models . |
| Outcome: | The proposed framework outperforms state-of-the-art methods on automatic and human evaluations. |
Copied to clipboard
| Challenge: | Large language models (LLMs) achieve strong performance across domains but remain difficult to deploy in resource-constrained environments due to their massive size. |
| Approach: | They propose an importance-aware activation reconstruction framework that links compression to its effect on model performance. |
| Outcome: | Experiments show that IMPACT reduces model size by 55.4% while maintaining accuracy comparable to or better than state-of-the-art models. |
Copied to clipboard
| Challenge: | CultureBank is a knowledge base built upon users’ self-narratives with 12K cultural descriptors sourced from TikTok and 11K from Reddit. |
| Approach: | They construct a pipeline to construct cultural knowledge bases from different online communities on a massive scale. |
| Outcome: | The proposed pipeline improves cultural awareness of language models by evaluating them on two cultural tasks in a zero-shot setting. |
Copied to clipboard
| Challenge: | Existing research on the allocation of public scarce resources has limitations due to data scarcity and data scariness. |
| Approach: | They propose a framework that integrates Large Language Models into economic simulations . they conduct extensive policy simulation experiments to verify the framework's effectiveness . |
| Outcome: | The proposed framework bridges the gap between theoretical models and real-world dynamics by integrating large language models into economic simulations. |
Copied to clipboard
| Challenge: | Hot news is one of the most popular topics in daily conversations. |
| Approach: | They propose a task where a dialogue system can lead the conversation based on key topics of the news. |
| Outcome: | The proposed method can lead conversations based on key topics of the news . it can also be used in information-seeking and chit-chat scenarios . |
Copied to clipboard
| Challenge: | Dongba pictographic is the only pictograph script still in use in the world. |
| Approach: | DongbaMIE is the first dataset focusing on multimodal information extraction of Dongbe pictographs. |
| Outcome: | The dataset contains 23,530 sentence-level and 2,539 paragraph-level high-quality text-image pairs. |
Copied to clipboard
| Challenge: | Medical coding is the process of translating unstructured clinical notes into standardized diagnostic and procedural codes. |
| Approach: | They propose a closed-loop framework that treats workflow design as a learning problem. |
| Outcome: | The proposed framework outperforms state-of-the-art workflows on benchmark datasets and produces interpretable, adaptable workflows that better reflect real coding practice. |
Copied to clipboard
| Challenge: | Existing watermarking methods reduce the fidelity of semantics in LLMs . |
| Approach: | They propose a low-entropy token partitioning mechanism and z-score-driven dynamic bias mechanism to enhance semantics. |
| Outcome: | The proposed framework improves semantic fidelity and robustness against bias sparsity attacks. |
Copied to clipboard
| Challenge: | Multimodal large language models have demonstrated remarkable performance in visual-language tasks, but their authenticity is often compromised by object hallucinations. |
| Approach: | They propose a multi-frequency perturbation method that leverages both low-frequency and high-frequency features of images to perturb visual feature representations and explicitly suppress redundant frequency-domain features during inference. |
| Outcome: | The proposed method significantly mitigates object hallucinations across various model architectures. |
Copied to clipboard
| Challenge: | Recent studies suggest that self-reflective prompting can significantly enhance the reasoning capabilities of Large Language Models (LLMs). |
| Approach: | They propose guidelines for when to implement self-reflection in Large Language Models. |
| Outcome: | The proposed approach improves the reasoning capabilities of Large Language Models under a more stringent evaluation setting, and reduces tendency toward majority voting. |
Copied to clipboard
| Challenge: | Empirical results show that AFT-trained models achieve substantial gains with test-time scaling. |
| Approach: | They introduce a supervised fine-tuning paradigm where models synthesize multiple draft responses into a single, refined answer. |
| Outcome: | Empirical results show that AFT-trained models outperform baseline models while eliminating external guidance. |
Copied to clipboard
| Challenge: | a goal of LLM alignment is to balance usefulness with harmlessness, but this conflictes when knowledge serves both legitimate and malicious purposes. |
| Approach: | They propose a framework that combines safety-research contexts with adversarial interactions to exploit a vulnerability in Jargon queries. |
| Outcome: | a framework outperforms existing methods in analyzing Jargon queries, a study shows . it achieves 93% of attacks across seven models, while remaining useful, the authors say . |
Copied to clipboard
| Challenge: | Retraining reward models to address privacy leaks and stereotypes is expensive . recent advances in large language models have led to improvements in understanding . |
| Approach: | They propose a lightweight intrinsic reward that can be used to prune existing LLMs to approximate an "untrust" and an ""untrust "" token distribution. |
| Outcome: | Experiments with two reward models and four LLMs show that selfRW improves trustworthiness with minimal impact on general utility benchmarks. |
Copied to clipboard
| Challenge: | Existing models fail to capture and model customer intention effectively because of insufficient information exploitation and only apparent information like descriptions and titles are used. |
| Approach: | They propose to exploit existing session data to capture and model intention in E-commerce product purchase sessions using a multimodal benchmark. |
| Outcome: | The proposed framework can bridge the gap between intention understanding in simplified research cases like co-buy intention and more complex yet practical scenarios like session history. |
Copied to clipboard
| Challenge: | Existing evaluations emphasize final accuracy or coarse token counts, and lack automated tools to separate essential logic from structural redundancy. |
| Approach: | They propose a graph-driven framework that quantifies reasoning efficiency by converting free-form CoTs into directed dependency graphs and extracting the Shortest Effective Path needed to reach a correct solution. |
| Outcome: | Evaluating 21 LRMs, the proposed framework quantifies reasoning efficiency by converting free-form CoTs into directed dependency graphs and extracting the Shortest Effective Path (SEP) needed to reach a correct solution. |
Copied to clipboard
| Challenge: | Large language models lack intrinsic cognition and cannot generalize to complex task-specific scenarios. |
| Approach: | They propose a framework that transitions from mechanistic interpretation to active intervention to map internal distributions of ToM features and implement it via targeted activation steering within ToM-critical layers. |
| Outcome: | The proposed framework significantly enhances human-like social reasoning capabilities and dialogue quality. |
Copied to clipboard
| Challenge: | Existing methods to measure scholarly impact of documents without citations only consider word frequency change. |
| Approach: | They propose a neural network framework that measures document influence without citations by using word frequency changes and word semantic shifts. |
| Outcome: | The proposed model outperforms existing models on document influence evaluation without citations. |
Copied to clipboard
| Challenge: | Generative AI has made rapid advances in multimodal understanding and code generation. |
| Approach: | They construct a first real-world benchmark for multimodal large language models that directly convert visual designs into code implementations by manually curating 484 diverse real-life webpages as test cases. |
| Outcome: | The proposed model can generate code implementations that directly render into the given reference webpages, given the screenshots as input. |
Copied to clipboard
| Challenge: | Existing deep learning paradigms focus on learning a model from training data of a single task and the learned model is also tested on the same task. |
| Approach: | They propose a Bayes-enhanced lifelong attention network to learn attention knowledge from a sequence of sentiment classification tasks and build lifelong ones. |
| Outcome: | The proposed model is able to learn attention knowledge from a set of sentiment classification tasks and build lifelong attentions. |
Copied to clipboard
| Challenge: | Existing efforts to mitigate catastrophic forgetting in continual learning have not been studied. |
| Approach: | They propose a rationale-guided replay framework that allows models to leverage their capabilities and provide partial external correct rationales to the original instructions. |
| Outcome: | The proposed framework mitigates pseudo forgetting while maintaining model plasticity. |
Copied to clipboard
| Challenge: | Existing methods to rate academic papers require a lot of feature engineering and can cause inequality. |
| Approach: | They propose to use a novel convolutional neural network to automatically rate academic papers . they propose to build a dataset to automatically determine whether to accept academic papers. |
| Outcome: | The proposed model outperforms baselines by a large margin. |
Copied to clipboard
| Challenge: | Textual backdoor attacks are increasingly challenging to detect due to the use of advanced generative models such as GPT-4. |
| Approach: | They propose a framework that harnesses advanced generative models to execute stealthier backdoor attacks on text classifiers. |
| Outcome: | The proposed framework achieves state-of-the-art attack success rate of 97.35% over four sentiment classification tasks and four human cognition stealthiness tests. |
Copied to clipboard
| Challenge: | Existing definition generation methods take the source word as an indecomposable semantic unit, but in parataxis languages like Chinese, word meanings can be composed using the word formation process. |
| Approach: | They propose to use word formation features to enhance Definition Generation (DG) in Chinese to generate an explanatory text. |
| Outcome: | The proposed model enhances Definition Generation (DG) in Chinese by decomposing the word meaning into different semantic components. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue generation models struggle with dull and repetitive outputs, a problem commonly termed as text degeneration. |
| Approach: | They propose a framework that allows the model to "cheat" the objective by duplicating knowledge segments in a superficial pattern matching based on overlap. |
| Outcome: | The proposed framework can be applied to a WoW dataset and shows that it works across models and decoding strategies. |
Copied to clipboard
| Challenge: | Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs. |
| Approach: | They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. |
| Outcome: | The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities. |
Copied to clipboard
| Challenge: | Existing methods to solve this problem can not satisfy the following three desiderata: (1) high translation quality, (2) high match accuracy, and (3) low latency. |
| Approach: | They propose a template-based method that can provide high translation quality and match accuracy and a low latency inference. |
| Outcome: | The proposed method outperforms baselines in lexically and structurally constrained translation tasks and can be used in a variety of applications. |