Papers by Huang Liu
Copied to clipboard
| Challenge: | ProUIE improves universal information extraction (UIE) without external information . many LLM-based methods rely on extra schema cues, external resources or complex alignment and verification pipelines . |
| Approach: | They propose a Macro-to-Micro progressive learning approach that improves UIE without external information. |
| Outcome: | ProUIE outperforms instruction-tuned baselines on average for NER and RE while using a smaller backbone. |
Copied to clipboard
| Challenge: | Existing methods for multi-turn function calling are limited by redundancy and lack explicit integration of progress awareness into training. |
| Approach: | They propose a framework that explicitly integrates progress awareness into LLM training for multi-turn function calling. |
| Outcome: | Empirical results show that Progra outperforms existing methods on two public benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to mitigating vision-knowledge conflict in Large Language Models (MLLMs) are not effective and can be further scaled. |
| Approach: | They propose a framework to generate inputs to simulate and evaluate vision-knowledge conflict in Multimodal Large Language Models (MLLMs) using original images and 1,122 high-quality question-answer pairs, they propose 'a diagnostic benchmark' |
| Outcome: | The proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in Multimodal Large Language Models (MLLMs). |
Copied to clipboard
| Challenge: | Existing approaches to hierarchical text classification are limited by lack of domain knowledge, which leads to mistakes in a variety of situations. |
| Approach: | They propose a Knowledge-enabled Hierarchical Text Classification model which integrates knowledge graphs into HTC to address the knowledge limitations of traditional methods. |
| Outcome: | The proposed model integrates knowledge graphs into the hierarchical text classification process, addressing the knowledge limitations of traditional methods. |
Copied to clipboard
| Challenge: | Existing methods aim to fully utilize the dynamic conversation context to enhance the semantic association between the user query and FAQ questions, but they are limited by noise and e.g., users may click questions they don't like, leading to inaccurate semantics modeling. |
| Approach: | They propose to introduce tags of FAQ questions to reduce noise in the conversation context and integrate them into a reinforcement learning framework to minimize the negative impact of irrelevant information. |
| Outcome: | The proposed method can eliminate irrelevant information and minimize negative impact of irrelevant information in the dynamic conversation context. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Language Modeling (RALM) is a popular approach for large language models. |
| Approach: | They propose a modular RALM that integrates large language models with documents from an external corpus to improve inference efficiency. |
| Outcome: | The proposed method improves inference efficiency with appending context pattern while maintaining decent performance after fine-tuning by Low-Rank Adaption. |
Copied to clipboard
| Challenge: | Existing benchmarks and MLLMs focus on single-image input scenarios, leaving performance of ML models when handling multiple images underexplored. |
| Approach: | They propose a benchmark to evaluate fine-grained abilities of multimodal large language models in multi-image scenarios. |
| Outcome: | The proposed benchmark categorizes the multi-image abilities into three scenarios: MII, MKS and MIC. |
Copied to clipboard
| Challenge: | Existing domain-specific knowledge of domain-related tasks is lacking in pre-trained language models. |
| Approach: | They propose a domain-adaptation method which can dynamically select domain-specific tokens and guide the discriminator to emphasize them, without introducing new training parameters. |
| Outcome: | The proposed method can capture domain-specific knowledge of domain-related tasks without introducing new training parameters. |
Copied to clipboard
| Challenge: | Large language models (LLMs) demonstrate remarkable performance across diverse tasks, yet their effectiveness often depends on costly commercial APIs or cloud services. |
| Approach: | They propose a dual-mode compatible approach that fine-tunes models through shortest-response preference optimization and a confidence-aware rejection mechanism. |
| Outcome: | The proposed approach reduces redundant outputs and response times while reducing computational costs by over 50% and cascade latency by over 80%. |
Copied to clipboard
| Challenge: | Existing methods focus on minimizing the number of questions required to assess ability, lacking clear and reliable explanations for the question selection process. |
| Approach: | They propose to use large language models to enhance computer adaptive testing (CAT) by providing human-like interpretability and explanations. |
| Outcome: | The proposed agent-based CAT performs comparably or superior to traditional CAT methods in accuracy and significantly improves student trust and satisfaction. |
Copied to clipboard
| Challenge: | Existing methods for code retrieval struggle to balance scalability and annotation quality. |
| Approach: | They propose a method that integrates functions called within the repository and information on third-party APIs to enhance the annotation context. |
| Outcome: | The proposed method improves the annotation context by incorporating functions called within the repository and information on third-party API functionalities. |
Copied to clipboard
| Challenge: | Existing methods focus on replicating dialogues in textual form, neglecting the role’s voice traits as a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. |
| Approach: | They propose a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. |
| Outcome: | The proposed model exhibits role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. |
Copied to clipboard
| Challenge: | Temporal knowledge graphs (TKGs) require predicting future facts by modeling structural dependencies within each snapshot and temporal evolution across snapshots. |
| Approach: | They propose an encoder-agnostic framework that provides persistent entity states . EST maintains a global state buffer and aligns structural evidence with sequential signals . |
| Outcome: | Experiments show that EST improves diverse backbones and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing benchmarks conflate coordination ability with role-based priors. |
| Approach: | They propose a role-free benchmark for evaluating free-form collaboration under information silos. |
| Outcome: | The proposed benchmark systematically probes coordination capabilities under information silos using 54 configurations and 3 frontier LLMs. |
Copied to clipboard
| Challenge: | Recent efforts to extract local subgraph information from click graphs have hindered collaboratively utilizing global click graph information. |
| Approach: | They propose a global-view long chain interests model that models a click graph with neighbor interest to enhance news recommendation. |
| Outcome: | The proposed method surpasses baseline methods on two real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods to improve performance of large language models rely on additional training objectives or language-specific parameters. |
| Approach: | They propose a bidirectional language projection framework that enables efficient multilingual alignment and language shift using the intrinsic parameters. |
| Outcome: | The proposed framework improves performance of non-dominant languages and improves internal representations. |
Copied to clipboard
| Challenge: | Existing defenses, including post-training alignment and prompt engineering, struggle with adaptability to out-of-distribution (OOD) attacks. |
| Approach: | They propose an adversarial game-based defense method that dynamically adjusts LLMs’ internal representations to achieve a balanced trade-off between helpfulness and harmlessness. |
| Outcome: | The proposed method improves LLMs’ safety over all baselines. |
Copied to clipboard
| Challenge: | a growing need for long document summarization datasets with 16k input is causing problems. |
| Approach: | They propose to use a dataset to analyze salient information in long document summarizations. |
| Outcome: | The proposed dataset outperforms existing models and LLMs in the distribution form of salient information and the distribution of salinal information is an indicator of quality. |
Copied to clipboard
| Challenge: | Recent studies have focused on enhancing reward models through data improvements, following the conventional training framework for reward models that directly optimizes the predicted rewards. |
| Approach: | They propose a hybrid alignment framework **HAF-RM** that incorporates additional constraint on token-level policy probabilities in addition to the reward score. |
| Outcome: | The proposed framework can supervise the internal preference model at the token level and optimize the mapping layer of the reward model at sequence level. |
Copied to clipboard
| Challenge: | Cross-modal retrieval tasks are used to retrieve data from one modality or another based on a query from another modality. |
| Approach: | They propose a generative cross-modal retrieval framework based on coarse-to-fine semantic modeling . they propose combining K-Means and RQ-VAE to discretize multimodal data into token sequences that support autoregressive generation. |
| Outcome: | The proposed framework achieves excellent performance and efficiency in multimodal retrieval tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have recently achieved remarkable progress on complex reasoning tasks by leveraging extended Chain-of-Thought (CoT) techniques. |
| Approach: | They propose a method that uses Extended Chain-of-Thought (EFT) to reduce the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
| Outcome: | The proposed method reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are currently used to evaluate scientific papers by assigning an absolute score to each paper independently. |
| Approach: | They propose a comparison-native framework for paper evaluation that integrates comparison into both data construction and model learning. |
| Outcome: | The proposed framework achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B, while exhibiting robust generalization to five previously unseen datasets. |
Copied to clipboard
| Challenge: | Chinese spelling correction (CSC) is a task which detects incorrect characters in Chinese text and corrects them. |
| Approach: | They propose to pre-train a Chinese spelling correction corrector under the detector-corrector architecture and propose to capture pronunciation and shape information in Chinese characters. |
| Outcome: | The proposed corrector achieves an average of 5.8% F1 improvements over state-of-the-art methods, verifying its effectiveness. |
Copied to clipboard
| Challenge: | Existing methods for creating versatile MLLMs rely on joint training with paired instruction data, which is resource-intensive and challenging to extend to new modalities. |
| Approach: | They propose a new paradigm for multimodal large language models by reusing modality encoders and merging LLM parameters. |
| Outcome: | The proposed model retains the modal understanding capabilities of each original model. |
Copied to clipboard
| Challenge: | Existing diffusion models fail to address the challenges of generating high-quality images from textual descriptions due to its large vocabulary size and complex character relationships. |
| Approach: | They propose a framework that integrates Chinese diffusion models with Alibaba Cloud's Platform for AI and enables the generation of contextually relevant images. |
| Outcome: | The proposed framework integrates with Alibaba Cloud’s Platform for AI, providing accessible and scalable solutions. |
Copied to clipboard
| Challenge: | Prior work typically decomposes inference into prefill and decode stages, with the decode stage dominating total latency. |
| Approach: | They propose an algorithm that detects threshold where information loss exceeds information gain during sparse decoding to reduce token consumption by up to 90% and a marginal accuracy degradation of less than 2%. |
| Outcome: | The proposed algorithm reduces token consumption by 90% with a marginal accuracy degradation of less than 2% across reasoning-intensive benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for LGT detection assume that it is a single homogeneous distribution. |
| Approach: | They propose a framework for LGT detection based on density-aware manifold learning and hybrid Mahalanobis energy. |
| Outcome: | The proposed framework outperforms baselines in detecting LLM-generated text (LGT) it is based on density-aware manifold learning and hybrid Mahalanobis energy . |
Copied to clipboard
| Challenge: | Existing models exhibit memorization and generalization behaviors in ways that are not easily interpretable or controllable. |
| Approach: | They propose to use a GPT-2 and LLaMA-3.2 model to identify distinct neuron subsets responsible for each behavior to steer the model toward memorization or generalization. |
| Outcome: | The proposed models show that inference-time interventions on these neurons can steer the model’s behavior toward memorization or generalization. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) do not accurately uncover safety vulnerabilities in LLMs. |
| Approach: | They propose a value alignment benchmark called Flames that encompasses both harmlessness principles and a unique morality dimension that integrates specific Chinese values such as harmony. |
| Outcome: | The proposed model performs poorly on Flames, particularly in safety and fairness dimensions. |
Copied to clipboard
| Challenge: | a comparative analysis of paper (meta-)reviews by large language models (LLMs) aims to identify and distinguish LLMs from human activities . |
| Approach: | They present a comparative analysis to identify and distinguish LLM activities from human activities. |
| Outcome: | The proposed analysis aims to improve recognition of instances when someone implicitly uses LLMs for reviewing activities. |
Copied to clipboard
| Challenge: | Existing textless speech-to-speech translation models have two main challenges: 1) learning cross-modal features and 2) learning alignment of difference languages in long sequences. |
| Approach: | They propose a unit language to overcome two main modeling challenges . they propose task prompt modeling to utilize the unit language in guiding the modeling process. |
| Outcome: | The proposed language improves over a strong baseline and achieves comparable performance to models trained with text. |
Copied to clipboard
| Challenge: | Recent large language model-based approaches often overlook graph context or depend on distillation from larger models, limiting generalisation. |
| Approach: | They propose a framework for zero-shot reasoning on text-rich networks . they use a Neighbour-aware Group Relative Policy Optimisation objective . |
| Outcome: | The proposed framework optimises base LLMs using a Neighbour-aware group relative policy optimisation objective based on a novel margin gain metric for the informativeness of neighbouring signals . |
Copied to clipboard
| Challenge: | Existing methods to extract features from images of entities overlook varying relevance of visual information across entities. |
| Approach: | a new model integrates structural and multimodal information of entities into a multimodal knowledge graph . a model evaluates the necessity of visual modality for each entity based on its attributes . |
| Outcome: | The proposed model improves on existing methods by adjusting visual data to different entity types. |
Copied to clipboard
| Challenge: | Existing models of robustness evaluation are incomprehensive, impractical, and invalid . |
| Approach: | They propose a framework for automatic robustness evaluation that shifts towards model-centric evaluation to further exploit the advantages of adversarial attacks. |
| Outcome: | The proposed framework is based on a model-centric evaluation protocol and a robustness evaluation protocol. |
Copied to clipboard
| Challenge: | Existing evaluation methods for transfer learning are limited in speech research . authors show that pre-trained models transfer well across multiple tasks . |
| Approach: | They propose a benchmark to evaluate pre-trained models by increasing task diversity and difficulty over SUPERB. |
| Outcome: | The proposed benchmark increases task diversity and difficulty over SUPERB-SG. |
Copied to clipboard
| Challenge: | Existing methods for handwriting generation capture global dependencies and can generate high-quality handwritten samples. |
| Approach: | They propose a Transformer-based model for ink generation, TrInk, which captures global dependencies. |
| Outcome: | The proposed model reduces character error rate and word error rate by 35.56% on the IAM-OnDB dataset compared to previous models. |
Copied to clipboard
| Challenge: | Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair . |
| Approach: | They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness. |
| Outcome: | The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms . |
Copied to clipboard
| Challenge: | Large Language Models encode vast factual knowledge, yet their inability to selectively forget specific information hinders privacy protection, bias mitigation, and post-deployment correction. |
| Approach: | They propose a LoRA-based negative-only unlearning framework that updates only low-rank adapters while freezing the backbone. |
| Outcome: | The proposed framework reduces computational cost by about an order of magnitude compared to full fine-tuning and memory-editing methods. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) improves the fine-tuning efficiency and performance of large language models. |
| Approach: | They propose a low-rank adaptive approach that decomposes update matrix into global and local adapters and assigns them to local and global adapters. |
| Outcome: | The proposed method achieves up to 2.7% accuracy improvement over LoRA and its variants on commonsense reasoning, mathematical reasoning, and code generation. |
Copied to clipboard
| Challenge: | Existing domain adaptation rumor detection methods ignore the data generalization differences and rely on a large amount of unlabeled target domain samples to achieve domain adaptation. |
| Approach: | They propose a Gradient Coherence guided Meta-Learning approach for emerging topics rumor detection that selectively learns more "generalizable" tasks that are more beneficial in adapting to the target domain. |
| Outcome: | The proposed method outperforms baselines on real-world datasets and significantly outperformed traditional methods on the in-domain condition. |
Copied to clipboard
| Challenge: | Existing benchmarks evaluate agents in simplified, idealized settings, relying on pre-packaged tool interfaces, overlooking critical steps, and assume inputs are clean and fully specified. |
| Approach: | They propose a framework that evaluates language agents in simplified, idealized settings . they show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 . |
| Outcome: | Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 . |
Copied to clipboard
| Challenge: | Experimental results show that, as the instruction data increases, LoRAMoE can significantly improve the ability to process downstream tasks, while maintaining the world knowledge stored in the LLM. |
| Approach: | They propose a framework that introduces several low-rank adapters and integrates them by using a router network to freeze the backbone model and force a portion of LoRAs to focus on leveraging world knowledge to solve downstream tasks. |
| Outcome: | The proposed framework freezes the backbone model and forces a portion of LoRAs to focus on leveraging world knowledge to solve downstream tasks, to alleviate world knowledge forgetting. |
Copied to clipboard
| Challenge: | High-quality scientific data is critical for advancing LLMs, yet academic literature remains underutilized. |
| Approach: | They construct a large-scale raw scientific corpus but identify a critical Learnability Gap . they develop a multi-stage pipeline featuring content cleaning and pedagogical augmentation . |
| Outcome: | The proposed approach boosts average performance by +2.12 (3B) and +2.95 (7B) on in-domain tasks. |
Copied to clipboard
| Challenge: | Existing systems that use a left-to-right completion paradigm are inefficient and expensive. |
| Approach: | They propose an open-source end-to-end interactive machine translation system platform . they propose to use a prefix-constrained decoding approach to achieve end- to-end evaluation . |
| Outcome: | The proposed system can guarantee high-quality, error-free translations . it uses prefix-constrained decoding and improves on previous systems . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on simple image-text interactions, overlooking complex visual formats like charts. |
| Approach: | They propose a semi-automatic framework for generating evaluation samples through multi-modal keypoint extraction, knowledge graph construction, and qa pair synthesis. |
| Outcome: | The proposed framework generates 4,738 question-answering pairs across 8 domains from real-world documents. |
Copied to clipboard
| Challenge: | Typical approaches to training large language models rely on limited contrasting patterns . contrasting data is limited and models are susceptible to harmful response tendencies . |
| Approach: | They propose a framework that integrates contrasting patterns across the prompt, model, and pipeline levels. |
| Outcome: | The proposed framework outperforms existing methods in the comparison of RQ1 and RQ2 . the proposed framework significantly outperformed existing methods, leading to more comprehensive alignment. |
Copied to clipboard
| Challenge: | Existing work on extracting events from news documents focuses on a set of pre-specified event types. |
| Approach: | They propose a latent variable neural model which is scalable to large corpus. |
| Outcome: | The proposed model performs better than the state-of-the-art method for event schema induction. |
Copied to clipboard
| Challenge: | Existing models capture cross-sentence relations with recurrent neural networks, but they are hard to capture sentence-level long-distance dependency. |
| Approach: | They propose a graph-based neural network for extractive summarization which contains semantic nodes apart from sentences. |
| Outcome: | The proposed graph-based neural network is the first to incorporate different types of nodes into it and perform a qualitative analysis. |
Copied to clipboard
| Challenge: | Context-DPO is the first alignment method specifically designed to enhance contextfaithfulness for large language models. |
| Approach: | They propose a benchmark that simulates Retrieval-Augmented Generation scenarios with knowledge conflicts to evaluate context-faithfulness. |
| Outcome: | The proposed method improves LLMs' context-faithfulness by 35% to 280% over open-source models. |
Copied to clipboard
| Challenge: | Existing methods for large language models suffer from two major issues: in-domain data are scarce compared with general domain-agnostic data. |
| Approach: | They propose a task-oriented in-domain data augmentation framework that uses in- domain data selection and task-orientated synthetic passage generation to adapt LLMs to two domains: advertisement and math. |
| Outcome: | The proposed framework improves LLM performance by 8% in the advertisement domain and 7.5% in the math domain. |
Copied to clipboard
| Challenge: | Triton is a high-level Python-like programming language for building efficient GPU kernels. |
| Approach: | They propose a TritonBench benchmark that provides a comprehensive evaluation of Tritonic operators on widely deployed GPUs. |
| Outcome: | The proposed benchmarks show that current LLMs struggle to generate efficient Triton operators on widely deployed GPUs aligned with industry applications. |
Copied to clipboard
| Challenge: | Existing methods for budget-constrained tool learning have been overlooked . et al., 2023b) compared tool learning with other methods to improve performance . |
| Approach: | They propose a method for budget-constrained tool learning by creating a preferable plan under the budget constraint before utilizing the tools. |
| Outcome: | The proposed method reduces the cost of tool learning and reaches competitive Pass Rate. |
Copied to clipboard
| Challenge: | Existing methods for GMNER fail to address semantic ambiguity caused by polysemy and long-tail distribution of datasets. |
| Approach: | They propose a framework for Grounded Multimodal Named Entity Recognition that leverages a Multimodal Large Language Model to address semantic ambiguity. |
| Outcome: | Extensive experiments show that the proposed framework outperforms existing methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for image-to-text generation store all knowledge within parameters, thus requiring computational-expensive fine-tuning. |
| Approach: | They propose a Retrieval-augmented Visual Language Model that stores all the knowledge within parameters and can be used to retrieve it from the external database. |
| Outcome: | The proposed model significantly boosts performance for image-to-text generation tasks with 4x less parameters compared with baseline methods. |
Copied to clipboard
| Challenge: | *entity-centric question generation (ECQG) is a task motivated by real-world applications such as topic-specific learning, assisted reading, and fact-checking. |
| Approach: | They propose a PLM-based framework GenCONE with two modules: content focusing and question verification. |
| Outcome: | The proposed framework outperforms baselines and is effective and complementary in generating high-quality questions. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) lack visual input, leading to errors in basic numerical comparisons. |
| Approach: | They propose a spatial OODA framework that integrates the OODAC cognitive loop into multiple control tasks and integrates it into LLMs. |
| Outcome: | The proposed model significantly improves the spatial reasoning capabilities of large language models across multiple scenarios including SPOD-Bench, SPACE and applications. |
Copied to clipboard
| Challenge: | Existing approaches to align pre-trained LLMs with instructions for one property are difficult to fine-tune. |
| Approach: | They propose a mixture-of-experts-based fusion mechanism that models alignment as a controllable drift within the subspace, guided by a drift-regularization loss to balance competing alignment dimensions. |
| Outcome: | Extensive evaluations of three benchmark datasets show that H3Fusion outperforms each individually aligned model by 11.37% and provides stronger robustness compared to the state-of-the-art LLM ensemble approaches by 13.77% and model-merging approaches by 6.18 %. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) systems focus on improving model performance, ignoring the need to quantify model uncertainty. |
| Approach: | They propose to introduce two uncertainty-guided loss terms to the conventional EDL and a series of uncertainty-guiding training strategies to solve these challenges. |
| Outcome: | The proposed method achieves better OOV/OOD detection performance and generalization ability on OOV entities compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of Large Language Models (LLMs). |
| Approach: | They propose a repository-level benchmark named DevEval to evaluate LLMs' coding abilities in real-world code repositories. |
| Outcome: | The proposed benchmarks show that the LLMs perform better in real-world code repositories than existing benchmarks. |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) improves large language models by incorporating non-parametric knowledge through evidence retrieved from external sources. |
| Approach: | They propose a training-free evidence compression technique that makes retrieved evidence more familiar to the target model while seamlessly integrating parametric knowledge from the model. |
| Outcome: | The proposed technique outperforms the most recent evidence compression baselines across open-domain QA datasets while achieving high compression rates. |
Copied to clipboard
| Challenge: | Current alignment approaches struggle with inconsistency and sparsity of human supervision signals. |
| Approach: | They propose a framework modeling hierarchical rewards in reinforcement learning from human feedback (RLHF) it integrates holistic rewards with aspect-specific rewards to enhance alignment of large language models with human preferences. |
| Outcome: | The proposed framework improves the alignment of large language models with human preferences by integrating holistic rewards with aspect-specific rewards. |
Copied to clipboard
| Challenge: | Existing methods for generating follow-up questions are limited to shallow contextual questions that are uninspiring and have a large gap to the human level. |
| Approach: | They propose a three-stage external knowledge-enhanced follow-up question generation method which generates questions by identifying contextual topics, building a knowledge graph online, and finally combining these with a large language model to generate the final question. |
| Outcome: | The proposed method generates questions by identifying contextual topics, building a knowledge graph (KG) online, and finally combining these with a large language model to generate the final question. |
Copied to clipboard
| Challenge: | Existing information retrieval benchmarks focus on general or specialized domains, such as medicine or finance, neglecting the unique linguistic complexity and diverse information needs encountered in disaster management scenarios. |
| Approach: | DisastIR is the first comprehensive IR evaluation benchmark specifically tailored for disaster management. |
| Outcome: | DisastIR covers 48 retrieval tasks derived from six search intents and eight general disaster categories . evaluations show no single model excelling universally . |
Copied to clipboard
| Challenge: | Computer-aided translation (CAT) is a form of software that assists a human translator in the translation process. |
| Approach: | They propose to use computer-aided translation (CAT) to assist a human translator in the translation process. |
| Outcome: | The proposed method can give significantly more accurate predictions than baseline methods on CAT datasets. |
Copied to clipboard
| Challenge: | Text-based image generation models, such as Stable Diffusion and DALL-E 3, hold significant potential in content creation and publishing workflows . however, considerable efforts are being made to prevent the generation of harmful content, such abusive, violent, or pornographic material. |
| Approach: | They propose a chain-of-jailbreak method which decomposes malicious queries into multiple sub-queries and iteratively edits images based on these sub-questions. |
| Outcome: | The proposed method can bypass safeguards of image generation models for over 60% cases, significantly outperforms other jailbreaking methods (14%) |
Copied to clipboard
| Challenge: | Existing methods for knowledge distillation use a two-stage paradigm: general distillation with a task-agnostic general corpus and task-specific distillation using augmented task- specific corpus. |
| Approach: | They propose a contextualized corpus that contextualizes task corpus with large-scale general corpus through relevance-based text retrieval to improve student learning. |
| Outcome: | The proposed model improves on the GLUE benchmark and shows that it is better than generalized corpus and augmented task-specific corpus. |
Copied to clipboard
| Challenge: | Existing interpretation methods only support tasks with specific inputs, limiting their practical applications. |
| Approach: | They propose an extensible module that matches different input data with interpretation methods and consolidates the interpreting outputs. |
| Outcome: | The proposed module can match different input data with interpretation methods and consolidate the interpreting outputs. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been used for general-purpose interfaces across multiple tasks and languages. |
| Approach: | They propose to use large language models as a general-purpose interface across multiple tasks and languages. |
| Outcome: | The proposed model performs better on 200K hours of 6-language data for voice generation applications. |
Copied to clipboard
| Challenge: | lack of reliable reward models for tool-use tasks has limited progress toward agentic AI . recent advances in agentic artificial intelligence are driven by tool-using capabilities of large language models. |
| Approach: | They propose a pipeline that constructs pairwise preference data using rule-based scoring and multidimensional sampling to build lightweight reward models. |
| Outcome: | The proposed model outperforms existing models on tool calling tasks with higher accuracy. |
Copied to clipboard
| Challenge: | Existing knowledge-enhanced pre-trained language models (KEPLMs) can capture internal knowledge, but can't understand external background knowledge. |
| Approach: | They propose to use Chinese knowledge-enhanced pre-trained language models to improve context-aware representations via learning from structured relations in knowledge bases. |
| Outcome: | Experiments show that Chinese knowledge-enhanced pre-trained language models outperform strong baselines over various benchmark NLP tasks and in different model sizes. |
Copied to clipboard
| Challenge: | Video dubbing systems use neural machine translation and text-to-speech technologies to translate original speech into visual media programs. |
| Approach: | They propose a preference optimization method to optimize video dubbing duration alignment . they propose combining segment-wise sampling and fine-grained loss to mitigate duration mismatches . |
| Outcome: | The proposed method achieves superior performance in duration alignment tasks. |
Copied to clipboard
| Challenge: | Existing approaches to personalize large language models (LLMs) rely on heuristic methods to compress user profiles but they ignore how LLMs process and prioritize different profile components. |
| Approach: | They propose an attention-guided context compression framework that leverages attention feedback from a marking model to mark important personalization sentences and guides a compression model to generate task-relevant compressed user contexts. |
| Outcome: | The proposed framework outperforms baselines across tasks, token limits, and settings while reducing token usage by 50 times. |
Copied to clipboard
| Challenge: | C-World enables users to build agent environments on demand. |
| Approach: | They propose a system that enables users to build agent environments on demand. |
| Outcome: | The proposed system outperforms baselines on 119k samples and achieves Spearman = 0.883 ranking correlation with real execution. |
Copied to clipboard
| Challenge: | Experimental results show that Flow-matching generative models can scale training by increasing data, computational resources, and model size. |
| Approach: | They propose a flow-matching transformer with masked generative modeling for scaling text-to-audio inference-time prediction. |
| Outcome: | The proposed model scales inference-time computations by masking generation and re-predicting them through iterative decoding. |
Copied to clipboard
| Challenge: | Text summarization aims to generate a short summary for an input text. |
| Approach: | They propose a non-autoregressive unsupervised summarization approach which performs edit-based search towards a heuristicically defined score and generates a summary as pseudo-groundtruth. |
| Outcome: | The proposed approach achieves state-of-the-art performance for unsupervised summarization, while improving inference efficiency. |
Copied to clipboard
| Challenge: | Existing approaches to multihop reasoning fail to address the problem of spurious paths . existing approaches neglect the internal semantic consistency of the reward function . |
| Approach: | They propose a framework that incorporates semantic consistency into the reward function to guide multi-hop reasoning. |
| Outcome: | The proposed framework outperforms baseline methods and facilitates more interpretable reasoning paths. |
Copied to clipboard
| Challenge: | Recent advances in Multi-modal Large Language Models (MLLMs) introduce significant variability in data quality. |
| Approach: | They propose to use human and LLM preference alignment to compress large corpus of machine-generated multimodal instructions into a compact and high-quality form. |
| Outcome: | The proposed algorithm outperforms LLaVA-series models in MLLM benchmarks by 90% . it uses human and LLM preference alignment to compress a large dataset . |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. |
| Approach: | They propose a multimodal safety awareness benchmark to evaluate MLLMs across 29 safety scenarios with 1,500 carefully curated image-prompt pairs. |
| Outcome: | The proposed model is able to identify unsafe content and avoid over-sensitivity that can hinder helpfulness. |
Copied to clipboard
| Challenge: | Puns add the challenge of fusing commonsense and world knowledge with the ability to interpret lexical-semantic ambiguity. |
| Approach: | They propose to augment existing datasets with detailed crowdsourced annotations of puns, keywords and fine-grained funniness ratings to challenge current models' ability to understand and generate humor. |
| Outcome: | The proposed tasks include explanation generation to aid with pun classification and keyword-conditioned pun generation to challenge state-of-the-art models' ability to understand and generate humor. |
Copied to clipboard
| Challenge: | Weakly supervised question answering usually has only final answers as supervision signals while correct solutions are not provided. |
| Approach: | They propose to explicitly exploit the semantic correlations between question-answer pairs and predicted answers by maximizing mutual information between question and answer pairs. |
| Outcome: | The proposed method significantly outperforms previous learning methods in terms of task performance and is more effective in training models to produce correct solutions. |
Copied to clipboard
| Challenge: | Text Image Machine Translation (TIMT) is a critical subfield of machine translation . it requires accurate optical character recognition, robust visual-text reasoning, and high-quality translation a challenge . |
| Approach: | They propose a multi-task optimization framework to specialize MLLMs into expert TIMT models. |
| Outcome: | The proposed model outperforms baselines on the latest in-domain MIT-10M benchmark. |
Copied to clipboard
| Challenge: | Existing literature on mechanistic interpretation (MI) treats it as an observational science, leaving practical applications underexplored. |
| Approach: | They propose a survey structured around the pipeline to identify and improve MI models. |
| Outcome: | The proposed framework enables tangible improvements in Alignment, Capability, and Efficiency. |
Copied to clipboard
| Challenge: | Rather than pursuing the reachless SOTA accuracy, researchers are focusing on model efficiency and usability. |
| Approach: | They propose an evaluation and a public leaderboard for efficient NLP models that depicts the Pareto Frontier for various language understanding tasks. |
| Outcome: | The proposed model outperforms or performs on par with SOTA compressed and early exiting models. |
Copied to clipboard
| Challenge: | Existing methods for mapping monaural audio to binaural signals lack flexibility and interactive control needed in complex multi-object user-interactive environments. |
| Approach: | They propose a text-guided audio spatialization framework that utilizes diverse text prompts to evaluate binaural audio models. |
| Outcome: | The proposed framework learns binaural differences guided by 3D spatial location and relative position prompts, enhanced with flipped-channel audio. |
Copied to clipboard
| Challenge: | Existing studies focus on developing models that exploit the unification of multiple modalities. |
| Approach: | They propose to maintain modality independence by using a multi-modal transformer model that fuses all modalities. |
| Outcome: | The proposed model outperforms state-of-the-art models in multi-modal emotion recognition. |
Copied to clipboard
| Challenge: | Existing approaches to extract triplets from sentences neglect the mutual information between aspects and have the problem of error propagation. |
| Approach: | They propose a Semantic and Syntactic Enhanced aspect Sentiment triplet Extraction model to exploit the syntactical and semantic relationships between the triplet elements and jointly extract them. |
| Outcome: | The proposed model outperforms existing methods on four benchmark datasets and significantly outperformed existing approaches. |
Copied to clipboard
| Challenge: | Existing models require a more expressive vocabulary to represent all languages . however, increasing the vocabulary size significantly slows down the pre-training speed . |
| Approach: | They propose an algorithm VoCap to determine the desired vocabulary capacity of each language. |
| Outcome: | The proposed algorithm improves cross-lingual model pre-training while reducing side effects of increasing vocabulary size. |
Copied to clipboard
| Challenge: | Z-Code++ is a pre-trained language model optimized for abstractive text summarization. |
| Approach: | They propose a pre-trained language model optimized for abstractive text summarization that uses a two-phase pre-training technique to improve model's performance. |
| Outcome: | The proposed model outperforms the competing models on low-resource summarization tasks in zero-shot and few-shot settings. |
Copied to clipboard
| Challenge: | Existing methods for addressing ambiguities in conversational search systems are one-size-fits-all and struggle to achieve effective domain transferability. |
| Approach: | They propose a method to provide search engines with strategies regarding when to ask clarification questions in a post-hoc manner. |
| Outcome: | The proposed method improves search performance 10% on four unseen domains. |
Copied to clipboard
| Challenge: | Recent research shows that themes and words within a conversation change across time, whereas topics and the patient's attitude towards their willingness to change might shift. |
| Approach: | They propose a method that models the temporal factor by using domain adaptation on clinical dialogue corpora, Motivational Interviewing (MI). |
| Outcome: | The proposed method improves on a college alcoholism dataset using a bi-LSTM and topic model to learn language usage change across different time sessions. |
Copied to clipboard
| Challenge: | Recent studies have fine-tuned judge models based on open-source LLMs to evaluate the quality of other LLM. |
| Approach: | They propose to use open-source LLMs to evaluate Large Language Models (LLMs) their empirical results show that the models underperform GPT-4 in several dimensions . |
| Outcome: | The proposed models outperform GPT-4 on several dimensions including generalizability, fairness and adaptability. |
Copied to clipboard
| Challenge: | Recent work shows that dense hierarchical retrieval (DHR) can outperform dense passage retrieval. |
| Approach: | They propose a framework that applies sparse, dense and a combination of them to document and passage retrieval. |
| Outcome: | The proposed framework can outperform dense hierarchical retrieval (DHR) and sparse retrievers (BM25) on open-domain question answering (ODQA) datasets with an average improvement of 4.69% on recall@100 over DHR. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations. |
| Approach: | They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions . |
| Outcome: | Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning. |
Copied to clipboard
| Challenge: | Existing methods for event temporal relation extraction ignore meaning of relations and wipe out their intrinsic dependency. |
| Approach: | They propose a unified event temporal relation extraction framework that transforms temporal relations into logical expressions of time points and completes the ETRE by predicting the relations between certain time points. |
| Outcome: | The proposed framework outperforms the state-of-the-art model on TB-Dense and MATRES by 0.3% on both datasets. |
Copied to clipboard
| Challenge: | Recent results report a surge in performance to nearhuman levels on the Winograd Schema Challenge (WSC) however, variations in task formulation across papers and evaluations makes it hard to understand the true degree of recent progress. |
| Approach: | They propose to use a model with multiple choice to frame the task as multiple choice and reuse a pretrained language modeling head to mitigate the model's extreme sensitivity to hyperparameters. |
| Outcome: | The proposed frameworks improve the model's reasoning ability by framing the task as multiple choice and reuse of a pretrained language modeling head. |
Copied to clipboard
| Challenge: | Existing approaches to query–document relevance assessment are limited . ambiguous user intent and asymmetric relevance are challenges for RAG platforms . |
| Approach: | They propose a decomposed reasoning model for relevance assessment that decomposes query intent into intent inference and evidence grounding. |
| Outcome: | The proposed model outperforms strong baselines on offline benchmarks and achieves significant gains in large-scale online A/B testing. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have reshaped code generation, but persistent challenges impede accurate assessment. |
| Approach: | They propose an online evaluation framework tailored for large language models to assess their coding capabilities. |
| Outcome: | a new evaluation framework for large language models (LLMs) provides unbiased, unbiased evaluations and open access to solutions and test cases. |
Copied to clipboard
| Challenge: | Existing multilingual benchmarks show severe drawbacks, such as overly translated content, the absence of difficulty control, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. |
| Approach: | They propose a multilingual benchmark that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening to provide a comprehensive, difficult multilingual assessment. |
| Outcome: | The proposed benchmark features 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. |
Copied to clipboard
| Challenge: | Existing LLMs often rely on complex prompting or extensive fine-tuning to introduce new capabilities while preserving strong generalizability. |
| Approach: | They propose a large-scale pre-training corpus to enhance LLM agents' capabilities . they use 103B agent-specific data encompassing 76,537 APIs . |
| Outcome: | The proposed training corpus outperforms open-source LLMs and commercial LLM agents on three agent benchmarks. |
Copied to clipboard
| Challenge: | Recent years have witnessed remarkable progress achieved by large language models in both natural language understanding and generation. |
| Approach: | They propose a large benchmark CMoralEval for moral evaluation of Chinese LLMs . they use a Chinese TV program discussing Chinese moral norms and Chinese moral anomies based on various sources . |
| Outcome: | The proposed dataset is characterized by diversity and authenticity. |
Copied to clipboard
| Challenge: | Recent transformer-based language models (LMs) provide reasoning over textual benchmarks . RAC is essential to understand and interact with the ever-changing environment . |
| Approach: | They propose to use a transformer-based language model to learn to reason over textual benchmarks. |
| Outcome: | The proposed model minimizes the influence of other linguistic requirements to focus on RAC. |
Copied to clipboard
| Challenge: | Existing methods for training reward models are vulnerable to context neglect and degraded accuracy. |
| Approach: | They propose distribution-aware reward modeling that augments the RM objective with a conditional mutual information regularizer that maximizes context and the predicted reward conditioned on the response. |
| Outcome: | The proposed model improves performance in RLHF and improves accuracy in other settings. |
Copied to clipboard
| Challenge: | Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or suffix tokens to achieve a specific representation manually or automatically. |
| Approach: | They propose a method that rewrites the original instruction to achieve a jailbreak . they propose rewriting the original instructions to improve the attack strategy . |
| Outcome: | The proposed method is more efficient and easier to identify since no additional features are introduced. |
Copied to clipboard
| Challenge: | Existing evaluations focus on isolated, short-term interactions, overlooking the inherently long-term nature of learning. |
| Approach: | They propose a benchmark for long-term personalized tutoring based on an annotated learning log . they propose an automated generator–verifier pipeline to enable benchmark expansion . |
| Outcome: | The proposed benchmarks evaluate LLMs across three progressive tasks: evidence acquisition, knowledge state diagnosis, and adaptive teaching action. |
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning with verifiable rewards (RLVR) rely on static objective functions and rigid clipping strategies that misalign with the model’s evolving reasoning capabilities. |
| Approach: | They propose to incorporate Power-Mean Policy Optimization (PMPO) and Feedback-Adaptive Clipping (FAC) to overcome limitations of static mechanisms. |
| Outcome: | Extensive experiments on nine reasoning tasks show the proposed paradigm outperforms state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing systems with opaque architectures are limiting deep search capabilities for web-augmented large language models. |
| Approach: | They propose a transparent and modular multi-agent framework to democratize deep search for LLMs. |
| Outcome: | The proposed framework outperforms open-source systems in deep reasoning tasks. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to memorize, recall, and reason in sustained interactions. |
| Approach: | They propose a multimodal real-world conversation benchmark for evaluating open-ended abilities of multimodal large language models. |
| Outcome: | The proposed benchmarks show that the models perform better in open-ended conversations. |
Copied to clipboard
| Challenge: | Creating lyrics and melodies in symbolic format requires expert knowledge of melody and an advanced understanding of lyrics. |
| Approach: | They introduce SongComposer, a music-specialized large language model that can create symbolic lyrics and melodies following instructions. |
| Outcome: | The proposed model outperforms existing models in symbolic song composition tasks. |
Copied to clipboard
| Challenge: | Existing Braille research focuses on isolated tasks while mixed-content Braille tasks face data scarcity and ambiguities. |
| Approach: | They propose a syntax tree-based augmentation method tailored for Braille data. |
| Outcome: | The proposed method improves Braille translation, formula-to-Braille conversion, and mixed-text translation. |
Copied to clipboard
| Challenge: | Task-oriented dialog systems typically manage structured knowledge to guide goal-oriented conversations. |
| Approach: | They propose a TOD system with hybrid knowledge management, HyKnow, which extends the belief state to manage both structured and unstructured knowledge. |
| Outcome: | The proposed model outperforms existing TOD systems in the evaluation of a multiWOZ dataset on unstructured knowledge with strong end-to-end performance. |
Copied to clipboard
| Challenge: | Existing benchmarks and evaluation protocols focus on surface-level factual recall. |
| Approach: | They propose a benchmark for assessing cognitive memory under cue–trigger semantic disconnect. |
| Outcome: | The proposed framework reveals failures not captured by existing benchmarks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are a promising alternative to expensive human evaluations. |
| Approach: | They propose a framework that iteratively aligns LLM-based evaluators with human preference . they decompose a given evaluation task into finer-grained criteria . |
| Outcome: | The proposed framework iteratively aligns LLM-based evaluators with human preference . it decomposes a given evaluation task into finer-grained criteria . the framework is efficient to train and more explainable than relying solely on prompts . |
Copied to clipboard
| Challenge: | Textual adversarial samples are often misrepresented in research on security, evaluation, explainability, and data augmentation. |
| Approach: | They propose to use adversarial samples to evaluate their methods on security tasks to demonstrate the real-world concerns rather than developing impractical methods. |
| Outcome: | The proposed method has higher practical value than the current benchmark. |
Copied to clipboard
| Challenge: | Existing unsupervised reinforcement learning methods lack the capacity to adapt to the model’s evolving reasoning capabilities during training. |
| Approach: | They propose an unsupervised reinforcement learning algorithm that adapts rewards to balance consensus and exploration based on the Free Energy Principle. |
| Outcome: | Empirical evaluations on nine datasets show that FREIA outperforms baseline methods on reasoning tasks. |
Copied to clipboard
| Challenge: | Language models have demonstrated remarkable performance in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. |
| Approach: | They propose a method to assess concept bias in models during fine-tuning and in-context learning using ChatGPT. |
| Outcome: | The proposed method outperforms token removal approaches and is validated through extensive testing. |
Copied to clipboard
| Challenge: | Current datasets cater to user-led systems and are limited to predefined specific scenarios and slots. |
| Approach: | They propose to use a Chinese dialogue dataset to train a model that authentically simulates human-computer dialogues in 30 popular life service scenarios. |
| Outcome: | The proposed model achieves a joint accuracy of 75.09% in out-of-domain evaluations . it also achieves notable abilities in slot filling and questioning . |
Copied to clipboard
| Challenge: | Existing video understanding evaluation frameworks that use fill-in-the-blanks do not reflect real-world tasks. |
| Approach: | They propose to use fill-in-the-blanks as a video understanding evaluation framework and introduce a novel dataset that collects multiple perspectives on the same video. |
| Outcome: | The proposed framework does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth. |
Copied to clipboard
| Challenge: | Large vision-language models are prone to hallucinations, where contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. |
| Approach: | They propose to automate the generation of hallucination-related questions using images . they propose to use three image manipulation strategies to induce hallucinosity . |
| Outcome: | The proposed approach reduces human bias in crafting such examples and improves accuracy. |
Copied to clipboard
| Challenge: | Existing rule retrieval methods suffer from low accuracy due to semantic gap between instantiated facts and abstract representations of rules. |
| Approach: | They propose a method that induces inferential rules that might offer benefits for reasoning by abstracting the underlying knowledge and logical structure in queries. |
| Outcome: | The proposed method improves retrieval effectiveness and accuracy across settings. |
Copied to clipboard
| Challenge: | Existing methods for achieving this alignment involve employing reinforcement learning from human feedback (RLHF) Existing approaches involve using RLHF to fine-tune LLMs based on human labels . however, RLRF is susceptible to instability during fine- tuning and presents challenges in implementation. |
| Approach: | They propose to use reinforcement learning from human feedback to fine-tune large language models with human preferences to achieve precise control of model behavior. |
| Outcome: | Experiments show that RAHF can be used to capture and manipulate representations to align with a broad spectrum of human preferences or values rather than being confined to a single concept or function. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have enabled strong performance in long-form writing, but current training paradigms remain limited. |
| Approach: | They propose an Adaptive Curriculum Reinforcement Learning framework to advance long-form writing capabilities beyond SFT. |
| Outcome: | Experiments on 7B-scale writer models show that Writing-RL improves long-form writing performance over strong SFT baselines. |
Copied to clipboard
| Challenge: | Proximal Policy Optimization (PPO) is central to aligning Large Language Models with verifiable rewards. |
| Approach: | They propose a scalable algorithm that harmonizes sample efficiency with stability of outcome-based updates. |
| Outcome: | The proposed algorithm outperforms standard PPO and matches the performance of computation-heavy group-based methods. |
Copied to clipboard
| Challenge: | Existing classification models only consider the temporal variations of existing data . current models focus on English corpora, leaving time as domains unexplored . |
| Approach: | They propose a framework to generalize classifiers over time on four languages, English, Danish, French, and German. |
| Outcome: | The proposed framework can generalize classifiers over time on four languages, English, Danish, French, and German. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely used in retrieval-augmented generation (RAG) when retrieved contexts are noisy, incomplete, or heterogeneous, a single generation process often struggles to reconcile evidence effectively. |
| Approach: | They propose a multi-agent synthesis approach to retrieval-augmented generation that structures evidence processing into multiple role-specialized agents. |
| Outcome: | Experiments on four benchmarks show that MASS-RAG consistently improves performance over strong RAG baselines. |
Copied to clipboard
| Challenge: | Existing RAG paradigms often overlook the cognitive step of applying knowledge, leaving a gap between retrieved facts and task-specific reasoning. |
| Approach: | They introduce a module extension that integrates application-aware reasoning into the RAG pipeline. |
| Outcome: | Experiments show that RAG+ outperforms standard RAG variants and achieves gains of 3–5% in complex scenarios. |
Copied to clipboard
| Challenge: | Existing tabular anomaly detection methods focus on detecting anomalies based on data distribution without considering regulatory compliance. |
| Approach: | They propose a task that leverages regulations to detect anomalies in tabular data . they also develop three new datasets to address this task . |
| Outcome: | The proposed method outperforms baselines on three new datasets. |
Copied to clipboard
| Challenge: | Existing approaches to improve LLM reasoning are limited in complex domains and lack external grounding makes verifiers unreliable on computation-intensive tasks. |
| Approach: | They propose a framework that transforms reward modeling into a multi-turn, tool-augmented deliberative process. |
| Outcome: | The proposed framework surpasses state-of-the-art ORMs by 25.2% under parallel and sequential TTS. |
Copied to clipboard
| Challenge: | Recent results from large pretrained models show that many datasets are saturated and unlikely to detect further progress. |
| Approach: | They evaluate 29 datasets using predictions from 18 pretrained Transformer models on individual test examples. |
| Outcome: | The proposed datasets are saturated and unlikely to detect future improvements. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can replicate insecure patterns from training data. |
| Approach: | They propose a framework that leverages distributed security-relevant cues by aggregating representations from multiple upper layers via an attention-based module. |
| Outcome: | Experiments show that the framework improves the secure-and-correct generation rate by 11.9% over baselines. |
Copied to clipboard
| Challenge: | a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling . |
| Approach: | They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution. |
| Outcome: | The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data. |
Copied to clipboard
| Challenge: | a rapid development of Chinese large language models poses big challenges for efficient LLM evaluation. |
| Approach: | They propose an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety. |
| Outcome: | The evaluation platform OpenEval benchmarks Chinese LLMs across capability, alignment and safety. |
Copied to clipboard
| Challenge: | a strict sum-to-one constraint forces attention sinks on irrelevant tokens, while probability mass disperses as sequence lengths increase. |
| Approach: | They propose a sink-free attention mechanism that achieves ultra-sparsity and improved robustness at longer sequence lengths without the computational overhead of projection methods. |
| Outcome: | The proposed mechanism produces >99 % exact zeros and eliminates attention sinks while maintaining competitive performance on standard and long-context benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) lack understanding of multi-image and interleaved inputs due to the visual features encoded by frozen encoders before being fed into the LLM backbone. |
| Approach: | They propose a two phase paradigm to enable in-depth multimodal context fusion prior to feeding the features into LLMs. |
| Outcome: | The proposed paradigm boosts the performance on 7 multi-image scenarios, contributing to increments on average accuracy by 2.13% and 7.60% against strong MLLMs baselines with 3B and 11B LLMs, respectively. |
Copied to clipboard
| Challenge: | Fine-tuning-as-a-service exposes models to harmful fine-tuneing attacks . however, inherent general adaptability of LLMs allows them to bypass selective unlearning by rapidly relearning or repurposing their general capabilities for harmful tasks. |
| Approach: | They propose a paradigm shift that inducing model collapse instead of selective removal by relearning or repurposing general capabilities for harmful tasks. |
| Outcome: | The proposed model collapse mechanism neutralizes the very general capabilities that attackers exploit, tackling the core issue unaddressed by selective unlearning. |
Copied to clipboard
| Challenge: | Masked diffusion models (MDMs) leverage bidirectional attention and a denoising process. |
| Approach: | They investigate the attention behaviors of Masked diffusion models by revealing the phenomenon of Attention Floating. |
| Outcome: | The proposed model doubles the performance of autoregressive models in knowledge-intensive tasks. |
Copied to clipboard
| Challenge: | Existing approaches to generating adversarial perturbations scale up the cost of training computational complexity by the number of gradient steps it takes to obtain the adversarials. |
| Approach: | They propose a flood method which aims at better generalization and a criterion to bring hyper-parameter-dependent flooding into effect with a narrowed-down search space by measuring how the gradient steps taken within one epoch affect the loss of each batch. |
| Outcome: | The proposed method improves BERT’s resistance to textual adversarial attacks by a large margin and achieves state-of-the-art robust accuracy on various text classification and GLUE tasks. |
Copied to clipboard
| Challenge: | Text-Centric Visual Question Answering (TEC-VQA) is a text-centric visual task understanding tool. |
| Approach: | They introduce a benchmark that features human expert annotations across 9 languages . they prioritize the text in question-answer pairs while disregarding visual text in images . |
| Outcome: | The proposed benchmarks prioritize the text in question-answer pairs while disregarding visual text in images. |
Copied to clipboard
| Challenge: | a recent study shows that media influence opinion via the inclusion or omission of partisan events. |
| Approach: | They develop a latent variable-based framework to predict the ideology of news articles by comparing multiple articles on the same story and identifying partisan events whose inclusion or omission reveals ideology. |
| Outcome: | The proposed framework validates the existence of partisan event selection and detects partisan events and article ideology better than baselines. |
Copied to clipboard
| Challenge: | Existing methods to learn user and item representations from reviews are limited . existing methods learn user representations based on ratings given by users . |
| Approach: | They propose a hierarchical user and item representation model with three-tier attention to learn user and items from reviews for recommendation. |
| Outcome: | The proposed model can learn user and item representations from reviews on four benchmark datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit prompt leakage vulnerabilities, raising intellectual property and confidentiality concerns. |
| Approach: | They use probing techniques to capture LLMs’ intent-related internal representations and show that they internalize prompt leakage intents in their hidden states before generating tokens. |
| Outcome: | The proposed probes achieve 90%+ AUROC across all tested models, even when applied to new system prompts and attacks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit significant but subtle weaknesses, such as mistakes in instruction-following or coding tasks. |
| Approach: | They propose a framework to automatically expose weaknesses in Large Language Models (LLMs) they use three LLM-powered agents to perform comprehensive weakness identification . |
| Outcome: | The proposed framework shows that it is more effective than untargeted data augmentation methods like Self-Instruct to identify weaknesses in LLMs. |
Copied to clipboard
| Challenge: | Existing approaches to generate intelligent open-domain dialogue agents only consider auxiliary commonsense stored in pure text, ignoring grounding information from the external visual world. |
| Approach: | They propose a VIsual Commonsense enhanced dialogue generaTOR that exploits auxiliary commonsense from images related to context to generate coherent and informative responses. |
| Outcome: | The proposed method outperforms the latest competitive methods in terms of coherence and diversity on two public datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for insurance claims adjudication are limited to information retrieval or simple multiple-choice setups. |
| Approach: | They propose a benchmark that provides complete reasoning traces linking factual inputs, relevant policy clauses, and final verdicts. |
| Outcome: | The proposed benchmark shows that models often produce correct decisions but fail to provide precise justifications, highlighting a critical discrepancy between decision accuracy and logical reasoning capabilities. |
Copied to clipboard
| Challenge: | Recent state-of-the-art (SOTA) effective neural network methods have been used in Chinese word segmentation (CWS) However, the robustness of the previous neural methods is limited by the large-scale annotated corpus. |
| Approach: | They propose a self-supervised Chinese word segmentation approach with a straightforward and effective architecture. |
| Outcome: | The proposed approach outperforms previous methods on 9 different CWS datasets with single criterion training and multiple criteria training and achieves better robustness. |
Copied to clipboard
| Challenge: | Prompt tuning has been proven to be successful on various tasks by incorporating a small number of trainable parameters while freezing large pre-trained language models. |
| Approach: | They propose a token-wise prompt tuning method that uses a bank of finer-grained soft prompt tokens to generate an instance-dependent prompt. |
| Outcome: | The proposed method performs far better than full parameter fine-tuned models and achieves state-of-the-art by tuning only 0.035% parameters on 14 datasets. |
Copied to clipboard
| Challenge: | Existing compression methods suffer from bottleneck issues when compression ratio is increased. |
| Approach: | They propose a novel approach to combine low-rank decomposition and quantization methods to reduce the compression bottleneck. |
| Outcome: | The proposed method reduces the computational and memory overhead of existing methods while maintaining model accuracy. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) has seen great success during recent years. |
| Approach: | They propose a metric that measures the fidelity of explanation methods on translation tasks . they use an efficient approximation to evaluate several explanation methods . |
| Outcome: | The proposed metric is efficient and can be used on translation tasks. |
Copied to clipboard
| Challenge: | Existing methods for singing voice synthesis are limited to fine-grained music scores . manual adjustment destroys regularity of note durations, making fine-grain music scores "crushed" |
| Approach: | They propose a method to synthesize singing voices given realistic music scores . they use real-music-score-based Singing Voice Synthesis to generate high-quality voices . |
| Outcome: | The proposed method eliminates manual annotation and simplifies phoneme-level mel-note alignment. |
Copied to clipboard
| Challenge: | Existing beam retrieval frameworks for multi-hop question answering were customized for two-hop questions and were poorly supervised. |
| Approach: | They propose an end-to-end beam retrieval framework for multi-hop question answering . they combine an encoder and two classification heads to optimize the retrieval process . |
| Outcome: | The proposed framework improves on MuSiQue-Ans and surpasses all previous retrievers on HotpotQA and achieves 99.9% precision on 2WikiMultiHopQA. |
Copied to clipboard
| Challenge: | Text-to-image (T2I) generation models have advanced in recent years, but effective interaction with these models is challenging for average users due to the need for specialized prompt engineering knowledge and the inability to perform multi-turn image generation. |
| Approach: | They propose to use off-the-shelf MLLMs and T2I models to build a multi-modal interactive dialogue system (MIDS) that can generate correct output modalities and coherence of output images. |
| Outcome: | The proposed pipeline can generate correct output modalities and coherent multi-modal outputs compared with other state-of-the-art models. |
Copied to clipboard
| Challenge: | Emerging AI-powered writing assistants focus on grammar fixes or simulating peer review with final scores, yet they fall short of providing concrete, actionable suggestions that help students improve their papers during drafting. |
| Approach: | They propose a human-centered writing assistant system that delivers actionable suggestions as Overleaf-native inline comments while leaving the actual writing entirely to human authors. |
| Outcome: | The proposed system outperforms a baseline with the skill library and provides actionable suggestions while leaving the actual writing to human authors. |
Copied to clipboard
| Challenge: | coding scaffolds that follow heterogeneous instructions remain under-examined in software engineering . coding models are capable software agents, but their ability to follow constraints remains under-explored . |
| Approach: | They introduce OctoBench, which benchmarks scaffold-aware instruction following in agentic coding. |
| Outcome: | The proposed benchmark aims to accelerate the development of more scaffold-aware agents. |
Copied to clipboard
| Challenge: | Chinese spelling check (CSC) tasks require that incorrect characters are usually similar to the correct ones in either phonetics or glyph. |
| Approach: | They propose a plug-and-play decoding intervention with similarity of characters module for Chinese spelling check (CSC) they propose to incorporate phonetic and glyph similarities only during the inference phase. |
| Outcome: | The proposed method significantly improves Chinese spelling check models on benchmarks and on benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods for extracting constituency trees from language models suffer from branching bias. |
| Approach: | They propose to measure the branching bias by comparing the performance gap on a language and its reversed language. |
| Outcome: | The proposed method is agnostic to language models and extracting methods, and it can be implemented with three factors to introduce the branching bias. |
Copied to clipboard
| Challenge: | Large language models are increasingly employed to empower autonomous agents to simulate human behavior. |
| Approach: | They propose to evaluate LLM-driven agents through multi-turn interactions using a bottom-up approach to create diverse social scenarios constructed from extensive scripts. |
| Outcome: | The proposed model evaluates LLM-driven agents through multi-turn interactions emphasizing goal completion and implicit reasoning. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded conversation models lack knowledge that occurs in training data, resulting in incomplete knowledge generation. |
| Approach: | They propose an Entity-Agnostic Representation Learning method to introduce knowledge graphs to informative conversation generation using context of conversations and relational structure of knowledge graph. |
| Outcome: | The proposed model generates more informative, coherent, and natural responses than baseline models. |
Copied to clipboard
| Challenge: | Existing prompting methods have been used to enhance multistep reasoning capabilities of large language models, but they have overlooked the potential of formulating higher-quality problems. |
| Approach: | They propose a method that starts from the problem side and refines problems to be more comprehensible and solvable for models. |
| Outcome: | The proposed method achieves notable and consistent effectiveness on five reasoning benchmarks across different models. |
Copied to clipboard
| Challenge: | Version updates are an indispensable requirement for Large Language Models . a large learning rate in the first stage and a complete learning decay process are crucial for version updates of LLMs. |
| Approach: | They propose a learning rate path switching training paradigm for version updates of Large Language Models. |
| Outcome: | The proposed paradigm reduces training cost to 58% when training four versions of LLMs compared to PTFS and CPT . |
Copied to clipboard
| Challenge: | Current research on in-image machine translation focuses on synthetic data with simple background, single font, fixed text position, and bilingual translation. |
| Approach: | They propose an end-to-end model to handle the challenge of practical conditions in PRIM . they annotate a real-world one-line text image with complex background, fonts, diverse text positions . |
| Outcome: | The proposed model improves translation quality and visual effect compared to other models. |
Copied to clipboard
| Challenge: | In-battle commentary is an important component of live streaming of e-sports competitions and is applicable to a wide range of scenarios like combat information analysis and live streaming. |
| Approach: | They propose a generative system for in-battle real-time commentary in mobile MOBA games and propose 'transform' method to convert match statistics and utterances into consistent encoding space. |
| Outcome: | The proposed system is based on real-time match statistics and events and can be used for live streaming, e-sports commentary and combat information analysis. |
Copied to clipboard
| Challenge: | Long-context inference is crucial for advancing large language models, but its prefill speed remains a bottleneck. |
| Approach: | They propose an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed. |
| Outcome: | The proposed framework achieves speedups of 9.2, 4.2, and 1.6 without any degradation in performance. |
Copied to clipboard
| Challenge: | Existing methods for generating large language models rely on student-generated outputs, which introduce generation errors and misguide the distillation process. |
| Approach: | They propose a multi-granularity semantic revision method for LLM distillation that corrects errors using teacher-generated tokens and re-generates the sequence to minimize errors. |
| Outcome: | The proposed method reduces errors and misguides distillation on student models and improves consistency between teacher and student outputs. |
Copied to clipboard
| Challenge: | MPLSandbox is an out-of-the-box multi-programming language sandbox designed to provide unified and comprehensive feedback from compiler and analysis tools for Large Language Models (LLMs). |
| Approach: | They propose a multi-programming language sandbox that provides unified feedback from compilers and analysis tools for Large Language Models. |
| Outcome: | The proposed multi-language sandbox can provide comprehensive feedback from compilers and analysis tools for large language models (LLMs). |
Copied to clipboard
| Challenge: | Prior work has explored the selection of examples for in-context learning, neglecting the internal relationships between examples and exist an inconsistency between training and inference. |
| Approach: | They propose a sequential-aware method that leverages the LLM’s feedback on varying context, aiding in capturing inter-relationships and sequential information among examples. |
| Outcome: | Experiments on 23 NLP tasks show that Se2 surpasses baselines and achieves 42% relative improvement over random selection. |
Copied to clipboard
| Challenge: | Existing models only output short phrases or sentences, raising doubts about their practical usability. |
| Approach: | They propose a dataset focused on document-level model editing that aims to correct errors and outdated knowledge in Large language models (LLMs) they propose to use document-based model editing to improve model capabilities in real-world scenarios. |
| Outcome: | The proposed model editing task improves model capabilities in real-world scenarios and reduces the cost of retraining. |
Copied to clipboard
| Challenge: | Compared to existing benchmarks, FinanceReasoning provides three key advancements: (1) credibility; (2) comprehensiveness; (3) numerical precision; (4) complexity; (5) complexity; and (6) complexity. |
| Approach: | They propose a benchmark to evaluate the reasoning capabilities of large reasoning models (LRMs) in financial numerical reasoning problems. |
| Outcome: | The proposed benchmark exceeds existing benchmarks in 67.8% of financial concepts and formulas and is credible, comprehensive, and challenging. |
Copied to clipboard
| Challenge: | Existing methods for learning incrementally do not address the problem of class-incremental learning. |
| Approach: | They propose a framework that can continuously learn new classes from a data stream without forgetting previously learned classes. |
| Outcome: | The proposed framework shows significant improvement over the state-of-the-art frameworks with up to 44.7% absolute F-score gain. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated better safety performance in high-resource languages than in low-resourced languages. |
| Approach: | They propose language-agnostic semantic alignment (LASA) which anchors safety alignment directly in semantic bottlenecks. |
| Outcome: | The proposed approach significantly improves safety across all languages: average attack success rate drops from 24.7% to 2.8% on LLaMA-3.1-8B-Instruct and remains within 3–4% across Qwen2.5 and Qwend3 Instruct models (7B–32B). |
Copied to clipboard
| Challenge: | Existing approaches to matching use Large Language Models as feature extractors, underutilizing their full modeling capabilities. |
| Approach: | They propose a matching paradigm that integrates two-tower, single-towing, and generative tasks within a unified LLM framework via attention-mask partitioning. |
| Outcome: | The proposed model achieves superior performance and strong practical value in an industrial search engine. |
Copied to clipboard
| Challenge: | Existing document-level relation extraction methods are sparse in relational entity pairs and the representation of entity pairs is insufficient. |
| Approach: | They propose a Pair-Aware and Entity-Enhanced(PAEE) model to solve two challenges . they propose predicting potential relational entity pairs and assembling directional entity pairs . |
| Outcome: | The proposed model can obtain state-of-the-art performance on four benchmark datasets . it can predict potential relational entity pairs and assemble directional entity pairs . |
Copied to clipboard
| Challenge: | Existing multimodal Retrieval-Augmented Generation (RAG) systems retrieve evidence at coarse granularities, making failures unverifiable. |
| Approach: | They propose a multimodal benchmark that features real-world landmarks with annotations across multiple viewpoints and a framework that treats visual elements as first-class retrieval units through three stages: element-level detection and classification, multi-granularity cross-modal alignment for evidence retrieval, and attribution-constrained generation. |
| Outcome: | The proposed framework achieves up to 29.2% improvement over six strong baselines for this task. |
Copied to clipboard
| Challenge: | Recent work has aimed to enhance reasoning capabilities of language models, but these methods are limited to domains with objectively verifiable answers. |
| Approach: | They propose a self-play framework to improve reasoning on general-domain data. |
| Outcome: | Experiments show that the proposed framework improves reasoning performance on general-domain data while maintaining competitive performance on verifiable academic benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to combine language modeling and knowledge graphs (KG) lack the context to provide a more precise understanding of the concepts. |
| Approach: | They propose to use external entity descriptions to provide contextual information for commonsense question answering models. |
| Outcome: | The proposed model achieves state-of-the-art among non-generative models in OpenBookQA and is the first of its kind in the field. |
Copied to clipboard
| Challenge: | Experimental results show that understanding attributes of mentions from text descriptions and visual images plays a vital role in multimodal entity linking. |
| Approach: | They propose to integrate attributes into multimodal entity linking using a text-image-based knowledge base. |
| Outcome: | The proposed approach integrates attributes into disambiguation. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) has become a standard paradigm for grounding Large Language Models (LLMs) however, performance degrades substantially when faced with noisy, outdated, or conflicting retrieved information. |
| Approach: | They propose a framework that explicitly elicits the model’s parametric knowledge as prior information to guide reasoning on retrieved documents. |
| Outcome: | The proposed framework achieves robust performance across varying degrees of external inconsistency and noise. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly being considered for high-stakes decision-making, yet their application in statistical risk analysis remains largely underexplored. |
| Approach: | They propose a method for extracting key information from raw data and translating it into structured contextual input within the LLM prompt. |
| Outcome: | The proposed approach significantly improves the LLM’s performance in risk assessment tasks. |
Copied to clipboard
| Challenge: | Existing open-domain question answering approaches follow a two-stage paradigm retriever then reader. |
| Approach: | They propose a novel reader-based generative approach that incorporates extractive and generative readers. |
| Outcome: | The proposed model improves on two benchmark datasets, Natural Questions and TriviaQA. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) do not align with psychiatric diagnostic protocols. |
| Approach: | They propose a framework that transforms the Mini International Neuropsychiatric Interview into automatic computational workflows through coordinated multi-agent collaboration. |
| Outcome: | The proposed framework transforms the gold-standard Mini International Neuropsychiatric Interview (MINI) into automatic computational workflows through coordinated multi-agent collaboration. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
Copied to clipboard
| Challenge: | Knowledge-grounded dialogue systems can generate informative responses based on dialogue history and external knowledge source. |
| Approach: | They conduct a thorough experiment to determine the optimal knowledge form, mutual effects between knowl- edge and model selection, and the few-shot performance of knowledge. |
| Outcome: | The proposed method combines knowledge-grounded dialogue with human-generated dialogues to generate informative and meaningful responses. |
Copied to clipboard
| Challenge: | Quantization studies have focused on instruction-tuned LLMs, leaving their performance on other benchmarks unclear. |
| Approach: | They propose a framework to evaluate quantized large language models using four dimensions . they propose to reduce the bits needed for model weights or activations with minimal performance loss . |
| Outcome: | The proposed framework can retain comparable performance to non-quantized LLMs on most benchmarks. |
Copied to clipboard
| Challenge: | Existing studies focus on detecting machine-generated text in open-source models, but their performance on closed-source large models is limited. |
| Approach: | They propose a method to detect rewritten text from large language models using a BERT encoder and propose to refine it to achieve semantic alignment. |
| Outcome: | The proposed method outperforms baseline methods on three text-generated datasets. |
Copied to clipboard
| Challenge: | Existing studies use legal facts to predict judgments, but legal facts are difficult to obtain in early stages of litigation. |
| Approach: | They propose a legal fact prediction task that takes evidence from trial as input to make predictions in the absence of ground-truth legal facts. |
| Outcome: | The proposed task can predict court rulings without ground-truth legal facts . the first benchmark dataset, LFPBench, is used to evaluate the task . |
Copied to clipboard
| Challenge: | Large language models demonstrate remarkable capabilities across various domains, including mathematics and logic reasoning. |
| Approach: | They propose a physics-based reasoning benchmark that includes physics theorems and constraints and a Physics Solution Auto Scoring Framework to evaluate physics based reasoning in large language models. |
| Outcome: | The proposed framework enables models to achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.99%). |
Copied to clipboard
| Challenge: | Existing methods to accelerate large language model inference have a fundamental limitation: candidates at the same tree layer share identical feature representations, constraining diversity and diminishing overall effectiveness. |
| Approach: | They propose a decoupled mixture of experts (MoE) into a draft model to generate diverse tokens from distinct feature spaces. |
| Outcome: | The proposed approach achieves significant speedups over strong baselines, with notable improvements in non-greedy scenarios where token diversity is crucial. |
Copied to clipboard
| Challenge: | retrieval-augmented generation (RAG) is a powerful tool for NLP applications . but it is challenging to encode large knowledge bases as compact offline structures . |
| Approach: | They propose a coarse-to-fine hierarchical graph inference method that uses random walks to retrieve information from a corpus of documents. |
| Outcome: | The proposed method reduces offline indexing costs and accelerates retrieval. |
Copied to clipboard
| Challenge: | rapid development of artificial intelligence (AI) technologies has inspired researchers to explore how AI can accelerate and enhance research. |
| Approach: | They organize the relevant studies into three main categories: hypothesis formulation, hypothesis validation, and manuscript publication. |
| Outcome: | The authors summarize the current state of research in three main areas: hypothesis formulation, hypothesis validation, and manuscript publication. |
Copied to clipboard
| Challenge: | Effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. |
| Approach: | They propose a multi-dimensional benchmark for evaluating LLMs’ alignment in Chinese with 8 main categories, 683 real-scenario rooted queries and corresponding human verified references. |
| Outcome: | The benchmark uses a human-in-the-loop data curation pipeline, 683 real-scenario rooted queries and human verified references. |
Copied to clipboard
| Challenge: | Existing Language Agents rely on a fixed mechanism or a set of mechanisms activated in a predefined order, limiting their adaptation to varied potential task solution structures. |
| Approach: | They propose to use language agents to learn to activate different mechanisms without relying on expert models to optimize their adaptation to different task solutions. |
| Outcome: | The proposed approach improves agent performance by enabling it to activate the appropriate mechanisms according to the potential characteristics of the task. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated proficiency in understanding and generating human natural languages. |
| Approach: | They propose a framework for scaling large language models using supervised fine-tuning, RLxF and test-time compute methodologies. |
| Outcome: | The proposed model can be used to understand and generate human natural languages. |
Copied to clipboard
| Challenge: | Experimental results show that Generative adversarial networks sacrifice sample diversity for quality and speed, while diffusion models exhibit outperformed sample quality and diversity at a high computational cost. |
| Approach: | They propose to combine GANs and diffusion probabilistic models to achieve better sample quality and diversity. |
| Outcome: | The proposed models outperform GANs and diffusion models in speech synthesis . the proposed models enjoy an efficient 4-step sampling process and exhibit better sample diversity . |
Copied to clipboard
| Challenge: | Existing studies show that some parameters in pre-trained language models can be pruned away without severe accuracy degradation. |
| Approach: | They propose a method which generates more features with very cheap operations from the remaining features and can be applied to unpruned BERT models to enhance their performance. |
| Outcome: | Empirical results on the GLUE benchmark on three backbone models (i.e., BERT, RoBERTa and ELECTRA) verify the efficacy of the proposed method. |
Copied to clipboard
| Challenge: | Existing approaches impose fixed cognitive structures that enhance performance in specific tasks but lack adaptability across diverse scenarios. |
| Approach: | They propose a test-time scaling framework based on meta-thoughts to improve performance . meta-thinkts are adaptive thinking strategies tailored to a given task . |
| Outcome: | Experimental results show that MetaScale outperforms standard inference approaches . it can scale more effectively with increasing sampling budgets and produces more structured responses . |
Copied to clipboard
| Challenge: | Existing methods for entity alignment fail to account for heterogeneity among KGs and distinction between KG entities and relations. |
| Approach: | They propose a Relation-gated Heterogeneous Graph Network (RHGN) that uses a relation-gate based convolutional layer to distinguish relations and entities in the KG. |
| Outcome: | Extensive experiments on four datasets show that the proposed method is superior to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive reasoning capabilities in financial tasks, but struggle with multi-step, goal-oriented scenarios in interactive financial markets. |
| Approach: | They propose a framework that integrates large language models with gradient-driven reinforcement learning (RL) policy optimization. |
| Outcome: | The proposed framework improves performance in trading and other financial domain tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to evaluate extremely long-context LLMs or analyze their limitations. |
| Approach: | They propose a Synthetic, Scalable, Systematic evaluation suite for LLMs using SQL execution. |
| Outcome: | The proposed evaluation suite is able to scale text length and difficulty across scenarios and provides strong correlations with real-world benchmarks. |
Copied to clipboard
| Challenge: | Experience-driven self-evolution has emerged as a promising paradigm for improving the autonomy of large language model agents, yet its reliance on self-curated experience introduces underexplored safety risks. |
| Approach: | They investigate how experience accumulation and utilization in self-evolving agents affect safety performance across web-based and embodied environments. |
| Outcome: | The findings expose inherent limitations of current self-evolving agents and call for more principled strategies to ensure safe and reliable adaptation. |
Copied to clipboard
| Challenge: | Simultaneous translation is notoriously dif- ficult due to word-order differences. |
| Approach: | They propose a prefix-to-prefix framework that implicitly learns to anticipate in a single translation model. |
| Outcome: | The proposed framework achieves low latency and reasonable qual- ity on 4 directions. |
Copied to clipboard
| Challenge: | Privacy-sensitive users require deploying large language models within their own infrastructure (on-premises) vulnerabilities in local environments can lead to unauthorized access and potential model theft. |
| Approach: | They propose a framework that secures a few bottom layers in a secure environment . they propose metric to optimize trade-off between protection and customization flexibility . |
| Outcome: | The proposed framework outperforms baselines on five models with 1.3B to 70B parameters. |
Copied to clipboard
| Challenge: | Several pre-training models of different modalities are showing a rising trend of homogeneity in their model structures. |
| Approach: | They propose a toolkit that supports pre-training models of different modalities. |
| Outcome: | The proposed toolkit can match the performance of the original implementations on text, vision, and audio benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to reward modeling in reinforcement learning tasks are limited when dealing with ambiguous preferences. |
| Approach: | They propose to use AAM to dynamically calibrate preference margins using the Bradley-Terry model's internal parameter knowledge to improve reward modeling in subjective tasks. |
| Outcome: | The proposed approach improves reward modeling by dynamically calibrating preference margins using the model’s internal parameter knowledge. |
Copied to clipboard
| Challenge: | Existing benchmarks on linguistic acceptability have been used to evaluate language models' ability to distinguish between acceptable and unacceptable sentences. |
| Approach: | They present the largest benchmark to date on linguistic acceptability: MELA . they establish LLM baselines on this benchmark and investigate cross-lingual transfer in acceptability judgements with XLM-R. |
| Outcome: | The proposed model outperforms open-source models on cross-lingual transfer in acceptability judgements. |
Copied to clipboard
| Challenge: | Extensive event extraction research has been conducted in many domains, including news, finance, and biology. |
| Approach: | They propose an end-to-end scientific event extraction framework for encoding nuggets into a grid matrix and simplifying complex event extraction as a nuggot-based grid modeling task. |
| Outcome: | The proposed framework performs well in scientific domain, demonstrating state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing studies employ autoregressive translation (AT) methods to encode sentences . however, the AT methods struggle with error accumulation when the length of sentences increases. |
| Approach: | They propose a context-aware non-autoregressive framework with the sentence-aligned connectionist temporal classification loss for document-level neural machine translation. |
| Outcome: | The proposed framework achieves 46X speedup on three benchmarks compared to strong baselines. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for natural language generation (NLG) tasks face the challenges on generalization ability and interpretability. |
| Approach: | They propose a metric that evaluates natural language generation tasks as an instruction-style question answering task and utilizes instruction-tuned pre-trained language models without training on evaluation datasets. |
| Outcome: | The proposed metric achieves state-of-the-art performance in untrained metrics for evaluating text summarization and dialogue generation, which exhibits strong dimension-level / task-level generalization ability and interpretability. |
Copied to clipboard
| Challenge: | Spotlighter is a lightweight token-selection framework that enhances accuracy and efficiency in prompt tuning. |
| Approach: | They propose a token-selection framework that enhances accuracy and efficiency in prompt tuning by preserving only the top-scoring tokens for downstream prediction. |
| Outcome: | The proposed framework outperforms CLIP by up to 11.19% in harmonic mean accuracy and achieves 0.8K additional FPS, with only 21 extra parameters. |
Copied to clipboard
| Challenge: | Existing approaches to prompt optimization trade off signal quality against computational cost. |
| Approach: | They propose a framework that uses a first-order gradient approximation to score segment importance in a continuous masking direction. |
| Outcome: | The proposed framework improves efficiency and robustness by using a first-order gradient approximation to score segment importance in a continuous masking direction. |
Copied to clipboard
| Challenge: | a recent HCI study has pointed to gaps in machine storytelling ability at the global level . authors show that LLMs have less suspense and less tension than human stories . |
| Approach: | They propose a computational framework to analyze narratives through three discourse-level aspects. |
| Outcome: | The proposed framework analyzes narratives through three discourse-level aspects . it shows that LLMs fall short of human abilities in discourse understanding . |
Copied to clipboard
| Challenge: | Large language models (LLMs) suffer from severe hallucination issues due to the knowledge misalignment between the pre-training stage and the supervised fine-tuning stage. |
| Approach: | They propose a training objective with an abstention mechanism that selectively rejects tokens that misalign with the desired knowledge distribution via a special [REJ] token. |
| Outcome: | The proposed model selectively rejects tokens that misalign with the desired knowledge distribution via a special [REJ] token. |
Copied to clipboard
| Challenge: | Recent researches focus on deep learning and reinforcement learning for multi-turn information seeking conversation systems. |
| Approach: | They propose an efficient and effective multi-turn conversation model based on convolutional neural networks and extend it to adapt the knowledge learned from a resource-rich domain to enhance the performance. |
| Outcome: | The proposed model performs better than the existing model on an industrial chatbot called AliMe Assist. |
Copied to clipboard
| Challenge: | Vision-Language Models struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. |
| Approach: | AgentThink integrates Chain-of-Thought reasoning with dynamic, agent-style tool invocation for autonomous driving tasks. |
| Outcome: | Experiments on the DriveLMM-o1 benchmark show AgentThink significantly boosts overall reasoning scores by 53.91% and enhances answer accuracy by 33.54% . |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) reasoning has improved the performance of large language models (LLMs) however, the detailed reasoning process in CoT often incurs long generation times and high computational costs due to the inclusion of unnecessary steps. |
| Approach: | They propose a method to identify critical reasoning steps using perplexity as a measure of their importance. |
| Outcome: | The proposed method achieves a better balance between reasoning accuracy and efficiency of CoT. |
Copied to clipboard
| Challenge: | Existing datasets suffer from outdated and insufficient challenging content, neglecting human-like reasoning, and limited reliability due to single-LLM generation. |
| Approach: | They propose a human-in-the-loop, multi-agent data generation framework that integrates reasoning-dense filters, multiagent collaboration, and human mathematicians’ evaluations to ensure the reliability and quality of the dataset. |
| Outcome: | The proposed framework improves accuracy and quality of the 2,000-synthesized datasets by integrating reasoning-dense filters, multi-agent collaboration, and human mathematicians’ evaluations. |
Copied to clipboard
| Challenge: | Role-playing Agents (RPAs) struggle to recognize and respond to hard queries that conflict with their role-play knowledge. |
| Approach: | They propose a lightweight representation editing approach that conveniently shifts conflicting requests to the rejection region, thereby enhancing the model’s refusal accuracy. |
| Outcome: | The proposed model improves RPAs’ refusal ability of conflicting requests while maintaining their general role-playing capabilities. |
Copied to clipboard
| Challenge: | Document retrieval techniques are used to compute semantic similarity between a query and documents, but the scalar similarity fails to reflect enough information, hindering the interpretation of retrieval results. |
| Approach: | They propose a method which improves the global document-query similarity through contrastive learning and integrates well-designed fusion and decoding modules. |
| Outcome: | The proposed method improves the global document-query similarity through contrastive learning and integrates well-designed fusion and decoding modules. |
Copied to clipboard
| Challenge: | Adaptive policies can balance translation quality and latency based on context information . previous methods on obtaining adaptive policies rely on complicated training process . |
| Approach: | They propose to obtain adaptive policies by a simple heuristic composition of fixed policies . they propose to use a heurism to obtain policies that can outperform fixed ones . |
| Outcome: | Experiments on Chinese -> English and German -> english show that adaptive policies outperform fixed policies by up to 4 BLEU points for the same latency. |
Copied to clipboard
| Challenge: | Existing methods for continual relation extraction (CRE) excel in preserving old knowledge but falter when confronted with contaminated data streams. |
| Approach: | They propose a noise-resistant contrastive framework for continual relation extraction (CRE) that preserves old knowledge while learning incremental corrupted relations. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on various benchmarks with increasing noise rates. |
Copied to clipboard
| Challenge: | Sequence Labeling (SL) is a long-standing field of natural language processing. |
| Approach: | They propose a framework that utilizes a conditional discrete diffusion model for generating discrete tag data. |
| Outcome: | The proposed framework outperforms gpt-3.5-turbo on multiple benchmark datasets and tasks. |
Copied to clipboard
| Challenge: | Existing methods for word-level segmentation (CWS) for the Chinese language have been successful in large-scale annotated corpora. |
| Approach: | They propose a method that integrates different segmentation criteria into one model . they use a transfer learning method to improve the performance of OOV words . |
| Outcome: | The proposed method achieves state-of-the-art performance on multiple benchmark datasets . it shows a competitive practicability and generalization ability for the CWS task . |
Copied to clipboard
| Challenge: | Existing approaches to balancing translation quality and latency are either too aggressive or too conservative. |
| Approach: | They propose an opportunistic decoding technique that always (over-)generates a certain mount of extra words at each step to keep the audience on track with the latest information. |
| Outcome: | The proposed technique reduces latency and increases BLEU with no over-generating . it also corrects mistakes in the overgenerated words when observing more context . |
Copied to clipboard
| Challenge: | Existing natural language processing systems are vulnerable to noisy inputs resulting from misspellings. |
| Approach: | They propose a stand-alone spelling correction problem that corrects the spelling of tokens without additional token insertion or deletion. |
| Outcome: | The proposed solution outperforms the state-of-the-art spelling correction model by 12.8% absolute F0.5 score. |
Copied to clipboard
| Challenge: | Existing models exhibit inconsistent reasoning abilities across different languages . existing models lack consistency across languages due to imbalance of training data . |
| Approach: | They propose a multilingual alignment-as-preference optimization framework to align reasoning processes in other languages with the dominant language. |
| Outcome: | The proposed framework improves multilingual reasoning across languages on three benchmarks. |
Copied to clipboard
| Challenge: | Existing methods rely on a single high-capacity LLM to represent an entire population of diverse learners. |
| Approach: | They propose an ability-aware student simulation framework that matches students with appropriate LLM backbones through cognitive alignment. |
| Outcome: | The proposed framework significantly reduces simulation bias and outperforms single-model baselines across the entire proficiency spectrum. |
Copied to clipboard
| Challenge: | Existing research on LLM biases has focused on direct questioning or general-purpose settings . pronounced behavioral biase despite their growing deployment in financial analysis, forecasting, and decision support. |
| Approach: | They propose a benchmark to evaluate behavioral biases of large language models in MFMD . they use a multilingual financial misinformation dataset to integrate these with misinformation claims . |
| Outcome: | The proposed benchmark evaluates behavioral biases of large language models across economic scenarios. |
Copied to clipboard
| Challenge: | Multimodal reasoning is a key capability for large vision-language models . however, the vanilla Chain-of-Thought method fails to address critical steps in multi-step reasoning tasks. |
| Approach: | They propose a bi-modal Behavioral Alignment method to augment multimodal reasoning . they use domain-specific language to integrate multimodal information into a precise alternative form . |
| Outcome: | The proposed method significantly improves GPT-4V(ision) on geometry problem solving, chess positional advantage prediction and molecular property prediction. |
Copied to clipboard
| Challenge: | Recent approaches to quantization of Large Language Models (LLMs) have been widely adopted due to activation outliers, which degrade model performance especially at lower bit precision. |
| Approach: | They propose a new metric for quantization that strategically distributes outlier magnitudes across matrix dimensions via optimized diagonal operations. |
| Outcome: | The proposed framework achieves less than 1% accuracy drop in W4A4 quantization on the LLaMA-3-8B model and reduces the performance gap by 39.1% on the more challenging W2A4KV16 model. |
Copied to clipboard
| Challenge: | Current approaches to simultaneous speech-to-speech translation accumulate more and more latencies in later sentences when the speaker talks faster. |
| Approach: | They propose a method which generates more fluent target speech latency than the baseline . they propose to use self-adaptive translation to adjust the length of translations to accommodate different source speech rates. |
| Outcome: | Xiong et al., 2019) show that the proposed method generates more fluent target speech latency than baseline . authors say it provides more natural communication process than speech-to-text translation . xiong and colleagues say the proposed technique is more efficient than current approaches . |
Copied to clipboard
| Challenge: | Existing methods to generate text in mental health are limiting, but they are effective for many tasks. |
| Approach: | They propose a task-adaptive tokenizer that allows for the integration of task-specific tokens into the pre-trained model's tokenization step. |
| Outcome: | The proposed tokenization approach improves generation performance on psychological question-answering tasks in Chinese and English while using 60% fewer tokens. |
Copied to clipboard
| Challenge: | Recent studies observe a phenomenon where reward models achieve high accuracy on static datasets but fail to generalize effectively during RLHF. |
| Approach: | They propose a method that combines rationale consistency with outcome accuracy to improve performance on RM-Bench and JudgeBench. |
| Outcome: | The proposed method surpasses baselines on RM-Bench and JudgeBench by an average of 5% and improves creative writing tasks by 7%. |
Copied to clipboard
| Challenge: | Medical-specific Large Language Models (LLMs) have demonstrated impressive performance on medical-related exams and tasks. |
| Approach: | They propose a framework for medical conversational data generation that uses Authentic Seed Data to ensure quality of the data. |
| Outcome: | The proposed model outperforms all baselines and human evaluations, and aligns with human preferences and clinical demands. |
Copied to clipboard
| Challenge: | Existing systems that provide detailed, constructive feedback on academic papers struggle with review fidelity. |
| Approach: | They explore factors that underlie the development of robust advising systems . large language models have shown remarkable progress in tasks from text generation to code synthesis . |
| Outcome: | The proposed model outperforms general-purpose language models in acceptance rates for self-ranked top-30% submissions to ICLR 2025. |
Copied to clipboard
| Challenge: | Recent advances in latent diffusion models (LDMs) have markedly enhanced text-to-audio generation, yet their iterative sampling processes impose substantial computational demands, limiting practical deployment. |
| Approach: | They propose to learn straight flow for fast simulation by using flashAudio with rectified flows and immiscible flow to minimize the total distance of data-noise pairs in a batch vias assignment. |
| Outcome: | The proposed method can learn straight flow for fast simulations and reduce noise distribution. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have made significant progress in knowledge-intensive applications, but they may face a multi-stage continuous learning scenario. |
| Approach: | They propose a multi-stage continuous learning paradigm that includes a preference-based learning bias to identify potential knowledge conflicts and a self-distillation-based data augmentation strategy to expand and enrich the training corpus. |
| Outcome: | The proposed learning paradigm achieves a significant improvement in accuracy after 7 stages of fine-tuning compared to previous methods while preserving general knowledge. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit impressive emergent abilities in natural language processing, but their democratization is hindered due to huge computation requirements and closed-source nature. |
| Approach: | They propose a tailored learning approach to distill the exclusive reasoning ability to smaller LMs to facilitate democratization. |
| Outcome: | The proposed approach enables the democratization of the exclusive reasoning ability by leveraging the black-box model as a reasoning teacher. |
Copied to clipboard
| Challenge: | Existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols. |
| Approach: | They propose a benchmark to assess whether generated software meets user needs . they use a fine-grained set of user requirements and a fully automated testing pipeline . |
| Outcome: | E2EDev is a benchmark to assess whether generated software meets user needs through mimicking real user interactions. |
Copied to clipboard
| Challenge: | Existing augmentation techniques manipulate words in the original text that break the semantic coherence of the text, or exploit generative models that ignore preserving entities in the text. |
| Approach: | They propose a novel Entity-to-Text based data augmentation technique called EnTDA to add, delete, replace or swap entities in the original text. |
| Outcome: | The proposed technique generates semantically coherent and entity preserving texts on thirteen NER tasks and two settings. |
Copied to clipboard
| Challenge: | Existing methods for concept expansion in MOOCs are inefficient because of the diversity of MOOC courses and rapid updates. |
| Approach: | They propose an end-to-end hierarchical reinforcement learning (HRL) model for concept expansion in MOOCs that employs a two-level mechanism of seed selection and concept expansion. |
| Outcome: | The proposed model improves on nine real MOOC datasets and maintains competitive performance under different settings. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated the potential to mimic human social intelligence, but most studies focus on static self-report or performance-based tests. |
| Approach: | They propose a framework to assess LLMs' ability to understand and manage intentions by mapping their ability to infer the intentions of others in a game setting. |
| Outcome: | The proposed framework assesses LLMs' ability to understand and manage intentions in a game setting. |
Copied to clipboard
| Challenge: | Large language models produce powerful text embeddings, but their causal attention mechanism restricts the flow of information from later to earlier tokens, harming performance. |
| Approach: | They propose a method that prepending a single summary token to reduce attention-level compression by partitioning the input into blocks and prepending blocks to subsequent blocks. |
| Outcome: | The proposed method achieves consistent performance gains across 11 retrieval datasets and 30 general embedding benchmarks. |
Copied to clipboard
| Challenge: | Automated theorem proving (ATP) benchmarks focus on symbolic inference but rarely involve understanding complex number combination reasoning. |
| Approach: | They propose a benchmark that requires a model to reduce a trigonometric expression with step-by-step proof and evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
| Outcome: | The proposed benchmark evaluates a generative LM’s reasoning ability on formulas and ability to manipulate, group, and factor number terms. |
Copied to clipboard
| Challenge: | Large Language models (LLMs) have remarkable abilities in understanding complex texts . however, understanding misalignment leads to LLMs mistakenly translating complex concepts . |
| Approach: | They propose a translation process that aligns the translation-specific understanding with the general understanding to improve translation quality and reduce translation literalness. |
| Outcome: | The proposed translation process improves translation quality and reduces translation literalness by -25% -51%. |
Copied to clipboard
| Challenge: | Existing IMT systems relying on lexical constrained decoding (LCD) are limited in translation efficiency and quality due to LCD. |
| Approach: | They propose a novel interactive neural machine translation system that uses lexical constraints to decode missing words in a manually revised translation. |
| Outcome: | The proposed system performs significantly better and faster than state-of-the-art IMT on three translation tasks. |
Copied to clipboard
| Challenge: | Current LLMs are achieving better performance on various benchmarks, but their performance in practical applications does not always match their benchmark results. |
| Approach: | They propose to detect and rewrite leaked benchmarks without altering their difficulties by using Inference-Time Decontamination (ITD) to mitigate performance inflation caused by memorizing leaked samples. |
| Outcome: | The proposed method reduces inflated accuracy by 22.9% on GSM8K and 19.0% on MMLU. |
Copied to clipboard
| Challenge: | Using paralinguistic cues is challenging for speech large language models, authors say . limited training data, annotation difficulty, and models exploiting lexical shortcuts are challenges . a recent study shows that modeling paralinguistic reasoning with multitask RL improves paralinguistics understanding . |
| Approach: | They propose multi-task reinforcement learning with chain-of-thought prompting that elicits explicit affective reasoning. |
| Outcome: | The proposed model improves paralinguistics understanding over baselines and strong proprietary models by 8-12% on Expresso, IEMOCAP, and RAVDESS. |
Copied to clipboard
| Challenge: | Aspect Sentiment Triplet Extraction (ASTE) is a thriving research area . current code-switching methods suffer from term boundary detection issues and out-of-dictionary problems. |
| Approach: | They propose a test-time code-switching framework which bridges the gap between bilingual training and monolingual test- time prediction. |
| Outcome: | The proposed framework achieves an average improvement of 3.7% on four cross-lingual datasets. |
Copied to clipboard
| Challenge: | Current paradigms for empowering Large Language Models with multilingual capabilities rely heavily on massive instruction tuning. |
| Approach: | They propose a hybrid cross-alignment approach that fuses a frozen NLLB encoder with a Qwen decoder via a closed-loop dual-adapter architecture. |
| Outcome: | The proposed model outperforms towerPlus-9B and Aya-101 on language-agnostic projection protocols. |
Copied to clipboard
| Challenge: | Clinical trials are expensive and time-consuming, and inappropriately designed studies can be devastating in a pandemic. |
| Approach: | They propose a model that takes a PICO-formatted clinical trial proposal and predicts the outcome from it. |
| Outcome: | The proposed model outperforms baseline models on a benchmark dataset with 10.7% relative gain over BioBERT. |
Copied to clipboard
| Challenge: | Existing methods to address the named entity recognition problem are limited and lack explicit optimization specific to the task. |
| Approach: | They propose a prototype-based representation alignment model for a cross-lingual named entity recognition task using labeled source language data. |
| Outcome: | The proposed model outperforms existing state-of-the-art methods in some challenging scenarios. |
Copied to clipboard
| Challenge: | Recent studies explore Large Language Models’ (LLMs) performance on Theory of Mind (ToM) reasoning tasks, but research on ToM abilities that require more nuanced social context is limited, such as white lies. |
| Approach: | They propose a novel English benchmark to evaluate Large Language Models’ ability to understand white lies within real-life conversations and reason about prosocial motivations behind them. |
| Outcome: | The proposed model outperforms state-of-the-art models on ToM reasoning tasks and reveals significant gaps between humans and LLMs. |
Copied to clipboard
| Challenge: | Tool-calling agents are increasingly deployed in real-world customer-facing workflows . but most studies on tool-callers focus on idealized settings with general, fixed, and well-specified tasks. |
| Approach: | They propose a tool-calling agent-based data pipeline that converts trajectories into user-facing tasks with controlled intent adaptations. |
| Outcome: | The proposed pipeline can be used to study tool use under three scenarios. |
Copied to clipboard
| Challenge: | Existing methods for integrating spatial layouts with text have limitations . existing methods produce overly long text sequences or lack autoregressive traits of LLMs . |
| Approach: | They introduce Interleaving Layout and Text in a Large Language Model (LayTextLLM) they use OCR-derived text and spatial layouts to integrate with LLMs for document understanding . |
| Outcome: | The proposed model shows an increase in performance in KIE and VQA tasks. |
Copied to clipboard
| Challenge: | Existing methods for video-text retrieval capture fine-grained semantic concepts . however, they lack the ability to capture finer-grain concepts such as objects and actions. |
| Approach: | They propose a dual-encoder architecture for fast video-text retrieval that learns lexicon representations to capture fine-grained semantics. |
| Outcome: | The proposed framework outperforms existing methods with 4.8% and 8.2% improvement on MSR-VTT and DiDeMo respectively. |
Copied to clipboard
| Challenge: | Existing methods for enhancing text or data are limited by lack of logical connections between generated texts and training data. |
| Approach: | They propose an encoder-decoder data augmentation framework that combines large language models and chain-of-thought prompting to summarize texts into target-specific if-then rationales, establishing logical relationships. |
| Outcome: | The proposed framework significantly improves over state-of-the-art methods on benchmark datasets while enabling interpretable rationale-based learning. |
Copied to clipboard
| Challenge: | Existing research on generative AI security is driven by mutually reinforcing attack and defense methodologies grounded in empirical experience. |
| Approach: | They propose a new algorithm that uses a random sampling algorithm to control risk. |
| Outcome: | The proposed algorithm improves robustness and utility while maintaining latency comparable to existing algorithms. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown great potential to enhance Natural Language Processing (NLP) models in areas such as predictive accuracy, fairness, robustness, and explainability. |
| Approach: | They evaluate or improve generative Large Language Models from a causal perspective in areas such as reasoning capacity, fairness and safety issues, explainability, and handling multimodality. |
| Outcome: | The proposed models can be used to perform causal relationship discovery and causal effect estimation tasks. |
Copied to clipboard
| Challenge: | Existing adaptive testing methods face several challenges due to mechanized nature of most algorithms and noisy response data. |
| Approach: | They propose to use large language models to enhance adaptive testing through interactive engagement to capture test-takers’ responses and anomalies. |
| Outcome: | The proposed agent achieves more accurate results with 20% fewer questions than state-of-the-art baselines and testers preferred it in speed, smoothness, and other dimensions. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual dependency parsing rely on large corpus size and cost. |
| Approach: | They propose a cross-lingual dependency parsing approach based on word reordering . they propose to train a model that transfers knowledge learned in one or multiple languages to target languages . |
| Outcome: | The proposed approach outperforms the baseline approach in Hindi and Latin by 15.3% and 6.7%. |
Copied to clipboard
| Challenge: | Large Language Models have shown immense potential in multimodal applications, but convergence between textual and musical domains remains unexplored. |
| Approach: | They propose a system that aligns music representations with a frozen LLM . they train the system on an extensive music caption dataset and fine-tune it with instructional data . |
| Outcome: | The proposed system bridges the gap between music audio and textual contexts by combining music captions with a frozen model . it performs well in generating music caption and composing music-related Q&A pairs . the proposed system is available for free download at http://www.musilingo.com/ . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been widely adopted in real-world dialogue applications, but their robustness is criticized all along. |
| Approach: | They propose to use play-by-play text commentary to build a multi-turn athletic real-world scenario dialogue benchmark to evaluate three critical aspects of multi-turned conversations: ultra multi- turn, interactive multi-twist, and cross-turn tasks. |
| Outcome: | The proposed benchmarks outperform open-source LLMs on three critical aspects of multi-turn conversations: ultra multi-turned, interactive multi- turn, and cross-turn tasks. |
Copied to clipboard
| Challenge: | Existing methods focus on correcting the output but overlook the ability of LLMs to detect and correct misleading content in the input itself. |
| Approach: | They propose a three-stage fine-tuning method that improves LLMs' ability to detect and correct misleading information in input queries. |
| Outcome: | The proposed method improves accuracy and factuality of LLM responses while also reducing hallucinations. |
Copied to clipboard
| Challenge: | Speculative decoding is a widely used method that accelerates the generation process of large language models (LLMs) drafting efficiency has become a bottleneck in the final speedup of speculative drafting, therefore generating longer drafts at less cost can lead to better speedup. |
| Approach: | They propose a method that uses existing model to drafting and target LLM to verify draft in a low-cost parallel manner. |
| Outcome: | The proposed method can achieve speedups of up to 2.4 over speculative decoding and 3.9 over vanilla decoding without fine-tuning draft and target models. |
Copied to clipboard
| Challenge: | Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world . |
| Approach: | They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions. |
| Outcome: | The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have driven the rise of agentic workflows . yet, how can we attribute performance gains to individual upgrades and their interactions? |
| Approach: | They propose a game-theoretic framework that models component upgrades as players and evaluates component coalitions to compute Shapley values. |
| Outcome: | The proposed framework provides interaction-aware attribution and recommendation for model allocation under a fixed workflow structure. |
Copied to clipboard
| Challenge: | Existing methods for machine translation require intensive keyboard interaction, which is inconvenient on mobile devices. |
| Approach: | They propose a touch-based editing method that is more flexible than keyboard-mouse-based translation postediting. |
| Outcome: | The proposed method significantly outperforms existing interactive translation methods on translation datasets and on post-editing datasets. |
Copied to clipboard
| Challenge: | Existing evaluation of Large Language Models on static benchmarks is vulnerable to data contamination and leaderboard overfitting. |
| Approach: | LLMEval-Fair framework provides a framework for dynamic evaluation of Large Language Models . evaluators use a proprietary bank of 220k graduate-level questions to analyze model data . |
| Outcome: | LLMEval-Fair provides robust and credible evaluation framework for Large Language Models . it provides a strong empirical validation for the dynamic evaluation paradigm . |
Copied to clipboard
| Challenge: | Existing studies either overlook temporal shifts or hardly capture rich shifting patterns of both semantic and knowledge. |
| Approach: | They develop a temporal adaptive learning framework that captures temporal shifts . they use medical ontology and other knowledge sources to integrate temporal adaptation . |
| Outcome: | The proposed framework improves classification tasks across multiple domains and domains with knowledge integration. |
Copied to clipboard
| Challenge: | Existing data selection methods for instruction-following large language models rely on unreliable scores or use downstream tasks for selection. |
| Approach: | They propose a method that utilizes the VLM itself as a filter to select high-quality instruction-tuning data. |
| Outcome: | The proposed method can reach better results compared to full data settings with merely about 15% samples and can achieve superior performance against competitive baselines. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have advanced legal intelligence, but the scarcity of scenario data impedes the progress toward interactive legal scenarios. |
| Approach: | They propose a Multi-agent Legal Simulation Driver to generate synthetic data by simulating interactive legal scenarios. |
| Outcome: | The proposed framework ensures consistency of legal attributes between participants and introduces a supervisory mechanism to align participants’ characters and behaviors as well as addressing distractions. |
Copied to clipboard
| Challenge: | Parameter-efficient fine-tuning (PEFT) is a low-cost alternative to full fine-timing due to the massive overhead. |
| Approach: | They propose a Mixture-of-Experts approach that enhances specialization while maintaining low resource overhead. |
| Outcome: | The proposed approach outperforms or matches state-of-the-art methods on GLUE, GSM8K, MBPP, and a text rewriting task from SmolTalk. |
Copied to clipboard
| Challenge: | Existing approaches to adapt pre-trained models with parameters frozen are based on input-side adaptation, which requires thousands of API queries. |
| Approach: | They propose to train a model-as-a-service (MaaS) setting to provide only the inference APIs for users . they argue that input-side adaptation could be arduous due to the lack of gradient signals . |
| Outcome: | The proposed model outperforms state-of-the-art algorithms with a 200x speed-up. |
Copied to clipboard
| Challenge: | Existing methods employ sentence-level retrieval and fusion methods, which may lead to similarity bias and interference from irrelevant information in unstructured knowledge sentences. |
| Approach: | They propose a segment-level and category-oriented network to solve similarity bias problem by segmenting and prompting knowledge retrieval methods and a category-based grounding method. |
| Outcome: | The proposed model eliminates similarity bias and improves the overall performance of the KB-REC task. |
Copied to clipboard
| Challenge: | Non-collaborative dialogue agents are expected to engage in strategic conversations with diverse users, and this poses two main challenges for existing dialogue agents: 1) the inability to integrate user-specific characteristics into the strategic planning; 2) the difficulty of training strategic planners that can be generalized to diverse users. |
| Approach: | They propose to integrate a user-aware strategic planning module and a population-based training paradigm into a non-collaborative dialogue agent for securing a mutual agreement that leans favorably towards the system's objectives. |
| Outcome: | The proposed model can be used to achieve a mutual agreement that leans favorably towards the system's objectives. |
Copied to clipboard
| Challenge: | Existing methods for psychiatric interviewing degenerate into rigid interrogation or aimless chitchat due to a lack of strategic planning. |
| Approach: | They propose a framework for psychiatric interviewing grounded in Speech Act Theory that integrates a large-scale dataset with fine-grained psychic speech act annotations. |
| Outcome: | The proposed framework outperforms baselines in psychiatric interviewing. |
Copied to clipboard
| Challenge: | Recent progress in large language models is driven by scaling of training compute through pre-training with nexttoken prediction (NTP) or post-training (RL) Pre-training using NTP enables models to acquire extensive knowledge and skills from general data, but it suffers from data inefficiency and catastrophic forgetting in continual learning settings. |
| Approach: | They propose to scale training compute through pre-training with next-token prediction (NTP) or post-training by scaling reinforcement learning (RL) to improve learning from general data. |
| Outcome: | Experiments on multiple benchmarks and models show that the proposed approach improves continual pre-training and provides a strong foundation for post-training on Qwen3-8B-Base. |
Copied to clipboard
| Challenge: | Existing evaluation methods for text summarization systems are limited to in-domain setting, where supervised pre-trained models are evaluated on the same dataset. |
| Approach: | They propose to use a cross-dataset evaluation approach to evaluate different summarization systems in a multi-domain setting. |
| Outcome: | The proposed model can be used to evaluate text summarization systems on different datasets. |
Copied to clipboard
| Challenge: | Existing studies show that textual unlearning does not achieve comparable safety performance with image-text alignment. |
| Approach: | They propose to use textual unlearning to align MLLMs with image-text pairs to explain this problem . they construct a visual leakless safety bench with 2.2k image- text pairs to test this problem. |
| Outcome: | The proposed model can refuse image-text pairs according to textual queries, leading to unreliable safety evaluations. |
Copied to clipboard
| Challenge: | Existing solutions for supervised fine-tuning often lead to catastrophic forgetting, where models lose their previously acquired knowledge and general capabilities. |
| Approach: | They propose a self-distribution alignment method that aligns input sequence logits to preserve the model’s semantic distribution, thereby mitigating catastrophic forgetting and improving downstream performance. |
| Outcome: | The proposed method achieves a superior balance between downstream learning and general capability retention. |
Copied to clipboard
| Challenge: | Existing Mixture-of-Experts training frameworks use a micro-batch to calculate LBL . micro-batches are restricted to a single sequence, preventing expert specialization . |
| Approach: | They propose to use a global-batch to loosen the load balance constraint for MoEs models . they propose to synchronize fi across micro-batches and then use it to calculate the LBL . |
| Outcome: | The proposed global-batch LBL improves the domain specialization of experts . the micro-battery LBL is almost at the sequence level, and the router is pushed to distribute the token evenly . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly aligned with human preferences through Reinforcement Learning from Human Feedback (RLHF). |
| Approach: | a new study proposes a domain-informed self-consistency policy optimization extension to GRPO that addresses inter-group imbalance. |
| Outcome: | a new extension of GRPO addresses inter-group imbalance with two key innovations . the proposed method outperforms existing GR PO variants by 5% on Qwen3 models . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated strong capabilities across various domains, but their large-scale deployment faces a major obstacle: the high computational cost of long-sequence inference. |
| Approach: | They propose an algorithm that retains key-value vectors until they are no longer needed to solve reasoning tasks. |
| Outcome: | The proposed algorithm achieves high accuracy with O(L) time but O(N) memory complexities. |
Copied to clipboard
| Challenge: | Experiments on GPT and other 23 LLMs indicate that tokens widely exist while GPT’s vocabulary behaves the worst: more than 23% long Chinese tokens (i.e., a token with more than two Chinese characters) are either porn or online gambling. |
| Approach: | They propose to locate Polluted Chinese (PoC) tokens in LLMs and build a PoC token detector to label them in vocabularies by considering each token’s semantics and related contents from the search engines. |
| Outcome: | The proposed method predicts that the ratio of “*” related webpages in GPT-4o's training data is around 0.5%. |
Copied to clipboard
| Challenge: | Commercial-scale language models (LMs) have taken APR to unprecedented levels, but they are limited by parameters and humans interact with them through explicit prompts. |
| Approach: | They propose a method that utilizes process supervision to improve program repair by allowing users to input feedback from compilers and test cases. |
| Outcome: | The proposed method outperforms large outcome-based generation methods and is inspired by strategies used in programming competitions. |
Copied to clipboard
| Challenge: | Existing research explores to enhance the two sublayers separately to improve the capability of Transformer for text representation. |
| Approach: | They propose to combine SAN and Feed-Forward Networks to create a dynamic mask attention network with a learnable mask matrix which can model localness adaptively. |
| Outcome: | The proposed model outperforms the original Transformer on translation and text summarization tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) traditionally represent text as sequences of discrete tokens . a long-context scaling problem requires processing more tokens more efficiently . |
| Approach: | They propose a framework that renders long texts into compact visual pages and processes them with a vision-language model. |
| Outcome: | The proposed framework renders long texts into compact visual pages and processes them with a vision-language model. |
Copied to clipboard
| Challenge: | Existing Reinforcement Learning from Verifiable Rewards (RLVR) methods exhibit limited exploration due to reliance on on-policy rollouts which are limited to the current policy’s distribution, resulting in narrow trajectory diversity. |
| Approach: | They propose a framework that leverages the in-context learning capability of Large Reasoning Models to provide expert guidance using existing datasets. |
| Outcome: | The proposed framework improves RLVR performance and training stability on mathematical reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to enhance performance of large language models (LLMs) on Text-to-SQL tasks rely on execution-based or LLM-based reward models. |
| Approach: | They propose a reward model framework for RL-based Text-to-SQL that employs the GMNScore outcome reward model. |
| Outcome: | The proposed reward model outperforms existing reward models on standard benchmarks including Spider and BIRD. |
Copied to clipboard
| Challenge: | Visual Question Generation (VQG) research focuses on natural images while neglecting diagrams, a critical component of educational materials. |
| Approach: | They propose a diagram-driven course questions generation task to generate diagram-relevant questions for specific courses. |
| Outcome: | The proposed framework outperforms existing models on DiagramQG while maintaining strong generalizability across natural image datasets. |
Copied to clipboard
| Challenge: | Recent success of neural language models on the Winograd Schema Challenge has called for further investigation of commonsense reasoning ability of these models. |
| Approach: | They propose a logic-based framework that focuses on high-quality commonsense knowledge. |
| Outcome: | The proposed framework focuses on high-quality commonsense knowledge. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive ability to role-play humans and replicate complex social dynamics. |
| Approach: | They propose an efficient agent communication language induction for social simulations that reduces token consumption by over 20%. |
| Outcome: | The proposed model reduces token consumption by over 20% while preserving human language. |
Copied to clipboard
| Challenge: | Existing approaches to visual chain-of-thought are limited by external tools or fail to generate high-fidelity diagrams. |
| Approach: | They propose a framework to enable large multimodal models with VCoT capabilities . they pre-train a model on a 15.2M-pair corpus and teach it how to leverage visual aids . |
| Outcome: | The proposed framework unlocks complex, human-like visual reasoning in large language models . it pre-trains the model on a 15.2M-pair corpus and fine-tunes it on MathCanvas-Instruct . |
Copied to clipboard
| Challenge: | Existing relevance models rely on query-keyword pairs but keywords are usually short texts with scarce semantic information, which may not accurately reflect the underlying advertising purposes. |
| Approach: | They propose a bidding-graph augmented triple-based relevance model with three towers to deeply fuse the bidding graphs and semantic textual data. |
| Outcome: | The proposed model outperforms existing models on a large industry dataset and consistently outperformed existing models. |
Copied to clipboard
| Challenge: | Extensive evaluation of modern large language models shows performance gain over component LLMs. |
| Approach: | They propose a diversityoptimized LLM ensemble method with three unique properties . they introduce the focal diversity metric to capture diversityperformance correlation . |
| Outcome: | The proposed method outperforms the best-performing ensemble on four benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated notable capabilities across various tasks, showcasing complex problem-solving abilities. |
| Approach: | They propose a benchmark to evaluate the rule-based logical reasoning capabilities of Large Language Models (LLMs) they create simulated scenarios in which models execute or plan operations to achieve specific outcomes. |
| Outcome: | The proposed benchmark evaluates the performance of large language models on a variety of scenarios with varying difficulty levels. |
Copied to clipboard
| Challenge: | Existing methods to improve lifelong event detection performance are limited by the limited stored examples. |
| Approach: | They propose to use Episodic Memory Prompts to explicitly retain the learned task-specific knowledge. |
| Outcome: | The proposed method can be used to update a model with new event types while retaining the capability on previously learned types. |
Copied to clipboard
| Challenge: | Existing methods focus on preventing catastrophic forgetting by making compromises between the original and new language pairs, leading to sub-optimal performance on both translation tasks. |
| Approach: | They propose a dual importance-based model division method to divide the model parameters into two parts and separate the translation of the original and new tasks. |
| Outcome: | The proposed method outperforms strong baselines under different incremental translation scenarios. |
Copied to clipboard
| Challenge: | Mistake Notebook Learning (MNL) is a new memory framework for large language model agents . it allows agents to distill shared error patterns into structured "mistake notes" |
| Approach: | They propose a new memory framework that enables agents to self-curate generalizable guidance from batch-clustered failures. |
| Outcome: | The proposed framework achieves competitive performance compared to existing memory mechanisms. |
Copied to clipboard
| Challenge: | Difficulties with social aspects of language are among the hallmarks of autism spectrum disorder (ASD). |
| Approach: | They propose a transformer-based framework for identifying linguistic features associated with social aspects of communication using a corpus of conversations between adults with and without ASD and neurotypical conversational partners. |
| Outcome: | The proposed framework yields strong accuracy overall, but performance is significantly worse for the language of participants with ASD, suggesting they use a more diverse set of strategies for some social linguistic functions. |
Copied to clipboard
| Challenge: | a benchmark is designed to evaluate the repository-level dependency understanding of large language models (LLMs) based on 2683 repositories from real-world websites. |
| Approach: | They propose a benchmark to evaluate repository dependency understanding for large language models . DEPENDEVAL evaluates models on three core tasks across 8 programming languages . |
| Outcome: | The benchmark evaluates models on three core tasks across 8 programming languages from real-world repositories. |
Copied to clipboard
| Challenge: | Existing information extraction (IE) tasks rely on in-context learning with large language models. |
| Approach: | They propose a Bayesian-based in-context learning framework that refines label representations across IE tasks using particle filtering and Bayes updates. |
| Outcome: | The proposed framework improves performance over existing methods (up to 30%) it underperforms one-shot prompting by a substantial margin on NER tasks and CodeIE fails on RE tasks with near-zero micro-F1. |
Copied to clipboard
| Challenge: | Existing methods for vision-language pre-training lack high-level semantics and text is not sufficiently involved in masked modeling. |
| Approach: | They propose a semantics-enhanced cross-modal MIM framework for vision-language representation learning that harvests high-level semantics from global image features via self-supervised agreement learning and transfers them to local patch encodings by sharing the encode space. |
| Outcome: | The proposed model achieves state-of-the-art or competitive performance on multiple vision-language tasks. |
Copied to clipboard
| Challenge: | a new paradigm for dialogue systems is being developed to mimic human interactions . the current single-step dialogue paradigm lacks the depth and fluidity of human interactions. |
| Approach: | They propose a step-by-step dialogue paradigm that mimics human interactions . they use a dataset to fine-tune existing language models . |
| Outcome: | The proposed system mimics the dynamic nature of human conversations . it is compared with existing paradigms and will be released later this year . |
Copied to clipboard
| Challenge: | Empirical evaluations show that Mixture of Expert Prompt Tuning outperforms state-of-the-art parameter efficient baselines on SuperGLUE. |
| Approach: | They propose a pretrain-then-fine-tune paradigm for manifold mapping using multiple prompt experts. |
| Outcome: | Empirical results show that the proposed approach outperforms state-of-the-art methods on SuperGLUE while reducing activated prompts by 79.25%. |
Copied to clipboard
| Challenge: | Existing methods for relation extraction with distant supervision generate plenty of training samples but noisy labels and imbalanced training data cause problems. |
| Approach: | They propose a method that automatically labels a sentence with relational triples from a knowledge base. |
| Outcome: | The proposed method outperforms existing methods even with false positive samples. |
Copied to clipboard
| Challenge: | Existing benchmark-based evaluations cannot accurately reflect the performance of real-world applications. |
| Approach: | They propose a reliable strategy for domains to choose more robust LLMs for real-world applications. |
| Outcome: | The proposed strategy addresses the challenges faced by domains to choose more robust LLMs for real-world applications. |
Copied to clipboard
| Challenge: | Decoding by contrasting layers (DoLa) is designed to improve the generation quality of large language models (LLMs) however, this approach does not work well on non-English tasks. |
| Approach: | They propose a contrastive decoding algorithm that uses amateur logits to contrast with the output of an expert model's early exit logits. |
| Outcome: | The proposed method outperforms baselines and significantly improves chain-of-thought reasoning accuracy across 11 languages. |
Copied to clipboard
| Challenge: | Existing approaches to speech-to-singing voice conversion are difficult to learn in text-free situations. |
| Approach: | They propose an STS model which views speech variance as different modalities . it uses a novel rhythm adaptor to predict the target rhythm representation . they also use the predicted rhythm representation to re-align the content . |
| Outcome: | The proposed model achieves superior performance in terms of objective and subjective metrics. |
Copied to clipboard
| Challenge: | Experimental results show that our approach can significantly improve the parsing accuracy of all baseline models, leading to new state-of-the-art results. |
| Approach: | They propose a deep hierarchical syntax understanding approach to improve the cross-lingual semantic memory capability of large language models by implicitly aligning linguistic knowledge between source and target languages. |
| Outcome: | The proposed approach improves the cross-lingual semantic memory capability of large language models by combining implicit multi-task fine-tuning and explicit label bank guiding. |
Copied to clipboard
| Challenge: | ternary quantization is a powerful solution for resource-constrained edge devices . current implementations suffer from a fundamental misalignment with commodity hardware . |
| Approach: | They propose a hardware-efficient ternary quantization framework that packs weights into five bits to restore power-of-two alignment. |
| Outcome: | The proposed framework reduces weights to -1, 0, +1 while preserving power-of-two alignment. |
Copied to clipboard
| Challenge: | Existing resources for visual knowledge remain confined to English, creating a "Worldwide Knowledge Gap" eric liu: production and dissemination of knowledge exhibit a distinct trend toward decentralization and linguistic fragmentation. |
| Approach: | They propose a dataset for multilingual visual knowledge seeking and updating across ten major languages. |
| Outcome: | The proposed dataset is the first dynamic-updating dataset for multilingual visual knowledge seeking and updating across ten major languages. |
Copied to clipboard
| Challenge: | Large language models are used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown. |
| Approach: | They propose a benchmark for evaluating large language models using a well-organized taxonomy. |
| Outcome: | The proposed model is based on a well-organized taxonomy and compares it with other models. |
Copied to clipboard
| Challenge: | Text-to-audio (T2A) models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio. |
| Approach: | They propose to use AI feedback learning to enhance basic capabilities of text-to-audio models . they use a large audio preference dataset to evaluate the model's capabilities . |
| Outcome: | The proposed model improves in simple and complex scenarios with AI feedback learning. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing studies have investigated knowledge poisoning attacks in medical RAG systems . knowledge poison attacks can disrupt model outputs and undermine system reliability . |
| Approach: | They propose a knowledge poisoning framework that injects misinformation into textual data . they propose to use paired visual data as a query-agnostic trigger to promote retrieval . |
| Outcome: | The proposed framework produces clinically plausible but incorrect generations on five LLMs and datasets. |
Copied to clipboard
| Challenge: | Recent techniques such as Generation-Augmented Retrieval (GAR) and Generative Document Retrieleval (GDR) leverage LLMs to enhance retrieval performance but face key challenges: GAR’s generated content may not always align with the target document corpus, while GDR limits the generative capacity of LLM. |
| Approach: | They propose a Context-Aware Generation-Augmented Retrieval approach which integrates corpus information into their generation process. |
| Outcome: | Experimental results show that CA-GAR outperforms existing methods on seven tasks and four non-English languages. |
Copied to clipboard
| Challenge: | Sentence embedding is essential for many NLP tasks, but reliance on manual labels limits scalability. |
| Approach: | They propose a method for controlling the generation direction of large language models in the latent space by integrating ranking information and semantic information. |
| Outcome: | The proposed method achieves new SOTA performance with a modest cost in ranking sentence synthesis. |
Copied to clipboard
| Challenge: | Automatic post-editing (APE) is an important task in natural language processing. |
| Approach: | They propose a method that explicitly models how to copy words from a machine translation to a correct translation. |
| Outcome: | The proposed method outperforms all published methods on the WMT 2016-2017 datasets. |
Copied to clipboard
| Challenge: | Prompt tuning for pre-trained language models has shown remarkable performance . however, prompt tuning is still not fully explored . |
| Approach: | They propose to pre-train prompts by adding soft prompts into the pre-training stage to obtain a better initialization. |
| Outcome: | The proposed framework outperforms full-model tuning under full-data and few-shot learning settings. |
Copied to clipboard
| Challenge: | Existing methods to compress context information ignore holistic contextual dependencies. |
| Approach: | They propose a method that adjusts position encodings to minimize the distance between context tokens and special tokens. |
| Outcome: | Enhanced Position Layout (EPL) improves compression of context information in large language models. |
Copied to clipboard
| Challenge: | Multi-aspect controllable text generation has attracted increasing attention . but the mutual interference of multiple prefixes limits its extensibility to training-time unseen combinations. |
| Approach: | They propose to use trainable gates to normalize the intervention of prefixes to restrain the interference. |
| Outcome: | The proposed approach outperforms baselines on constraint accuracy, text quality, and extensibility. |
Copied to clipboard
| Challenge: | Existing methods for complex instruction-following with elaborate constraints rely on a weaker model, especially GPT-4, limiting their application. |
| Approach: | They propose a Multi-granularity Self-Contrastive Training framework to improve instruction alignment without relying on a stronger model. |
| Outcome: | The proposed framework improves instruction-following with elaborate constraints without external supervision on coarse and fine granularity. |
Copied to clipboard
| Challenge: | Consistency identification in task-oriented dialog usually consists of three subtasks . a proposed model for consistency identification in dialog is based on an explicit interaction paradigm . |
| Approach: | They propose a cycle guided interactive learning model that makes information exchange explicit from all the three tasks. |
| Outcome: | The proposed model achieves state-of-the-art performance pushing the overall score to 56.3% (5.0% point absolute improvement) |
Copied to clipboard
| Challenge: | Existing ODQA datasets consist mainly of Wikipedia corpus, and are insufficient to study models’ generalizability across diverse domains. |
| Approach: | They propose a benchmark to evaluate ODQA's domain robustness using Wikipedia corpus . they annotate QA pairs in retrieval datasets with rigorous quality control . |
| Outcome: | The proposed benchmark improves model performance on annotated QA pairs in retrieval datasets with rigorous quality control. |
Copied to clipboard
| Challenge: | Existing tool environments face challenges in balancing stability, scale, and realism, especially for benchmarking purposes. |
| Approach: | They propose a framework that trains specialized LLMs to accurately simulate real API responses by supervised fine-tuning and chain-of-thought reasoning. |
| Outcome: | The proposed framework achieves superior accuracy and stability compared to state-of-the-art methods on the newly constructed MirrorAPI-Bench and its integration into StableToolBench. |
Copied to clipboard
| Challenge: | Personalization in conversational AI requires persona profiles and contextual understanding to create meaningful conversations. |
| Approach: | They propose a method that softly prompts LLMs for personalized conversations in a selective way. |
| Outcome: | The proposed approach improves response diversity by up to 90% on the CONVAI2 dataset. |
Copied to clipboard
| Challenge: | Existing methods for long-video inference use compression or sparse attention . existing methods restrict LMMs from handling longer, more complex videos . |
| Approach: | They propose a sequence-parallel framework with optimized attention that accelerates long-video inference across multiple GPUs. |
| Outcome: | The proposed framework delivers speedups of 12.72x, 1.70x, and 1.18x over FlashAttn, ZigZagRing, and APB without significant performance loss. |
Copied to clipboard
| Challenge: | We introduce AI-Press, an automated news drafting and polishing system based on multi-agent collaboration and Retrieval-Augmented Generation. |
| Approach: | They introduce AI-Press, an automated news drafting and polishing system based on multi-agent collaboration and Retrieval-Augmented Generation. |
| Outcome: | The proposed system generates public responses considering demographic distributions. |
Copied to clipboard
| Challenge: | With the development of medical digitization, the extraction and structuring of electronic medical records (EMRs) have become challenging but fundamental tasks. |
| Approach: | They propose a speaker-aware dialogue encoder with multi-task learning which takes the speaker's identity into account and a co-attention fusion network to aggregate the utterance information. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods on the public medical dialogue extraction datasets to demonstrate its superiority. |
Copied to clipboard
| Challenge: | Large language models are often not well aligned with human intents, which requires additional training. |
| Approach: | They propose to use Black-Box Prompt Optimization (BPO) to perform alignments on large language models that are not well aligned with human intents. |
| Outcome: | The proposed model outperforms existing models and is model-agnostic. |
Copied to clipboard
| Challenge: | Existing literature on knowledge extraction for question answering questions whether it is still relevant for question answerrs. |
| Approach: | They extend an existing benchmark with knowledge extraction annotations and evaluate commercial and open-source LLMs of varying sizes. |
| Outcome: | The proposed model can achieve high QA accuracy, but can still benefit from knowledge extraction through augmentation with extracted triples and multi-task learning. |
Copied to clipboard
| Challenge: | a new dataset is being developed to improve the capabilities of mobile GUI-control agents. |
| Approach: | They propose a dataset designed for generalist mobile GUI-control agents . they use screenshots from popular mobile applications to create a detailed GUI-annotated dataset . |
| Outcome: | The Android Multi-annotation EXpo (AMEX) is a large-scale dataset for generalist mobile GUI-control agents . it includes screenshots from popular mobile applications, which are annotated at multiple levels . |
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has emerged as a paradigm for enhancing the reasoning capabilities of large language models. |
| Approach: | They propose a positive-advantage reweighting approach that regulates model entropy by adjusting the loss weights assigned to tokens with positive advantages during RLVR training. |
| Outcome: | The proposed approach regulates model entropy by adjusting loss weights assigned to tokens with positive advantages during RLVR training while maintaining competitive performance. |
Copied to clipboard
| Challenge: | Long-form audio understanding poses significant challenges due to the extreme length of audio sequences and the need to reason over heterogeneous acoustic cues distributed over time. |
| Approach: | They propose a retrieval-augmented generation framework for scalable long-form audio understanding . planRAG-Audio explicitly plans which modalities and temporal spans are required for a given query . |
| Outcome: | Experiments show that planRAG-Audio reduces the length of inputs for long-form audio models . the proposed framework can efficiently reason over long-term speech data . |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios . |
| Approach: | They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios. |
| Outcome: | The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm. |
Copied to clipboard
| Challenge: | a typical way to polish sentences is to add engaging modifiers, which enhance the meaning of the sentence. |
| Approach: | They propose a task that requires polishing sentences while maintaining fluency . they remove engaging modifiers from public resources and fine-tune LongLM to reconstruct original sentences from corrupted ones. |
| Outcome: | The proposed model generates more engaging sentences with suitable modifiers than strong baselines while keeping fluency. |
Copied to clipboard
| Challenge: | Recent studies have shown that adversarial examples can be easily fooled by adversarially perturbed examples. |
| Approach: | They propose a pluggable defense module PlugAT to provide robust predictions by adding a few trainable parameters to the model inputs while keeping the original model frozen. |
| Outcome: | The proposed model improves robustness over several strong baselines whilst training only 9.1% parameters. |
Copied to clipboard
| Challenge: | Existing RAG systems often underutilize the retrieved documents, authors say . they fail to extract and integrate key clues needed to support faithful and interpretable reasoning . |
| Approach: | a new framework extracts key clues from retrieved content and generates multiple reasoning paths . the framework optimizes the model by selecting the most appropriate reasoning path . |
| Outcome: | Experiments show that ClueAnchor outperforms baseline RAG frameworks in completeness and robustness. |
Copied to clipboard
| Challenge: | Recent advances in large language models showcase varied multilingual capabilities across tasks . previous assessments focused on fundamental natural language processing (NLP) or isolated capability-specific tasks. |
| Approach: | They propose a multilingual multitask benchmark to assess multilingual capabilities . they use a large-scale benchmark covering fundamental and capability-specialized datasets . |
| Outcome: | The proposed benchmark compares models and tasks across languages and tasks and examines knowledge transfer from English to other languages. |
Copied to clipboard
| Challenge: | Existing benchmarks that focus on knowledge-intensive tasks do not reflect diverse educational scenarios. |
| Approach: | They propose a benchmark that incorporates 9 major scenarios and 4,000 educational contexts. |
| Outcome: | The proposed model performs comparable to state-of-the-art large models on the test set. |
Copied to clipboard
| Challenge: | Existing studies have shown that adversarial samples are more vulnerable than normal ones to textual adversarials. |
| Approach: | They propose a simple and effective sharpness-based detector that can distinguish adversarial samples by maximizing the loss increment within the region where the inference sample is located. |
| Outcome: | The proposed method outperforms previous detection methods by large margins on three text classification tasks. |
Copied to clipboard
| Challenge: | Speculative sampling is an efficient way to accelerate the auto-regressive generation process of large language models. |
| Approach: | They propose a frequency-ranked speculative sampling framework that optimizes draft candidate selection through vocabulary space compression. |
| Outcome: | Experiments show that FR-Spec reduces LM Head computation overhead by 75% while ensuring the equivalence of the final output distribution. |
Copied to clipboard
| Challenge: | Text-to-Image Synthesis (TIS) is a popular task to convert natural language texts into realistic images. |
| Approach: | They propose a transformer-based Chinese text-to-image synthesizer for high-resolution image generation that incorporates linguistic and relational knowledge facts into the model to ensure better performance without the usage of ultra-large models. |
| Outcome: | The proposed model outperforms existing models in Chinese with linguistic and relational knowledge facts. |
Copied to clipboard
| Challenge: | Existing approaches to enhance the context-faithfulness of Large Language Models (LLMs) ignore the fundamental mechanism of how contextual information is processed within LLMs’ internal states. |
| Approach: | They propose a method that enhances the utilization of contextual knowledge within LLMs’ internal representations by employing V-usable information analysis. |
| Outcome: | The proposed method improves context-faithfulness generation in Question-Answering tasks, particularly in scenarios involving unknown or conflicting contextual knowledge. |
Copied to clipboard
| Challenge: | ECom-Bench is a benchmark framework for evaluating LLM agent with multimodal capabilities in e-commerce customer support domain. |
| Approach: | They introduce a benchmark framework for evaluating LLM agent with multimodal capabilities in the e-commerce customer support domain. |
| Outcome: | The proposed benchmark features dynamic user simulation based on persona information from real e-commerce customer interactions and a realistic task dataset derived from authentic ecommerce dialogues. |
Copied to clipboard
| Challenge: | Mixture-of-Experts (MoE) LLMs achieve higher performance with fewer active parameters, but are still difficult to deploy due to their immense parameter sizes. |
| Approach: | They propose expert-level sparsification techniques to enhance the deployment efficiency of large language models by introducing plug-and-play expert pruning and skipping techniques. |
| Outcome: | The proposed methods reduce model sizes and increase inference speed while maintaining satisfactory performance across a wide range of tasks. |
Copied to clipboard
| Challenge: | Neural machine translation models are sensitive to noises in input sentences . one special kind of noise is the homophone noise, where words are replaced by other words with similar pronunciations. |
| Approach: | They propose to embed phonetic and textual information into neural machine translation datasets to improve robustness to homophone noises. |
| Outcome: | The proposed method improves the robustness of neural machine translation to homophone noises on clean test sets. |
Copied to clipboard
| Challenge: | Recent studies have demonstrated that inference-time scaling increases performance of Large Language Models (LLMs) in various reasoning tasks such as mathematics and complex question answering by increasing the length of Chain-of-Thought (CoT). |
| Approach: | They propose a model which synthesizes longer CoT data and iteratively improves performance through self-training by incorporating a few demonstration examples. |
| Outcome: | The proposed model achieves an average improvement of more than +2.5 points across five reasoning tasks: MMLU, GSM8K, ARC-C, HellaSwag, and BBH on two backbone models. |
Copied to clipboard
| Challenge: | Existing methods to accelerate inference speed are model compression and dynamic computation (e.g., dynamic token pruning). |
| Approach: | They propose a two-stage knowledge distillation framework that produces a customized small language model for dynamic token pruning. |
| Outcome: | The proposed framework can make the small language model more customized for dynamic token pruning and achieve better speed-performance trade-off. |
Copied to clipboard
| Challenge: | Existing methods to measure instance difficulty use generalization and threshold-tuning . a new approach to learn to exit is based on hash functions to assign tokens to a fixed exiting layer. |
| Approach: | They propose a Hash-based Early Exiting approach that replaces learn-to-exit modules with hash functions to assign each token to a fixed exiting layer. |
| Outcome: | The proposed approach improves on learning to exit and predicting instance difficulty. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting relies on the initial decisions, causing errors in early steps to accumulate and impact the final answers. |
| Approach: | They propose a divide-and-conquer style algorithm that leverages large language models to raise and answer sub-questions until collecting enough information to tackle the original one. |
| Outcome: | The proposed algorithm is more robust to errors and errors than CoT prompting and Tree-of-Thought prompting methods. |
Copied to clipboard
| Challenge: | Existing methods to optimize LLM for long sequences for long documents are slow and consume memory. |
| Approach: | They propose a method that starts with a small memory size and gradually increases it . they propose Decremental Chunk based on Incremental Memory (IMDC) which reduces chunk size while increasing memory size . |
| Outcome: | The proposed method is faster (1.45x) and reduces GPU memory consumption by 23.3% compared to fixed-size memory. |
Copied to clipboard
| Challenge: | Existing watermarking methods are fact-agnostic and cause "faithfulness hallucinations" a novel framework to enforce knowledge loyalty is proposed to improve watermarks . |
| Approach: | They propose a new framework that enforces knowledge loyalty by spoofing terms from retrieved contexts and prompt-based semantic guidance to protect against factual corruption. |
| Outcome: | The proposed framework reduces the Knowledge Corruption Rate while maintaining its original high security and robustness. |
Copied to clipboard
| Challenge: | Existing methods to identify phenotypes using electronic health records (EHRs) are expensive and difficult to transfer models from one disease to another. |
| Approach: | They propose a task-oriented dialogue system framework to make diagnosis for patients automatically, which can converse with patients to collect additional symptoms beyond their self-reports. |
| Outcome: | The proposed system can collect additional symptoms from conversation and improve disease identification accuracy. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods focus on editing triple-based facts such as entity-relation pairs and events (multiple triplets). |
| Approach: | They propose a Knowledge Localization for Free-Text method which uses a Dynamics-aware Module to locate the parameter positions corresponding to commonsense knowledge and a knowledge editing module to update knowledge. |
| Outcome: | The proposed method exploits the potential of the MLP and Attention layers and edits commonsense knowledge based on free-text. |
Copied to clipboard
| Challenge: | Existing research on visual question generation is focused on training models to fit the annotated data set that makes them indifferent from other language generation tasks. |
| Approach: | They propose to use two discriminators to enhance the training of a visual question generator to ask natural questions about an image. |
| Outcome: | The proposed model outperforms state-of-the-art models in terms of automatic and human evaluation metrics. |
Copied to clipboard
| Challenge: | Existing context window extension methods obstruct scaling external knowledge input. |
| Approach: | They develop a multi-agent framework to overcome two core bottlenecks in existing agent orchestration designs. |
| Outcome: | The proposed framework overcomes two core bottlenecks and improves inference-time knowledge integration without longer-context training. |
Copied to clipboard
| Challenge: | Recent studies have shown that pretrained models are effective for low-resource downstream tasks. |
| Approach: | They propose a two-stage finetuning method to transfer pretrained models to MSG tasks by concatenating multiple sources into a single long sequence. |
| Outcome: | The proposed model outperforms baselines on the WMT17 APE task and multi-source translation task using the WTM14 test set. |
Copied to clipboard
| Challenge: | Existing methods that optimize for scalar scores or ranking reward ignore multi-dimensional nature of human preferences. |
| Approach: | They propose to extend the preference of Direct Preference Optimization to two dimensions: segments and aspects. |
| Outcome: | The proposed framework decomposes the overall objective into multi-segment and multi-aspect objectives. |
Copied to clipboard
| Challenge: | Existing approaches to formalizing mathematical statements face limitations in accuracy, especially in the context of complex, highlevel problems that involve sophisticated mathematical reasoning. |
| Approach: | They propose a CriticLean framework that elevates the role of the critic from a passive validator to an active learning component and introduce a benchmark to measure models’ ability to distinguish semantically correct from incorrect formalizations. |
| Outcome: | The proposed framework outperforms open- and closed-source benchmarks and shows that it significantly outperformed existing models. |
Copied to clipboard
| Challenge: | Experimental results show that our model achieves competitive results with the state-of-the-art classification-based model OneIE on ACE 2005. |
| Approach: | They propose a generative template-based event extraction method with dynamic prefix . they integrate context information with type-specific prefixes to learn a context-specific name for each context . |
| Outcome: | The proposed method achieves competitive results with state-of-the-art model OneIE on ACE 2005 and performs well on ERE. |
Copied to clipboard
| Challenge: | Experiments on five text classification benchmarks and five backbone models have shown that our methods reduce the error rate over Mixup variants in a significant margin (up to 31.3%), especially in low-resource conditions (upto 17.5%). |
| Approach: | They propose to add a small adversarial perturbation to the mixing coefficients rather than the examples to relax locally linear constraints. |
| Outcome: | Experiments on five text classification benchmarks and five backbone models show that the proposed methods reduce the error rate over Mixup variants by 31.3%, especially in low-resource conditions. |
Copied to clipboard
| Challenge: | Despite LLMs' impressive capabilities in musical knowledge, music reasoning remains an unsolved task. |
| Approach: | They propose an open-source large language model (LLM) that integrates intrinsic musical abilities into LLaMA2 and GPT-3.5. |
| Outcome: | The proposed model can understand and generate music with a pure text tokenizer without external multi-modal neural structures or tokenizers. |
Copied to clipboard
| Challenge: | kNN-BOX enables quick development and visualization for novel generation paradigm . Currently, knn-BOx has provided implementation of seven popular kN-MT variants . |
| Approach: | They propose a framework which decomposes the datastore-augmentation approach into three modules . they apply kNN-BOX to machine translation and three other tasks . |
| Outcome: | The proposed framework decomposes the datastore-augmentation approach into three modules . it provides implementation of seven popular kNN-MT variants, covering research from performance enhancement to efficiency optimization. |
Copied to clipboard
| Challenge: | RFC-Bench evaluates large language models on financial misinformation under realistic news . current models struggle to maintain coherent belief states without external grounding, study finds . |
| Approach: | They propose a benchmark for evaluating large language models on financial misinformation under realistic news. |
| Outcome: | The proposed model performs better when context is available, while reference-free settings expose significant weaknesses. |
Copied to clipboard
| Challenge: | Large language models have demonstrated impressive reasoning capabilities across multiple languages, but the relationship between capabilities in different languages is less explored. |
| Approach: | They decompose the process of reasoning tasks into two separate components: knowledge retrieval and knowledge-free reasoning. |
| Outcome: | The proposed model can be transferred across source-target languages despite secondary impact of resource in some specific target languages, while cross-lingual knowledge retrieval significantly hinders the transfer. |
Copied to clipboard
| Challenge: | Existing retrieval-augmented generation methods rely on multiple calls of large language models (LLMs) Large-language models lack knowledge underrepresented in training data and still face hallucinations. |
| Approach: | They propose an efficient retriever for multi-hop question answering that generates new queries iteratively without the need for LLM calls. |
| Outcome: | The proposed method surpasses existing methods on three open-domain multi-hop question-answering datasets. |
Copied to clipboard
| Challenge: | Existing word embeddings only consider contextual information, which is suboptimal when used in various tasks due to a lack of task-specific features. |
| Approach: | They propose a task-oriented word embedding method that regularizes the distribution of words to enable a clear classification boundary. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on a text classification task. |
Copied to clipboard
| Challenge: | Existing evaluation methods for mobile GUI agents rely on static frame assessments or offline static apps. |
| Approach: | They propose an evaluation system that leverages large language models as reward models to verify task completion and process achievement. |
| Outcome: | The proposed system addresses the limitations of traditional function based evaluation methods on online dynamic apps. |
Copied to clipboard
| Challenge: | Existing approaches to integrating graph and language models face two key limitations: achieving robust semantic alignment and ensuring interpretability in outputs. |
| Approach: | They propose a framework to integrate graph and language modalities while enhancing transparency. |
| Outcome: | Extensive experiments on three benchmark datasets show that the proposed framework surpasses existing methods in efficiency and generates outputs that are significantly more interpretable. |
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models (MLLMs) have garnered significant attention, offering a promising pathway toward artificial general intelligence (AGI). |
| Approach: | They propose a benchmark to evaluate associative ability while circumventing the inherent ambiguity in association tasks by decomposing ambiguities into two types and propose 'assoCiAm' they conduct extensive experiments on MLLMs, revealing a strong positive correlation between cognition and association. |
| Outcome: | The proposed method shows that ambiguity in association evaluations makes MLLMs more random-like and the model's behavior more random. |
Copied to clipboard
| Challenge: | Existing models for solving math word problems rely on predefined rules or feature engineering. |
| Approach: | They propose to incorporate copy and alignment mechanism into the sequence-to-sequence model to address two shortcomings . they use model output as a feature and incorporate it into the feature-based model to explore the effectiveness . |
| Outcome: | The proposed model outperforms the state-of-the-art models on the problem solving task. |
Copied to clipboard
| Challenge: | Large Language Model (LLM) watermarking is radioactive and enables the detection of watermarks inherited by student models when trained on the outputs of watermarked teacher models. |
| Approach: | They propose two types of watermark removal attacks that allow student models to perform untraceable knowledge distillation while avoiding watermark inheritance. |
| Outcome: | The proposed attacks eliminate inherited watermarks while maintaining knowledge transfer efficiency and low computational overhead. |
Copied to clipboard
| Challenge: | Existing back-translation methods focus on in-domain lexical knowledge, which may lead to poor translation of unseen in- domain words. |
| Approach: | They propose an iterative constrained back-translation method to incorporate in-domain lexical knowledge into synthetic parallel data from BT. |
| Outcome: | The proposed method improves the BLEU score by up to 3.08 on four domains. |
Copied to clipboard
| Challenge: | Existing methods for modifying large language models focus on individual models, resulting in errors and hallucinations. |
| Approach: | They propose an ensemble-based approach that employs a plug-in model as the editing module and a dynamic weight mechanism to enhance its effectiveness. |
| Outcome: | The proposed approach outperforms existing methods while achieving superior editing efficiency. |
Copied to clipboard
| Challenge: | Existing defense approaches focus on developing new model structures or training algorithms, but they do little to tap the potential of training instances. |
| Approach: | They propose a method that can distinguish between robust and non-robust instances according to the model’s sensitivity to perturbations on individual instances during training. |
| Outcome: | The proposed method can distinguish between robust and non-robust instances according to the model’s sensitivity to perturbations on individual instances during training. |
Copied to clipboard
| Challenge: | TrickCatcher generates test cases that pass existing tests yet contain bugs . a recent study found that tricky bugs are not detected by test suites . |
| Approach: | They propose an LLM-powered approach to generating test cases for uncovering bugs in plausible programs . they use a PUT and specification to generate program variants, an input generator and an Llm to construct test inputs . |
| Outcome: | The proposed approach achieves recall, precision, and F1 scores that are 1.80, 2.65, and 1.66 . trickCatcher generates program variants based on the program under test and its specification . |
Copied to clipboard
| Challenge: | Existing methods for strategic reasoning face challenges in adaptability, scalability, and transferring strategies to new contexts. |
| Approach: | They propose an explicit policy optimization model that provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior. |
| Outcome: | The proposed model provides strategies in open-ended action space and can be plugged into arbitrary LLM agents to motivate goal-directed behavior. |
Copied to clipboard
| Challenge: | Recent methods to reduce the KV cache size fail to identify crucial KVs for generation while excluding others accurately, resulting in severe information loss. |
| Approach: | They propose an intention-aware KV cache eviction method that identifies and retains crucial KVs according to the attention distribution of intention, which semantically reflects the user’s goal and determines which part of the context is relevant. |
| Outcome: | The proposed method can maintain the model performance while reducing the KV cache size from 128K to 2K, leading to a 6.3x increase in decoding speed and 7.8x enhancement in memory efficiency compared to the default setting. |
Copied to clipboard
| Challenge: | FinReporting is an agentic workflow for localized cross-jurisdiction financial reporting . existing approaches assume a single-market setting and overlook structural differences across jurisdictions . |
| Approach: | They propose a workflow that decomposes financial reporting into auditable stages . they use Large Language Models to extract and summarize corporate disclosures . |
| Outcome: | The proposed system decomposes reporting into auditable stages . it improves consistency and reliability under heterogeneous reporting regimes. |
Copied to clipboard
| Challenge: | Large language models struggle when dealing with complex, ill-formed, or noisy inputs . open-source models are less robust, while closed-source ones are more robust . |
| Approach: | They propose to use GSM-Noise to refine inputs before engaging in in-depth analysis to improve LLM robustness under noisy conditions. |
| Outcome: | The proposed model can achieve consistent performance gains under noisy conditions with prompt engineering, supervised finetuning, and reinforcement learning. |
Copied to clipboard
| Challenge: | Existing training data detectors fail to detect clean samples from contaminated test sets . existing methods fail to identify clean samples due to black-box nature of LLMs . |
| Approach: | They propose a framework that detects and filters contaminated evaluation data . they propose 'failure detection' to reduce the proportion of contaminated samples mistakenly retained . |
| Outcome: | The proposed framework reduces false discovery rate (FDR) under valid FDR control while maintaining evaluation consistency. |
Copied to clipboard
| Challenge: | Existing knowledge graphs lack the ability to integrate structural information into LLMs and output predictions deterministically. |
| Approach: | They propose a method which encodes structural information of KGs and merges it with LLMs to enhance KGC performance. |
| Outcome: | The proposed method improves the performance of KG Completion datasets on KGs by integrating structural information with LLMs. |
Copied to clipboard
| Challenge: | Visual Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images. |
| Approach: | They propose to use 1,200 images paired with detailed geographic metadata to evaluate VLMs' performance. |
| Outcome: | The models achieve 53.8% accuracy in city prediction, but exhibit significant biases in regional tasks. |
Copied to clipboard
| Challenge: | Contemporary practices in instruction tuning often hinge on enlarging data scaling without a clear strategy for ensuring data quality. |
| Approach: | They propose a method that leverages one-shot learning to discern and select high-quality instruction data from extensive datasets. |
| Outcome: | Nuggets outperforms existing methods on MT-Bench and Alpaca-Eval benchmarks. |
Copied to clipboard
| Challenge: | Existing personalized dialogue models focus on dialogue history and personality information, reducing the responses’ consistency. |
| Approach: | They propose a Persona-Aware LLM-enAnCEd(PALACE) framework that generates responses consistent with dialogue history and personality information across multiple sessions to engage users’ interest in the dialogue. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods in automatic and human evaluation metrics on the MSC and DuLeMon datasets. |
Copied to clipboard
| Challenge: | Existing models that describe concepts in everyday situations are difficult to summarize in a single sentence. |
| Approach: | They propose DimonGen, which generates sentences describing concept relationships in everyday scenarios. |
| Outcome: | The proposed model outperforms baseline models in terms of quality and diversity of generated sentences. |
Copied to clipboard
| Challenge: | Existing methods for hierarchical text classification are lacking in the field of natural language processing. |
| Approach: | They propose a hierarchy-aware T5 model with path-adaptive attention mechanism to exploit hierarchical dependency across different levels. |
| Outcome: | The proposed model outperforms state-of-the-art models especially in Macro-F1 and low Macro. |
Copied to clipboard
| Challenge: | Conventional approaches to learning sentence embeddings from dialogues employ the siamese-network for this task, but such architecture yields a large gap between training and evaluating. |
| Approach: | They propose a dialogue-based contrastive learning approach to learn sentence embeddings from dialogues using a siamese-network. |
| Outcome: | The proposed model outperforms baseline methods on three multi-turn dialogue datasets in terms of MAP and Spearman’s correlation measures. |
Copied to clipboard
| Challenge: | Unsupervised neural machine translation methods have been observed to make particular errors in comparison to supervised machine translation, such as confusing nouns that pertain to the same semantic category. |
| Approach: | They propose a method that incorporates images at the word level to augment lexical mappings. |
| Outcome: | Experiments on a multi-lingual dataset show that the proposed method generates more accurate translations with only monolingual data. |
Copied to clipboard
| Challenge: | Existing generative language models (LMs) can generate new or reusable theorems, but their ability to generate new theorels is under-explored. |
| Approach: | They propose to use Metamath library to generate new theorems that can be saved as reusable knowledge for future theoretical proving. |
| Outcome: | The proposed benchmark evaluates whether an agent can generate valuable (and possibly brand new) theorems that are applicable for downstream theoretic proving as reusable knowledge. |
Copied to clipboard
| Challenge: | Existing work on instruction tuning has focused on task level, without considering that tasks are artificially defined and, to LLMs, merely consist of tokens and representations. |
| Approach: | They propose a training data arrangement framework that allows for continual learning and loss reduction. |
| Outcome: | The proposed framework promotes continual learning and loss reduction on unseen tasks. |
Copied to clipboard
| Challenge: | Existing models are susceptible to reward hacking, leading to a substantial overestimation of a model's reasoning ability. |
| Approach: | They propose a Rubric Reward Model that rewards the entire reasoning trajectory against problem-specific rubrics. |
| Outcome: | The proposed model outperforms outcome-only supervision on four math benchmarks and boosts Verified Pass@1024 from 26.7% to 62.6% and reduces the incidence of Miracle Steps by 71%. |
Copied to clipboard
| Challenge: | Generative Large Language Models (LLMs) produce untruthful outputs, referred to as hallucinations, which are often referred as false positives. |
| Approach: | They propose an open-source Python library with over 30 truthfulness prediction methods. |
| Outcome: | The proposed methods span diverse trade-offs in computational cost, access level, grounding document requirements, and supervision type (self-supervised or supervised). |
Copied to clipboard
| Challenge: | Pre-Trained Models (PTMs) have reshaped the development of natural language processing (NLP) but it is not easy to obtain high-performing PTMs without a large amount of labeled training data and deploy them online with fast inference speed. |
| Approach: | They propose to make it easy to build NLP applications with knowledge-enhanced pre-training and knowledge distillation. |
| Outcome: | EasyNLP supports a comprehensive suite of NLP algorithms and features knowledge-enhanced pre-training, knowledge distillation and few-shot learning functionalities. |
Copied to clipboard
| Challenge: | Recent studies show that LLMs’ intrinsic self-correction fails without oracle labels as feedback. |
| Approach: | They propose to use one simple task and three complex tasks with state-of-the-art LLMs like ChatGPT, Llama, and DeepSeek to interpret LLM's intrinsic self-correction. |
| Outcome: | The proposed methods reveal the dark side of LLMs’ intrinsic self-correction for different tasks, especially for those failure cases. |
Copied to clipboard
| Challenge: | Existing methods for predicting protein-protein interactions oversimplify the problem of PPI prediction in a semi-supervised manner. |
| Approach: | They propose a multimodal large language model that integrates proteins and PPI networks. |
| Outcome: | Experiments show that LLaPA can predict protein-protein interactions (mPPI) types and affinities based on sequence data. |
Copied to clipboard
| Challenge: | Existing methods to improve robustness of models focus on a single dataset . but, there are few studies on how to combine merits of different datasets . |
| Approach: | They propose a federated learning framework that could unify datasets and tasks . they propose MV-Encoder as backbone of the framework to provide multi-granularity text representations . |
| Outcome: | The proposed framework improves on two SLU benchmark datasets and federated learning settings. |
Copied to clipboard
| Challenge: | Existing methods for harmful meme detection are limited due to the dynamic nature of memes . eliciting knowledge-revising behavior within the LMM agent is a key factor in achieving this goal . |
| Approach: | They propose an agency-driven framework for low-resource harmful meme detection . they use annotated memes to leverage label information as auxiliary signals for model . |
| Outcome: | The proposed framework achieves superior performance than state-of-the-art methods on the low-resource harmful meme detection task. |
Copied to clipboard
| Challenge: | Existing methods for labeling relational facts require significant expert labor to write relation-specific patterns, which makes them too sophisticated to generalize quickly. |
| Approach: | They propose a neural pattern diagnosis framework that can summarize and refine relation-specific patterns with human experts in the loop. |
| Outcome: | The proposed framework can summarize and refine high-quality relational patterns from noise data with human experts in the loop. |
Copied to clipboard
| Challenge: | Existing methods for generating complex instructions are resource-intensive and lack diversity. |
| Approach: | They propose a framework to generate complex instructions with constraints using a document-generated initial instruction and an iterative refinement framework to incorporate LLM-as-judge guidance. |
| Outcome: | The proposed framework significantly outperforms existing methods for generating complex instructions, and outperformed existing methods. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have advanced machine translation (MT) a meta-evaluation dataset focused on non-literal translations is lacking . experimental results show the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge. |
| Approach: | They propose a meta-evaluation framework that leverages sub-agents to evaluate machine translation metrics. |
| Outcome: | The proposed framework improves on the knowledge cutoff and score inconsistency problem. |
Copied to clipboard
| Challenge: | Current evaluation methods focus on one dataset, e.g., Newstest dataset in each year’s WMT Metrics Shared Task. |
| Approach: | They propose to use a single dataset to evaluate the performance of automatic translation metrics. |
| Outcome: | The results show that the rankings of metrics vary when the evaluation is conducted on different datasets. |
Copied to clipboard
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Copied to clipboard
| Challenge: | Existing approaches to RLVR train LMs based on their own on-policy responses and are constrained by the initial capability of LM. |
| Approach: | They propose an approach that hints LMs with their self-made mistakes without external guidance. |
| Outcome: | The proposed approach outperforms the normal group relative policy optimization and requires no external guidance. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models can translate multiple language pairs in a single model but lacks ability to capture language-specific features. |
| Approach: | They propose a token-level feature mixing method that captures different features and dynamically determines feature sharing across languages. |
| Outcome: | The proposed method outperforms baselines and can be extended to zero-shot translation. |
Copied to clipboard
| Challenge: | Existing methods to model emotion-relevant content are based on rule-based and statistics-based approaches. |
| Approach: | They propose a semi-supervised graph-based algorithm to produce rich structural descriptors . they use word embeddings to evaluate the algorithm on emotion recognition tasks . |
| Outcome: | The proposed method outperforms state-of-the-art methods on emotion recognition tasks. |
Copied to clipboard
| Challenge: | Extensive experiments show that typed decoders outperform state-of-the-art baselines and can generate more meaningful questions. |
| Approach: | They devised two typed decoders that generate questions with different types of interrogatives, topic words, and ordinary words. |
| Outcome: | Extensive experiments show that the typed decoders outperform state-of-the-art baselines and can generate more meaningful questions. |
Copied to clipboard
| Challenge: | Existing models for text understanding fail to adapt to domain shifts in real-world applications . current models do not improve themselves as they are applied to new domains . |
| Approach: | They propose a continual test-time adaptation framework that adapts to evolving domains . they propose accumulating domains and a refine-then-filter framework to calibrate teacher predictions . |
| Outcome: | The proposed model excels in a teacher-student framework adaptable to evolving domains. |
Copied to clipboard
| Challenge: | Existing zero-shot methods fail to align speech and text into a shared semantic space . Existing methods require expensive and expensive parallel ST data . |
| Approach: | They propose a method that uses a shared discrete vocabulary space to align speech and text into a common space. |
| Outcome: | The proposed method significantly improves the SOTA and even performs on par with the strong supervised ST baselines. |
Copied to clipboard
| Challenge: | Recent prompt tuning (PT) has gained increasing attention as a parameter-efficient way of tuning pre-trained language models (PLMs). |
| Approach: | They propose a prompt tuning algorithm that uses a small-scale partial PLM and progressively expands its depth and width until the full-model size. |
| Outcome: | The proposed method could save over 30% of training computations while achieving comparable performance. |
Copied to clipboard
| Challenge: | Existing benchmarks for long-form generation assess real-world queries with hard-to-verify metrics or use synthetic setups that overlook real-life intricacies. |
| Approach: | They propose a new approach that balances verifiable and real-world assessment with Target-Anchored Evaluation. |
| Outcome: | The proposed model balances real-world and verifiable assessment with Target-Anchored Evaluation (TAE) it generates queries, textual materials, and anchors based on verifier targets within real-life scenarios . |
Copied to clipboard
| Challenge: | Generated infographics may appear correct at first glance but contain easily overlooked issues, such as distorted data encoding or incorrect textual content. |
| Approach: | They propose to evaluate reliability of text-to-infographic generation using IGenBench . they employ multimodal large language models to verify each question . |
| Outcome: | The proposed framework decomposes reliability verification into atomic yes/no questions based on a taxonomy of 10 question types. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Large Multimodal Models have exceeded general human capabilities in various tasks. |
| Approach: | They present an Olympiad-level bilingual multimodal scientific benchmark featuring 8,476 problems from Olympiad level mathematics and physics competitions. |
| Outcome: | The best performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning. |
Copied to clipboard
| Challenge: | Existing approaches focus on global planning, which plans toward the target before the conversation. |
| Approach: | They propose to generate a global path as a natural language sentence instead of a sequence of nodes. |
| Outcome: | The proposed method has fewer turns, more coherent semantics, and higher success rate than baselines. |
Copied to clipboard
| Challenge: | Recent advances struggle to train a separate model for each language pair, which is costly and unaffordable when the number of languages increases in the real world. |
| Approach: | They propose to train different MMT models to support translations between different languages. |
| Outcome: | The proposed model is able to handle the above issues by providing a shared semantic space for multiple languages. |
Copied to clipboard
| Challenge: | Existing approaches to address address standardization are lacking in the current field. |
| Approach: | They propose a framework that incorporates spatial knowledge into address texts and achieves efficient address standardization. |
| Outcome: | The proposed framework incorporates spatial knowledge into address texts and achieves efficient address standardization. |
Copied to clipboard
| Challenge: | Existing evaluations of LLMs in finance are text-only, monolingual, and largely saturated by current models. |
| Approach: | They propose a multilingual and multimodal benchmark for evaluating LLMs in real financial contexts. |
| Outcome: | The first expert-annotated multilingual and multimodal benchmark is released . it evaluates 21 leading LLMs and shows they perform better in multilingual settings . |
Copied to clipboard
| Challenge: | Recent advances in news summarization have created problems with “hallucinations” that are factually inconsistent with the source text. |
| Approach: | They propose to disentangle LLMs’ propensities to generate faithful and fake content by adopting a probing-based specific training method to improve their capacity of distinguishing two types of propensity. |
| Outcome: | The proposed method disentangles LLMs’ propensities to generate faithful and fake content and improves their ability to distinguish between two types of propensity. |
Copied to clipboard
| Challenge: | Recent ECG Self-Supervised Learning methods mitigate this by learning features without extensive labels but fail to capture fine-grained clinical semantics and require extensive task-specific fine-tuning. |
| Approach: | They propose a supervised pre-training framework for Multimodal ECG representation learning that combines structured diagnostic labels with large language models to help denoise, standardize cardiac concepts and improve clinical representation learning. |
| Outcome: | The proposed framework improves on six downstream datasets covering 106 cardiac conditions and achieves a zero-shot AUC performance of 77.20% over state-of-the-art eSSLs. |
Copied to clipboard
| Challenge: | Existing methods, tasks and benchmarks to measure model’s effective memory length are limited. |
| Approach: | They propose a method called forgetting curve to measure the memorization capability of long-context models. |
| Outcome: | The proposed method is robust to the tested corpus and experimental settings, and can be applied to any model size. |
Copied to clipboard
| Challenge: | Existing methods for multitask learning typically use a dataset name as input prefix, which limits the effectiveness of multitask training. |
| Approach: | They propose compositional task configurations, a set of prompts prepended to the encoder to improve cross-task generalization of unified models. |
| Outcome: | The proposed model outperforms the UnifiedSKG baseline by noticeable margins in both in-domain and zero-shot settings. |
Copied to clipboard
| Challenge: | Recent years have seen success in the use of deep neural networks on text summarization, but there is no clear understanding of why they perform so well or how they might be improved. |
| Approach: | They propose to use different types of model architectures to improve extractive summarization systems. |
| Outcome: | The proposed framework achieves state-of-the-art on CNN/DailyMail by a large margin based on observations and analysis. |
Copied to clipboard
| Challenge: | Large language models exhibit limitations when handling complex mathematical reasoning and logical inference tasks. |
| Approach: | They propose a sparsification strategy to reduce token costs within Multi-agent Debate (MAD) this strategy minimizes ineffective exchanges of information and unproductive discussions among agents . |
| Outcome: | The proposed approach reduces token costs by up to 94.5% while maintaining performance degradation below 2.0%. |
Copied to clipboard
| Challenge: | Direct Preference Optimization (DPO) has proven effective in aligning LLM behavior with human preferences across various tasks, but is limited in multi-turn social interactions. |
| Approach: | They propose a method which dynamically selects key segments within interactions to optimize multi-turn agent behavior. |
| Outcome: | The proposed methods outperform existing methods and proprietary LLMs on the SOTOPIA benchmark and show that they can improve social intelligence. |
Copied to clipboard
| Challenge: | Prompt-based tuning for pre-trained language models has shown its effectiveness in few-shot learning. |
| Approach: | They propose a prototypical verbalizer which learns prototype vectors as verbalizes by contrastive learning. |
| Outcome: | The proposed verbalizer outperforms existing verbalizing methods on topic classification and entity typing tasks. |
Copied to clipboard
| Challenge: | Pre-trained language models alleviate segmentation ambiguity and out-of-vocabulary (OOV) words. |
| Approach: | They propose a semisupervised neural method which distills knowledge from unlabeled data to a student model to improve both in-domain and out-of-domain CWS. |
| Outcome: | The proposed method can keep practicability of the lightweight student model and improve segmentation effectively on downstream Chinese NLP tasks. |
Copied to clipboard
| Challenge: | Existing methods to protect privacy of sensitive data are differential privacy (DP) and DP is used to protect users from privacy leakage. |
| Approach: | They propose an LDP-based Dynamic Text sanitization for privacy-preserving LLM inference that dynamically constructs semantic-aware adjacency lists of sensitive tokens to sample non-sensitive tokens for perturbation. |
| Outcome: | The proposed model excels on three datasets. |
Copied to clipboard
| Challenge: | Large Language Models have demonstrated remarkable capabilities in natural language understanding, reasoning, and generation. |
| Approach: | They present a comprehensive synthesis of large language models and their applications . they dissect a four-module agent architecture and review representative designs . |
| Outcome: | The proposed models address fundamental challenges in traditional recommender systems . they include limited comprehension of complex user intents, insufficient interaction capabilities . |
Copied to clipboard
| Challenge: | Generating long-term texts using artificial intelligence has always been a challenge . however, the generated novels exhibit poor logical coherence and appeal in their plots and deficiencies in character and event depiction, ultimately compromising the overall narrative quality. |
| Approach: | They propose a method for extracting excelsior and expanding from novel data to generate arbitrarily long novels using large language models. |
| Outcome: | The proposed method produces high-quality long-form novels with a high level of logical coherence and appeal despite the use of large language models. |
Copied to clipboard
| Challenge: | Existing approaches for goal-oriented dialogue policy learning focus on the target agent policy and treat the opposite agent policy as part of the environment. |
| Approach: | They propose a framework for policy learning in goal-oriented dialogues that uses the opposite agent's policy estimation to improve the target agent by regarding it as part of the target policy. |
| Outcome: | The proposed framework shows superior performance over state-of-the-art models on cooperative and competitive dialogue tasks. |
Copied to clipboard
| Challenge: | Emotional support is a crucial ability for many conversation scenarios, including social interactions, mental health support, and customer service chats. |
| Approach: | They propose an Emotional Support Conversation task and an ESC Framework to train emotional support into dialog systems. |
| Outcome: | The proposed framework provides an example of an Emotional Support Conversation task and shows that it is more effective than existing models. |
Copied to clipboard
| Challenge: | Large Action Models (LAMs) face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. |
| Approach: | They propose a framework for online exploration of agentic tasks with high-quality feedback . they use a dynamic task query generator and an extensive collection of tools to create a high-level feedback environment for LLM Agents. |
| Outcome: | The proposed framework achieves 49.3% performance improvement over baselines on toolbench and CRMArena. |
Copied to clipboard
| Challenge: | X-Eval is a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users. |
| Approach: | They introduce a two-stage instruction tuning framework to evaluate text in both seen and unseen aspects customized by end users. |
| Outcome: | The proposed framework improves the model’s ability to follow evaluation instructions and enhances the learning stage to better assess text quality. |
Copied to clipboard
| Challenge: | Existing approaches to domain adaptation only use reliable pseudo instances, i.e., pseudo instances with high prediction confidence, to retrain the model. |
| Approach: | They propose a domain adversarial learning enhanced self-training framework that uses meta-learning to estimate the importance of each pseudo instance and a meta constructor to construct the meta-validation set. |
| Outcome: | The proposed framework reduces label noise and preserves hard examples while maintaining accuracy. |
Copied to clipboard
| Challenge: | Existing models lack multimodal understanding capabilities, resulting in closed-source model that does not support multimodal interleaved sequences. |
| Approach: | They propose a foundation model built on multimodal tokens capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. |
| Outcome: | The proposed model is able to understand speech, text, images, and videos in an end-to-end, autoregressive manner. |
Copied to clipboard
| Challenge: | Alignment of large language models (LLM) is a process that ensures the model’s responses to user prompts align with human intentions and social values. |
| Approach: | They propose an alignment method based on a two-agent game consisting of an adversarial agent and a defensive agent. |
| Outcome: | The proposed method improves on a two-agent game with an adversarial agent and a defensive agent. |
Copied to clipboard
| Challenge: | Large language models exhibit cultural bias from over-represented viewpoints in training data, yet cultural alignment remains a challenge due to limited cultural knowledge and a lack of exploration into effective learning approaches. |
| Approach: | They propose a cost-efficient method for fine-tuning large language models on native speakers’ word-association norms and a preference optimization method to improve cultural alignment. |
| Outcome: | The proposed model trains Llama-3.1-8B and Qwen-2.5-7B on native speakers’ word-association norms and shows that such associations capture cultural knowledge. |
Copied to clipboard
| Challenge: | Recent studies have explored fine-tuning Large Language Models with synthetic data to enhance their long-context capabilities. |
| Approach: | They propose a framework that leverages a Multi-Armed Bandit rollout strategy to identify the most informative chunks from the given long context for sampling high-quality and diverse responses. |
| Outcome: | The proposed framework achieves 4% improvement on long-context reasoning benchmarks on Llama and Qwen. |
Copied to clipboard
| Challenge: | Existing methods for calculating distinct scores have evident biases that assign higher penalties to longer sequences. |
| Approach: | They propose to scale the number of distinct tokens based on their expectations. |
| Outcome: | The proposed metric removes evident biases in the original distinct score . the proposed meter correlates better with human judgment in evaluating response diversity . |
Copied to clipboard
| Challenge: | Multi-agent systems (MAS) inherit general task-solving and instruction-following capabilities, but their interconnectivity introduces significant security risks. |
| Approach: | They propose a framework that reshapes the flow of malice via risk-aware topological evolution. |
| Outcome: | Empirically, TopoSHIELD reduces toxicity by 58% on GPT-4o while preserving high utility (>90% success rate). |
Copied to clipboard
| Challenge: | Existing approaches for event extraction focus on sentence-level event extraction, but they lack a broader view of the document context. |
| Approach: | They build graphs with candidate event filler extractions enriched by sentential embeddings as nodes and use graph attention networks to identify event regions in a document and aggregate event information. |
| Outcome: | The proposed method performs well on two languages and shows that it is faster than previous methods. |
Copied to clipboard
| Challenge: | Scientific innovation is driven by detailed workflows, which include critical steps such as contextualizing literature, generating ideas, validating ideas, and planning new research. |
| Approach: | They propose to use large language models to extract five key aspects from scientific publications to optimize scientific workflows. |
| Outcome: | The proposed dataset includes more than 152,000 peer-reviewed publications from 17 leading computer science conferences spanning the past 50 years. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating LLMs’ tool usage face several limitations: limited evaluation scenarios, lacking assessments in real multi-turn dialogue contexts; narrow evaluation dimensions, with insufficient detailed assessments of how LLM use tools; and reliance on LLM or real API executions for evaluation, which introduces significant overhead. |
| Approach: | ACEBench is a benchmark for evaluating tool usage in Large Language Models . it categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. |
| Outcome: | ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. |
Copied to clipboard
| Challenge: | Existing methods to improve factuality of large language models (LLMs) rely on human-engineered instructions. |
| Approach: | They propose a retrieval-augmented generation framework that trains the model with distractor-aware QA pairs mixing gold evidence with subtle distractor passages and instills reasoning-centric habits that make the LLM plan, rationalize, and synthesize without extensive human engineered instructions. |
| Outcome: | The proposed framework outperforms state-of-the-art solutions across 12 open-book RAG QA benchmarks and is being deployed in production. |
Copied to clipboard
| Challenge: | Text-to-Image Synthesis (TIS) aims to generate images based on textual inputs . but, current diffusion-based models lack entity knowledge and low inference speed . |
| Approach: | They propose a framework for training and deploying latent diffusion models with rich entity knowledge injected and optimized networks. |
| Outcome: | The proposed framework improves image quality and inference speed and can be used in industrial applications. |
Copied to clipboard
| Challenge: | Existing evaluations treat visual understanding and generation in isolation or overlook tasks that inherently couple them. |
| Approach: | They propose a benchmark that examines the bidirectional synergy between generation and understanding across eight reasoning-centric domains. |
| Outcome: | The proposed model systematically unfolds the bidirectional synergy between generation and understanding across eight reasoning-centric domains. |
Copied to clipboard
| Challenge: | Dialogue embeddings are a critical prerequisite for semantically understanding dialogues. |
| Approach: | They propose a self-guided contrastive learning approach called dial2vec that captures interaction patterns between interlocutors and leverages them to guide the learning of the embeddings corresponding to each interlocuter. |
| Outcome: | The proposed approach achieves 8.7, 9.0, and 13.8 points absolute improvements over the strongest baseline on the three evaluation tasks respectively. |
Copied to clipboard
| Challenge: | Existing methods for text generation still suffer from incoherence problems . Neural sequence-to-sequence (seq2sequ) models generate fluent results . |
| Approach: | They propose a novel generation framework that leverages autoregressive self-attention mechanism to conduct content planning and surface realization dynamically. |
| Outcome: | The proposed framework outperforms baseline models and generates more coherent texts with richer contents. |
Copied to clipboard
| Challenge: | Conventional calibration methods treat answer correctness as binary and do not work for long-form generation where an answer can be partially correct. |
| Approach: | They propose a framework where correctness of LLMs' responses and associated confidence levels are treated as distributions across a range of scores. |
| Outcome: | The proposed framework treats the correctness of the LLMs’ responses and their associated confidence levels as distributions across a range of scores. |
Copied to clipboard
| Challenge: | Existing Process Reward Models (PRMs) output evaluation scores directly, limiting both learning efficiency and evaluation accuracy. |
| Approach: | They propose a Reasoning-Driven Process Reward Modeling (R-PRM) which activates inherent reasoning to enhance process-level evaluation. |
| Outcome: | The proposed model outperforms baseline models on ProcessBench and PRMBench by 13.9 and 8.5 F1 scores. |
Copied to clipboard
| Challenge: | Existing work on contextual speech recognition (ASR) systems focuses on recognizing words that are not frequently seen in training data, such as rare words, but word error rate on rare words remains over 20%. |
| Approach: | They propose to use public-domain earnings calls and supplementary materials to evaluate contextual ASR approaches grounded on real-world applications. |
| Outcome: | The proposed frameworks are noisier than artificially synthesized contexts that contain the ground truth, yet still make great room for future improvement of contextual ASR technology. |
Copied to clipboard
| Challenge: | Existing benchmarks for Emotional Intelligence (EI) focus on emotion recognition, neglecting essential EI capabilities. |
| Approach: | They propose a benchmark that proposes a comprehensive definition for machine EI . they propose 400 hand-crafted questions in English and Chinese to evaluate EI. |
| Outcome: | The proposed benchmarks focus on emotion recognition, neglecting EI capabilities . they are constructed from existing datasets, which include frequent patterns and errors . the proposed benchmark includes questions in English and Chinese that require thorough reasoning and understanding . |
Copied to clipboard
| Challenge: | Existing multimodal benchmarks overlook linguistic and visual ambiguities, authors say . ambiguity resolution between modalities is lacking in multimodal large language models . |
| Approach: | They propose a benchmark to evaluate multimodal ambiguity resolution across multilingual and cross-modal scenarios. |
| Outcome: | a new benchmark evaluates multimodal ambiguity resolution across multilingual and cross-modal scenarios . the benchmark shows that MLLMs can resolve ambiguities in image-text alignment . however, existing benchmarks often overlook linguistic and visual ambiguties . |
Copied to clipboard
| Challenge: | Existing studies show that large language models (LLMs) can handle multilingual machine translation (MMT) However, the multilingual translation ability of LLMs remains under-explored. |
| Approach: | They evaluate eight popular LLMs including ChatGPT and GPT-4 to determine their performance in multilingual machine translation. |
| Outcome: | The proposed model can generate moderate translation even on zero-resource languages and cross-lingual exemplars can provide better task guidance for low-resourced translation than exemplar in the same language pairs. |
Copied to clipboard
| Challenge: | Existing methods address this by adding intrinsic rewards, but they fail to provide meaningful guidance in long-horizon decision-making tasks with large state and action spaces lacking purposeful exploration. |
| Approach: | They propose a multi-modal model-based RL approach that integrates the proposed hinting subgoals into the model rollouts to encourage goal discovery and reaching in challenging tasks. |
| Outcome: | The proposed model outperforms existing methods in challenging, sparse-reward environments such as HomeGrid, Crafter, and Minecraft by 41.8%, 21.1%, and 9.9%. |
Copied to clipboard
| Challenge: | Existing medical benchmarks for diagnostic reasoning are limited in their ability to perform complex tasks. |
| Approach: | They propose to benchmark diagnostic capabilities of large language models to assess their accuracy and generalization bottlenecks. |
| Outcome: | The proposed model achieves 45.82%, 31.09%, and 17.79% accuracy, compared to current models, o3-mini, e1 and DeepSeek-R1 . |
Copied to clipboard
| Challenge: | Existing feature alignment methods are susceptible to task interference during training. |
| Approach: | MONTROSE is a cross-domain rumor detection method that generates high-quality synthetic data for the target domain and a domain-sharpness-aware approach to train models with these synthetic data. |
| Outcome: | Experiments show that MONTROSE improves in cross-domain rumor detection. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks, but their ability to extrapolate from image sequences has been less investigated. |
| Approach: | They propose a new benchmark to assess MLLMs’ sequential image reasoning abilities. |
| Outcome: | The proposed benchmark features 4,761 diverse image sequences with varying lengths. |
Copied to clipboard
| Challenge: | Personalized dialogue generation is a popular approach for conversational AI applications . however, persona profiles may not provide comprehensive descriptions of the persona . |
| Approach: | They propose a method that leverages persona profiles and dialogue context to generate personalized dialogues by leveraging personas and persona profile. |
| Outcome: | The proposed method outperforms baselines on the CONVAI2 dataset . it is expected to generate personalized dialogues based on persona profiles and dialogue context . |
Copied to clipboard
| Challenge: | Existing schema linking methods are not able to handle complex SQL queries. |
| Approach: | They propose a new algorithm that transforms SQL queries into grammatically verifiable sub-queries which are arranged sequentially to reflect single-hop reasoning steps. |
| Outcome: | The proposed algorithm achieves significant performance gains on the BIRD dataset and surpasses schema linking methods at comparable or better cost. |
Copied to clipboard
| Challenge: | Existing methods for MLLMs struggle with fine-grained temporal reasoning . despite advances in video understanding, current methods struggle with time-sensitive tasks . |
| Approach: | They propose a time-stamp-aware multi-segment grounding method that enhances temporal understanding by introducing timestamps. |
| Outcome: | The proposed method outperforms existing methods on time-sensitive tasks and generalizes well across diverse temporal understanding scenarios. |
Copied to clipboard
| Challenge: | Existing methods for Chinese word segmentation have high performance on benchmarks but are limited by the small-scale annotated corpus. |
| Approach: | They propose a framework that incorporates a lexicon-based graph convolutional network into the Transformer encoder to improve Chinese word segmentation (CWS) Chinese word is an essential and pre-processing step for many downstream NLP tasks. |
| Outcome: | The proposed framework captures the information of candidate words and improves performance on benchmarks and datasets. |
Copied to clipboard
| Challenge: | Existing work on how to effectively capture multi-document relationships remains an open question . Existing techniques to mitigate this problem include hierarchical summarization of semantically related chunks or integrating Knowledge Graphs (KGs). |
| Approach: | They propose a method which constructs a local knowledge graph from retrieved documents . they use propositional claims to construct a knowledge graph and contextualize a small language model . |
| Outcome: | The proposed method outperforms RAG on biomedical benchmarks and is generalizable and effective. |
Copied to clipboard
| Challenge: | Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored. |
| Approach: | They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain. |
| Outcome: | MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows. |
Copied to clipboard
| Challenge: | Currently, knowledge graphs are decoupled from their downstream application, resulting in suboptimal graph structures. |
| Approach: | They propose a framework to directly optimize KG construction for task performance using Reinforcement Learning (RL). |
| Outcome: | The proposed framework improves performance across multiple QA benchmarks and consistently achieves significant performance gains over task-agnostic baseline graphs. |
Copied to clipboard
| Challenge: | Recent studies show that encoding more syntactic information does not lead to better performance. |
| Approach: | They propose a method to optimize pareto-optimal models by formalizing it as a multi-objective optimization problem. |
| Outcome: | The proposed method is better than a baseline method on two NLP tasks. |
Copied to clipboard
| Challenge: | C2Rust is a system programming language that enforces strict memory and type safety guarantees. |
| Approach: | They propose a raw pointer rewriting technique that lifts raw pointers in individual functions to appropriate Rust data structures. |
| Outcome: | The proposed technique eliminates 18.57% of local raw pointers and improves memory safety on 28 real-world C projects. |
Copied to clipboard
| Challenge: | Existing prompt tuning methods only introduce prompts at the input layer, limiting performance and leaving large room for improvement. |
| Approach: | They propose a method that involves tuning a small set of soft prompts for pre-trained language models. |
| Outcome: | The proposed method outperforms state-of-the-art methods with pre-trained models on the SuperGLUE benchmark. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can be used in psychotherapy to overcome challenges such as shame, distrust, and resource scarcity. |
| Approach: | They propose a cognitive reframing therapy method that uses empathetic dialogue to address deep-rooted negative thoughts and fosters rational, balanced perspectives. |
| Outcome: | The proposed model outperforms other models in terms of empathy, guidance, and logical coherence, demonstrating its effectiveness and potential positive impact on psychotherapy. |
Copied to clipboard
| Challenge: | Existing general-domain benchmarks do not capture complexity of real-world judicial cognition and decision-making. |
| Approach: | They propose a benchmark specifically designed to evaluate LLM Agents in the legal domain. |
| Outcome: | The proposed benchmark includes 17 corpora from real-world legal scenarios and provides 37 tools for interacting with external knowledge. |
Copied to clipboard
| Challenge: | Existing length control methods involve fine-tuning the parameters of LLMs, which is inefficient and suboptimal for practical use. |
| Approach: | They propose an iterative sampling framework that regulates LLMs to generate length-constrained text without modifying the underlying parameters. |
| Outcome: | The proposed method achieves 100% success rates on Llama3.1 tasks with minimal additional computational overhead. |
Copied to clipboard
| Challenge: | InfiMM-WebMath-40B is a dataset of interleaved image-text documents . it consists of 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
| Approach: | InfiMM-WebMath-40B is a high-quality dataset of interleaved image-text documents . it contains 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
| Outcome: | InfiMM-WebMath-40B is a high-quality dataset of interleaved image-text documents . it consists of 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
Copied to clipboard
| Challenge: | Existing evaluation methods rely on rigid pipelines that overlook user needs and provide numerical results without clear explanations. |
| Approach: | They propose an evaluation framework that employs human-like strategies for efficient, dynamic, multi-round evaluations using only a few samples per round. |
| Outcome: | The evaluation agent framework reduces evaluation time to 10% of traditional methods while delivering comparable results. |
Copied to clipboard
| Challenge: | Large language model agents rely on in-context policy documents to act as effective user assistants. |
| Approach: | They propose an agentic benchmark generator with Controllable Complexity in agent policy across four levels to evaluate agents under increasing complexity. |
| Outcome: | The proposed method outperforms the baseline in data-sparse and high-complexity settings. |
Copied to clipboard
| Challenge: | Existing methods for supervised fine-tuning focus on unit test feedback to construct preference pairs. |
| Approach: | They propose a preference alignment framework that mimics human iterative debugging to refine Code LLMs. |
| Outcome: | Experiments show that Preference Learning improves on BigCodeBench and BigCodeBind tasks. |
Copied to clipboard
| Challenge: | Long-context Multimodal Large Language Models (MLLMs) require substantial computational resources for inference . the growth of their multimodal Key-Value (KV) cache challenges memory and time efficiency. |
| Approach: | They propose a fine-tuning-free approach that efficiently reduces the multimodal KV cache size while maintaining performance comparable to a full cache. |
| Outcome: | The proposed method reduces the multimodal KV cache size while maintaining performance comparable to a full cache. |
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) focus on task completion, but neglect a crucial capability: the ability to devise and adjust cost-optimal plans in response to changing environments. |
| Approach: | They propose a scalable, cost-centric benchmark to evaluate agents’ economic reasoning and replanning abilities. |
| Outcome: | Evaluating leading open-sourced and proprietary models on CostBench reveals a substantial gap in cost-aware planning . |
Copied to clipboard
| Challenge: | Existing methods to parse natural language into structured logical expressions have limitations due to paucity of labeled data. |
| Approach: | They propose a scoring model to automatically learn a model-based reward . they also propose introducing a Chinese-PL/FOL dataset to compensate for paucity of labeled data . |
| Outcome: | The proposed model outperforms competitors on several datasets. |
Copied to clipboard
| Challenge: | Existing quantization solutions are integer-based and struggle with bit widths below 8 bits. |
| Approach: | They propose a method for quantizing weights and activations in large language models down to 4-bit floating-point values in a post-training manner. |
| Outcome: | The proposed method outperforms existing methods on common sense zero-shot reasoning tasks by 12.7 points. |
Copied to clipboard
| Challenge: | Existing work on long document visual question answering is based on Retrieval-Augmented Generation (RAG) where textual or visual content is encoded into embeddings and relevance is determined by similarity scores with respect to the original query. |
| Approach: | They propose a framework that employs an agentic, vision-aware workflow to address long document visual question answering through iterative information discovery and synthesis. |
| Outcome: | The proposed framework outperforms existing RL systems by 10.4% on the MMLongbench-Doc benchmark and demonstrates superior training performance over GRPO. |
Copied to clipboard
| Challenge: | Existing studies on K-LLMs systems focus on declarative knowledge and procedural knowledge (rules) . |
| Approach: | They propose to build a toolkit that supports comprehensive heterogeneous knowledge collaborative enhancement for Large Language Models (LLMs). |
| Outcome: | The proposed toolkit provides unified knowledge integration and joint knowledge retrieval methods to achieve more comprehensive heterogeneous knowledge collaborative enhancement. |
Copied to clipboard
| Challenge: | Knowledge graph embedding is a new form of knowledge graphing that allows for better link prediction. |
| Approach: | They propose to use relational embedding to fit symmetry/antisymmetry and combination relationships. |
| Outcome: | The proposed model can fit symmetry/antisymmetry and combination relationships. |
Copied to clipboard
| Challenge: | Context gates are effective to control the contributions from the source and target contexts in the recurrent neural network (RNN) based neural machine translation. |
| Approach: | They propose a method to identify source and target contexts and introduce a gate mechanism to control the contributions from source and targets in the advanced Transformer architecture. |
| Outcome: | The proposed model achieves an averaged gain of 1.0 BLEU score over a strong transformer baseline. |
Copied to clipboard
| Challenge: | Existing methods for IE are task-specific, resulting in specialized and isolated approaches for different tasks. |
| Approach: | They propose a method to retrieve task-specific knowledge from pretrained language models to enhance universal IE by using a Meta-Pretraining Algorithm. |
| Outcome: | The proposed method achieves the new state-of-the-art on 4 IE tasks, 12 datasets under fully-supervised, low-resource and few-shot scenarios. |
Copied to clipboard
| Challenge: | Text-to-speech (TTS) performance has improved with the advent of denoising Diffusion Probabilistic Models . however, perceived quality of audio depends on content, pitch, rhythm, and energy . |
| Approach: | They propose a visual TTS model with scalable diffusion transformers that complement phoneme sequences with visual information to generate high-perceived audio. |
| Outcome: | The proposed model outperforms existing models regardless of visibility of the scene . it can generate high-perceived audio, opening up new avenues for AR and VR applications . |
Copied to clipboard
| Challenge: | Existing tasks for evaluating story understanding and generation focus on reasoning plots from context, but they focus on bridging plots with implied morals. |
| Approach: | They propose two understanding tasks and two generation tasks to assess machines' ability to bridge story plots and implied morals. |
| Outcome: | The proposed tasks are based on a dataset of Chinese and English moral stories . they show that the proposed models can perform better than existing models . |
Copied to clipboard
| Challenge: | Current medical benchmarks have limitations in question design, data sources and evaluation methods. |
| Approach: | They propose a new benchmark covering five core medical areas . it includes 2,996 questions created from real-world electronic health records . |
| Outcome: | The proposed model covers five core medical areas and includes 2,996 questions created from real-world electronic health records and expert-designed clinical scenarios. |
Copied to clipboard
| Challenge: | Multi-intent utterances processing remains a persistent challenge due to intricate intent-slot dependencies and semantic ambiguities. |
| Approach: | They propose a label-aware contrastive attention network (LCAN) that integrates label-based attention and contrastive learning strategies to improve semantic understanding and generalization in multi-intent scenarios. |
| Outcome: | The proposed model improves intent recognition and slot filling performance in multi-intent dialogue systems. |
Copied to clipboard
| Challenge: | Existing methods for maximizing preference optimization on all available tokens are noisy and inefficient. |
| Approach: | They propose a selective alignment strategy that centers on efficient key token selection without strong, fine-grained supervision signals. |
| Outcome: | The proposed strategy outperforms baseline methods on three benchmarks with up to 60% reduction in training hours. |
Copied to clipboard
| Challenge: | Recent advances in large language model (LLM) agents have accelerated deployment of multi-agent systems for complex tasks. |
| Approach: | They propose an open-source toolkit for instantiating, probing, and measuring emergent risks in LLM-based multi-agent systems under controlled conditions. |
| Outcome: | The proposed toolkit is based on a structured topology–environment–protocol–agent–task quintuple enabling reproducible studies of how communication structure, coordination mechanisms, and incentives shape system-level risks. |
Copied to clipboard
| Challenge: | Existing text generation methods only use heuristic replacement strategies or language models to generate replacement words at the word level. |
| Approach: | They propose a search and learning framework for Adversarial Text Generation by Search and Learning to evaluate the robustness of natural language processing models. |
| Outcome: | The proposed methods are significantly superior to the most advanced methods in terms of attack efficiency and adversarial text quality. |
Copied to clipboard
| Challenge: | Personality is a crucial factor that shapes human communication patterns, thereby regulating the personalities of large language models (LLMs). |
| Approach: | They propose a method that uses an Unsupervisedly-Built Personalized Lexicon (UPL) during the decoding phase to manipulate LLM’s personality traits. |
| Outcome: | The proposed method can modulate the personality expression of large language models by dynamically altering their predicted probability of upcoming words in a pluggable fashion. |
Copied to clipboard
| Challenge: | Existing methods for empathetic response generation ignore hierarchical relationships between different factors, leading to a weak ability of empathy modeling. |
| Approach: | They propose a multi-factor hierarchical framework for empathetic response generation which models the above three key factors in a hierarchically structured way. |
| Outcome: | The proposed model generates more empathetic responses than previous methods. |
Copied to clipboard
| Challenge: | Existing models for LLM role-playing lack high-quality datasets with explicit reasoning traces and reliable reward signals aligned with human preferences. |
| Approach: | They propose a unified framework for cognitive-level persona simulation that strictly distinguishes characters’ first-person thinking processes from LLMs’ third-person reasoning. |
| Outcome: | The proposed framework outperforms the Qwen3-32B baseline model and achieves a 30.26% and 14.97% performance on the minimax benchmarks. |
Copied to clipboard
| Challenge: | A moral dialogue system aligned with users’ values could enhance conversation engagement and user connections. |
| Approach: | They propose a framework to train and evaluate moral dialogue systems based on communication mechanisms of morality and a method to construct moral discussions between simulated users and the dialogue system. |
| Outcome: | The proposed framework can train and evaluate moral dialogue systems based on simulated users and their values . |
Copied to clipboard
| Challenge: | Existing toolsets that use large language models are limited to single-task settings. |
| Approach: | They propose a framework that dynamically constructs and evolves a hierarchical graph of reusable tools across multiple scenarios. |
| Outcome: | The proposed framework achieves up to 4.3 faster milestone completion in Minecraft compared to the previous state-of-the-art method and provides an average improvement of 9.23% over existing tool-making methods in code generation tasks and 10.03% in agent tasks. |
Copied to clipboard
| Challenge: | Chinese Grammatical Error Correction (CGEC) is a challenging NLP task and a common application in human daily life. |
| Approach: | They propose a linguistic rules-based approach to construct large-scale CGEC training corpora with automatically generated grammatical errors. |
| Outcome: | The proposed method improves performance of existing CGEC models and the benchmark is excellent resource for further development. |
Copied to clipboard
| Challenge: | Standard evaluation metrics, e.g., BLEU, TER and METEOR, focus on the quality of translations at the sentence level and do not consider discourse-level features. |
| Approach: | They propose to use a metric to take discourse coherence into consideration by categorizing discourse-related spans and calculating the similarity-based F1 measure of categorized spans. |
| Outcome: | The proposed metric possesses better selectivity and interpretability at the document-level, and is more sensitive to document- level nuances. |
Copied to clipboard
| Challenge: | Existing methods for ERC lack interpretability and shallow semantics capture deep semantics. |
| Approach: | They propose a Fast-Slow thinking framework for Emotion Recognition in Conversation . they use fine-grained emotion reasoning chains to capture deep semantics . |
| Outcome: | The proposed framework achieves state-of-the-art in explanation and judgment on a benchmark dataset. |
Copied to clipboard
| Challenge: | Existing methods of prompt-tuning for Aspect-based Sentiment Analysis (ABSA) are crude and simple. |
| Approach: | They propose a Syntax-aware Enhanced Prompt method which mines syntactic information related to aspect words from the syntaktic dependency tree. |
| Outcome: | The proposed method exploits the syntactic knowledge embedded in PLMs and achieves favorable results on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing studies find VH instances only in existing image datasets, which results in biased understanding of MLLMs’ performance under VH. |
| Approach: | They propose a tool called VHTest to generate a diverse set of VH instances from existing image datasets and a text-to-image generative model to generate VH images based on the text descriptions. |
| Outcome: | The proposed tool finds VH instances in existing image datasets and generates images based on the text descriptions. |
Copied to clipboard
| Challenge: | Existing Chinese preference datasets suffer from limited scale, restricted domain coverage, and insufficiently rigorous data validation. |
| Approach: | They propose an LLM-based data annotation pipeline with no human intervention to annotate Chinese preference datasets. |
| Outcome: | The proposed pipeline outperforms existing Chinese preference datasets on AlignBench and Chinese Reward Benchmark. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks do not support arbitrarily interleaved images and text for both inputs and outputs. |
| Approach: | They propose to use a benchmark to evaluate interleaved text-and-image generation . they define five evaluation aspects for InterleavatedEval, a reference-free metric . |
| Outcome: | The proposed benchmarks cover a limited number of domains and use cases and lack comparableity-based metrics. |
Copied to clipboard
| Challenge: | Existing studies on in-context learning (ICL) focus on the selection of individual examples and ignore correlations among examples. |
| Approach: | They propose a method to capture positive and negative correlations using the determinantal point process . they optimize the method via kernel decomposition-based MLE to fit a constructed pseudo-labeled dataset . |
| Outcome: | The proposed method outperforms baselines in ICL example selection. |
Copied to clipboard
| Challenge: | Annotated training data is costly to obtain in many languages . |
| Approach: | They propose a semantic contrastive loss to align parallel sentences that share the same semantics in different languages and a language contrastive gain to leverage parallel sentence pairs to remove language-specific information from non-parallel corpora. |
| Outcome: | The proposed model improves retrieval performance while requiring less computational effort. |
Copied to clipboard
| Challenge: | Existing methods for rationalization use spurious correlations in data to compose rationales and make predictions. |
| Approach: | They propose a method to discover the causal rationales by using a structural causal model. |
| Outcome: | The proposed method is based on the causal theory and validates on three real-world datasets. |
Copied to clipboard
| Challenge: | Event extraction is of practical utility in natural language processing . it is common that multiple events exist in the same sentence, causing difficulties in extracting them . |
| Approach: | They propose a framework to jointly extract multiple event triggers and arguments . they introduce syntactic shortcut arcs to enhance information flow and attention-based graph convolution networks to model graph information. |
| Outcome: | The proposed framework achieves competitive results compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing datasets are often criticized for their lack of granularity, which can mask deficiencies in basic syntactic elements that humans care about. |
| Approach: | They propose a new program translation metrics that address basic syntax errors . they propose BLUE, CodeBLUE and computation accuracy metrics which address these errors based on a highly interpretable evaluation harness. |
| Outcome: | The proposed model passes the unit tests with a 26.15% pass rate compared to previous models . |
Copied to clipboard
| Challenge: | Developing non-collaborative dialogue agents traditionally requires manual codification of expert strategies. |
| Approach: | They propose a method that formalizes expert knowledge into a Strategy Forest from raw transcripts. |
| Outcome: | The proposed method outperforms existing methods by 9%-10% in two benchmarks. |
Copied to clipboard
| Challenge: | Rapid adoption of LLMs has overshadowed the potential advantages of traditional BERT-like models in text classification. |
| Approach: | They compare BERT-like models fine-tuning, LLM internal state utilization, and LLM zero-shot inference across six datasets. |
| Outcome: | The proposed method outperforms LLMs on six challenging datasets. |
Copied to clipboard
| Challenge: | Existing strategies focus on planning the next-turn dialogue strategies, while external strategy planners focus on generating empathetic responses. |
| Approach: | They propose a Markov-driven emotion anticipation framework with emotion trajectory-aware retrieval for LLM-based ESC, which anticipates future emotion states to guide strategy planning and achieve sustained emotional support. |
| Outcome: | The proposed framework can anticipate future emotions and achieve sustained emotional support on two datasets with two models. |
Copied to clipboard
| Challenge: | Genetic Prompt combines genetic algorithms with Large Language Models to augment synthetic data generation. |
| Approach: | They propose a framework that combines genetic algorithms with LLMs to augment synthetic data generation. |
| Outcome: | The proposed framework outperforms state-of-the-art models and shows robust performance across generator models. |
Copied to clipboard
| Challenge: | Existing table understanding methods struggle with low initialization accuracy and coarse rewards in tabular contexts. |
| Approach: | They propose a three-stage RL framework that enhances multimodal table understanding through: (1) Warm-up that prompts initial perception and reasoning capabilities; (2) Perception Alignment GRPO (PA-GRPO); (3) Hint-Completion GR PO (HC-GRP); |
| Outcome: | The proposed framework outperforms existing models on held-in and held-out datasets, outperforming SFT and GRPO largely. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) generate responses that are plausible but incorrect or unsupported—commonly referred to as hallucinations. |
| Approach: | They propose a representation-level intervention framework that modulates hallucination-related features during inference by probing their encoded features. |
| Outcome: | The proposed framework reduces hallucinations while maintaining the performance and generalization capabilities of Large Vision-Language Models (LVLMs). |
Copied to clipboard
| Challenge: | Existing knowledge representation learning methods suffer from immaturity on tackling potentially-imperfect knowledge graphs and highly-imbalanced positive-negative instances during training. |
| Approach: | They propose a framework for knowledge representation learning that incorporates two functional components to achieve robust embedding for each entity/relation. |
| Outcome: | The proposed framework achieves better convergence against state-of-the-art methods on several benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have improved search engines and recommendation systems through their text understanding capabilities. |
| Approach: | They propose a token-level proximal policy optimization approach to empower LLMs to perform better in query generation through fine-tuning. |
| Outcome: | The proposed approach outperforms existing LLMs on an open-source and industrial dataset. |
Copied to clipboard
| Challenge: | Existing LLM-based tools struggle with insufficient assessment cues, weak narrative coherence, limited stylistic diversity, and poor support for creative thinking. |
| Approach: | They propose an evolutionary tree-based psychometric context generator that integrates rule-guided outline planning, sentence-level MCTS generation, MAP-Elites quality-diversity optimization and assessment-guide refiner simulation. |
| Outcome: | The proposed tool outperforms strong LLMs and structured frameworks on 7 evaluation dimensions and shows higher alignment with expert-designed contexts. |
Copied to clipboard
| Challenge: | Existing wild vision TTA methods fail to handle speech data due to the unique characteristics of high-entropy speech frames, which are unreliably filtered out even when containing crucial semantic content. |
| Approach: | They propose a method for acoustic foundation models to perform confidence-based adaptation in wild acustic test settings. |
| Outcome: | The proposed method outperforms baselines under Gaussian noise, environmental sounds, accent variations, and sung speech in the wild. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating large language models use static datasets, leading to data leakage or overlooking the complexities of multi-agent interactions. |
| Approach: | They propose a framework that evaluates the diverse capabilities of LLM agents in multi-agent dynamic environments. |
| Outcome: | The proposed framework assesses the diverse capabilities of LLM agents in multi-agent dynamic environments. |
Copied to clipboard
| Challenge: | a lack of systematic studies on the robustness of language understanding models in task-oriented dialog systems is limiting . authors propose a model-agnostic toolkit LAUG to approximate natural language perturbations . |
| Approach: | They propose a model-agnostic toolkit LAUG to approximate natural language perturbations for testing the robustness of language understanding models in task-oriented dialog systems. |
| Outcome: | The proposed toolkit reveals critical robustness issues in state-of-the-art models. |
Copied to clipboard
| Challenge: | Using reinforcement learning from human feedback, large language models perform poorly when applied to colloquial subtitle translation tasks. |
| Approach: | They propose an adversarial training framework that iteratively updates the offline reward model and the online LLM to improve training outcomes. |
| Outcome: | The proposed training framework significantly improves upon translation baselines. |
Copied to clipboard
| Challenge: | Translationese is a linguistic property that is often introduced in the translation process that is different from those of original texts. |
| Approach: | They propose to use synthesized translations and translations in the wild to evaluate T-index's generalizability in cross-domain settings and its validity against human judgments. |
| Outcome: | The proposed measure can generalize to unseen genres, authors, and language pairs. |
Copied to clipboard
| Challenge: | Existing knowledge poisoning attacks against RAG systems require multiple poisoned documents or can only function effectively on simplistic queries. |
| Approach: | They propose a more realistic knowledge poisoning attack that poisons only a single document while remaining effective for complex multi-hop questions involving complex relationships between multiple elements. |
| Outcome: | The proposed attack achieves success by poisoning only a single document while remaining effective for complex multi-hop questions involving complex relationships between multiple elements. |
Copied to clipboard
| Challenge: | Low-resource languages, like Tibetan, remain underrepresented in large language models' evaluations. |
| Approach: | They propose a Tibetan Language Understanding Evaluation Benchmark to assess LLMs' proficiency in Tibetan . they use a multi-task understanding benchmark and a safety benchmark to evaluate models . |
| Outcome: | The proposed benchmark shows that most large language models perform below the random baseline, especially in Tibetan language processing. |
Copied to clipboard
| Challenge: | Recent studies in deep learning have shown significant progress in named entity recognition (NER) . however, most existing works assume clean data annotation, while real-world data typically involve a large amount of noises. |
| Approach: | They propose a confidence estimation approach for named entity recognition using noisy labels using local and global independence assumptions. |
| Outcome: | The proposed method marginalizes out labels of low confidence with a CRF model and integrates it into a self-training framework for boosting performance. |
Copied to clipboard
| Challenge: | Existing reasoning based on chains of thought (CoTs) fails to find logical connections between reasoning steps . |
| Approach: | They propose a method to match KG reasoning chains with CoTs based on semantic similarity . they use a knowledge graph to find relevant information "within" each reasoning step . |
| Outcome: | The proposed method outperforms baselines on multi-answer questions with 5.1% improvement over baselines. |
Copied to clipboard
| Challenge: | Chinese Spell Checking (CSC) aims to detect and correct Chinese spelling errors. |
| Approach: | They propose a framework which renders Chinese Spell Checking model to learn heterogeneous knowledge from the dictionary in terms of phonetics, vision, and meaning. |
| Outcome: | The proposed framework renders the CSC model to learn heterogeneous knowledge from the dictionary in terms of phonetics, vision, and meaning. |
Copied to clipboard
| Challenge: | Membership inference attacks are a promising tool for auditing training data of LLMs . existing methods rely on the assumption that LLM's assign higher confidence scores to training samples than to non-training ones. |
| Approach: | They propose a membership inference framework that can be robust against adversarial MIAs. |
| Outcome: | The proposed framework can be robust against adversarial MIA methods and AIGT detectors while maintaining the performance of baselines. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have led to significant success in using LLMs as agents. |
| Approach: | They propose a cognitive framework that incorporates first-order and second-order perspective transitions into LLMs to enhance their ability to identify and counteract deceptive information. |
| Outcome: | The proposed framework enhances LLMs’ ability to identify and counteract deceptive information without extra fine-tuning and data. |
Copied to clipboard
| Challenge: | Pretraining-based (PT) evaluation metrics are not effective for training grammatical error correction systems. |
| Approach: | They propose a pretraining-based GEC evaluation metric which only uses PT-based metrics to score the corrected parts of the system. |
| Outcome: | The proposed evaluation metric outperforms existing methods on a CoNLL14 evaluation task. |
Copied to clipboard
| Challenge: | Multimodal large language models (MLLMs) have made rapid progress in perception and alignment, but their reasoning ability often lags behind strong text-only LLMs. |
| Approach: | They propose a method that transfers reasoning knowledge in the gradient space while preserving multimodal alignment. |
| Outcome: | Experiments on multimodal reasoning benchmarks show that DRIFT outperforms naive merging and standard SFT. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are limited by their parametric knowledge, leading to hallucinations in knowledge-extensive tasks. |
| Approach: | They propose an end-to-end extract-and-restructure paradigm that leverages a single decoder-only LLM to adaptively extract query-relevant contents verbatim along with the necessary context. |
| Outcome: | Experiments show that a trained Refiner outperforms state-of-the-art RAG and compressing approaches in multiple tasks. |
Copied to clipboard
| Challenge: | Existing methods for enhancing dense retrieval with query augmentation ignore the alignment between generation and ranking objectives. |
| Approach: | They propose a unified LLM-augmented dense retrieval framework that jointly optimizes both the LLM and the retriever. |
| Outcome: | Experimental results show that ExpandR outperforms strong baselines, achieving more than 5% improvement in retrieval performance. |
Copied to clipboard
| Challenge: | Currently, most of the neural extractive summarization systems score and extract sentences individually and model the relationship between sentences. |
| Approach: | They propose to instantiate a neural extractive summarization task as a semantic text matching problem and use it to match a source document and candidate summaries in a semantic space. |
| Outcome: | The proposed framework is faster and more efficient than existing frameworks. |
Copied to clipboard
| Challenge: | Parameter Efficient Fine-Tuning (PEFT) has gained significant attention for its ability to achieve competitive results while updating only a small subset of trainable parameters. |
| Approach: | They propose a new approach to fine-tuning neural models that scales and biases the representation produced at each layer. |
| Outcome: | The proposed approach reduces the number of trainable parameters by a factor of 25,700 compared to full parameter fine-tuning and by . 32 compared with LoRA. |
Copied to clipboard
| Challenge: | Existing studies have shown that LLMs finetuned on incorrect completions can exhibit harmful behaviors, which is called emergent misalignment. |
| Approach: | They investigate whether LLMs finetuned on incorrect completions can exhibit harmful behaviors . they find that 1% of misalignment data is sufficient to decrease honest behavior . |
| Outcome: | The proposed model can be misaligned on errors within narrow domains to exhibit harmful behaviors . the proposed model is able to exhibit dishonest behavior with only 10% biased user population . |
Copied to clipboard
| Challenge: | Recent studies have evaluated and shown limitations in specific capabilities such as visual understanding, but a systematic evaluation of VLMs’ fundamental WM abilities remains absent. |
| Approach: | They propose a framework that assesses perception and prediction to provide an atomic evaluation of VLMs as WMs. |
| Outcome: | The proposed framework assesses perception and prediction abilities on 15 latest VLMs and compares them to human-level models. |
Copied to clipboard
| Challenge: | Existing tagging systems that use sentence-level data are not well understood. |
| Approach: | They propose a larger-context approach to tagging tasks that incorporates contextual information into existing tapping systems. |
| Outcome: | The proposed aggregators improve on four tagging tasks and 13 datasets. |
Copied to clipboard
| Challenge: | Recent singing-voice-synthesis methods lack ability to control style attributes of synthesized singing. |
| Approach: | They propose a singing-voice-synthesis method that enables attribute controlling on singer gender, vocal range and volume with natural language. |
| Outcome: | The proposed method achieves favorable control ability and audio quality. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data. |
| Approach: | They propose a method to disentangle risks through step-by-step reasoning within multimodal inputs. |
| Outcome: | The proposed approach improves safety alignment in MLLMs by fine-tuning and iterative Reinforcement Learning from AI feedback. |
Copied to clipboard
| Challenge: | State-of-the-art large multimodal models face challenges when processing high-resolution images, as these inputs are converted into enormous visual tokens, many of which are irrelevant to the downstream task. |
| Approach: | They propose a multi-turn grounding-based policy optimization framework that enables LMMs to iteratively focus on key visual regions by automatically cropping sub-images based on model-predicted grounding coordinates within a multiple-turn conversation framework. |
| Outcome: | The proposed framework improves on Qwen2.5-VL-7B with 21K samples and surpasses OpenAI’s o1 and GPT-4o models on the out-of-distribution (OOD) V* Bench. |
Copied to clipboard
| Challenge: | Existing Large Language Model (LLM) enabled agents lack flexibility to respond to users’ varying needs and preferences. |
| Approach: | They propose a test-time user-preference alignment strategy that optimizes the persona prompt, ensuring real-time preference alignment through textual loss feedback between simulated and ground-truth responses. |
| Outcome: | The proposed framework outperforms baseline methods in real-time and in real applications. |
Copied to clipboard
| Challenge: | Existing MDMs employ uncertainty-based decoding strategies that limit their reasoning ability and ultimately degrade generation quality. |
| Approach: | They propose a framework that regularizes uncertainty-based decoding by incorporating two complementary priors to shape global decoding trajectories and promote content informativeness. |
| Outcome: | The proposed framework outperforms existing decoding strategies by more than 7% while achieving comparable performance to autoregressive models of similar parameter scales. |
Copied to clipboard
| Challenge: | Existing knowledge demonstrates the superiority of TM-based neural machine translation only on TM specialized tasks . |
| Approach: | They propose a translation memory-based approach to machine translation using a single bilingual sentence as its TM. |
| Outcome: | The proposed approach surpasses baselines on two general tasks and improves on the TM-specialized translation tasks. |
Copied to clipboard
| Challenge: | Recent advances in slow-thinking reasoning models have shown exceptional performance in complex reasoning tasks. |
| Approach: | They propose a framework that enables models to automatically adjust Chain-of-Thought (CoT) length based on problem difficulty. |
| Outcome: | The proposed framework penalizes inefficiency on simple problems while incentivizing deep reasoning for complex ones. |
Copied to clipboard
| Challenge: | Existing question answering systems use a retriever-reader framework to answer multi-hop questions . existing models lack retrieval, selector, and reasoner capabilities . |
| Approach: | They propose a three-stage text tableQA framework which comprises of retriever, selector, and reasoner. |
| Outcome: | The proposed framework outperforms baseline methods in the few-shot setting and ranks first on the HybridQA leaderboard. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable success across diverse tasks such as instruction following, code generation, and medical diagnosis. |
| Approach: | They propose a supervised fine-tuning-based auxiliary loss for Q-value estimations during supervised refinement. |
| Outcome: | The proposed method outperforms beam search on GSM8K, MATH, and GAOKAO on reasoning benchmarks. |
Copied to clipboard
| Challenge: | MobileLLM-Flash is a family of foundation models for efficient on-device use with strong capabilities. |
| Approach: | They propose a method for designing on-device large language models under mobile latency constraints using hardware-in-the-loop architecture search. |
| Outcome: | The proposed model is amenable to industry-scale deployment and is compatible with mobile runtimes like Executorch. |
Copied to clipboard
| Challenge: | In this paper, we examine the generalization behaviour of summarization models . we propose several properties of datasets that matter for generalization . |
| Approach: | They propose several properties of datasets which matter for generalization of summarization models. |
| Outcome: | The proposed approach improves the state-of-the-art model by rethinking the model design process on a typical dataset. |
Copied to clipboard
| Challenge: | Existing approaches to grammatical error correction (GEC) are sequence-to-sequence and sequence-edit. |
| Approach: | They propose a unified decoding intervention framework that employs an external critic to assess the appropriateness of the token to be generated incrementally. |
| Outcome: | The proposed framework outperforms baselines and state-of-the-art methods on English and Chinese datasets. |
Copied to clipboard
| Challenge: | Existing research explores different text features of reply comments on word level and ignores interactions between participants. |
| Approach: | They propose a co-attention mechanism based neural network to capture interactions between participants on argument level to better model dialogical argumentation. |
| Outcome: | The proposed model outperforms state-of-the-art methods on a publicly available dataset showing that it extracts interactive argument pairs from the original post and the reply. |
Copied to clipboard
| Challenge: | HKVE selectively accepts gradient optimization results based on the distribution of attention scores across different layers, ensuring that every optimization step positively contributes to the attack. |
| Approach: | They propose a framework that selectively accepts gradient optimization results based on the distribution of attention scores across different layers and selectively takes them into account when calculating the attack success rate. |
| Outcome: | The proposed framework outperforms existing methods by achieving success rates of 75.08% on MiniGPT4, 85.84% on LLaVA and 81.00% on Qwen-VL. |
Copied to clipboard
| Challenge: | Social networking services (SNS) have experienced rapid growth, which has proposed significant challenges for platform content management and interaction quality improvement. |
| Approach: | They propose a domain-specific LLM to break the performance bottleneck of single-task baselines and establish a comprehensive foundation for social networking services. |
| Outcome: | The proposed model achieves an average improvement of 14.02% across 8 major tasks and 7.56% in bilingual evaluation benchmark, compared with baseline models. |
Copied to clipboard
| Challenge: | Existing studies on human-LLM competitive programming use scattered, application-specific human feedback. |
| Approach: | They propose a taxonomy of human feedback consolidating the entire programming process, which promotes fine-grained evaluation. |
| Outcome: | The proposed benchmark pinpoints strengths and weaknesses of existing methods and will be openly released. |
Copied to clipboard
| Challenge: | Existing approaches to resolve explicit knowledge conflicts are based on semantic decoding and auxiliary embedding. |
| Approach: | They propose a framework that adjudicates conflicts by structuring the underlying logic. |
| Outcome: | Experiments show that the proposed framework improves on existing models. |
Copied to clipboard
| Challenge: | Existing multilingual neural machine translation models perform poorly on language pairs with no parallel corpus. |
| Approach: | They propose a two-stage approach that encourages original models to acquire language-agnostic multilingual representations from new data and preserves the model architecture without introducing parameters. |
| Outcome: | The proposed approach improves performance in translation directions where existing models are weak and mitigates degeneration in the well-performing translation directions, offering flexibility in the real-world scenario. |
Copied to clipboard
| Challenge: | Existing methods for evaluating expressive speech focus on word accuracy, naturalness, signal quality, or emotional intensity at the utterance level. |
| Approach: | They propose a framework for Evaluating Expressive Appropriateness in speech that assesses whether a speech sample aligns with the underlying communicative intent implied by its discourse-level narrative context. |
| Outcome: | The proposed framework outperforms existing speech evaluation and analysis systems on a human-annotated test set. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have paved the way for complex tasks such as role-playing. |
| Approach: | They propose a framework to benchmark, elicit, and enhance role-playing abilities in Large Language Models. |
| Outcome: | The proposed framework improves role-playing abilities with 168,093 samples. |
Copied to clipboard
| Challenge: | Existing reward models concatenate contexts and responses, but they often ignore crucial segments of the context that are important for evaluating the response quality. |
| Approach: | They propose a reward model that evaluates the response quality based on a given context and assigns a rewards reward. |
| Outcome: | The proposed framework significantly improves preference modeling by increasing attention to relevant information within the context and achieves better generalizability. |
Copied to clipboard
| Challenge: | Existing methods for event extraction neglect grammatical incorrectness, structure misalignment, and semantic drifting . et al., 2004; Ahn, 2006) show that the proposed method generates more diverse text representations for event extracting compared with the state-of-the-art. |
| Approach: | They propose a framework for event extraction that generates additional training data and iteratively selects the effective subset from the generated training data. |
| Outcome: | The proposed method generates more diverse representations of training data and achieves comparable results with the state-of-the-art. |
Copied to clipboard
| Challenge: | AEGIS examines whether current models can effectively audit AI-generated images in academic papers. |
| Approach: | They propose a holistic benchmark for forensic analysis of AI-Generated academic ImageS that reveals limitations in academic image forensics. |
| Outcome: | AEGIS compared with existing benchmarks on seven academic categories and features key advances in forensic analysis. |
Copied to clipboard
| Challenge: | Existing distantly supervised entity relation extractors rely on noisy data for training and evaluation. |
| Approach: | They propose a criterion for clean instance selection based on influence functions to collect sample-level evidence for recognizing good instances. |
| Outcome: | The proposed method shows strong performance on real and synthetic noisy datasets. |
Copied to clipboard
| Challenge: | Existing sampling methods that are sensitive to temperature scaling fail to distinguish between diversity and noise. |
| Approach: | They propose a method that identifies informative tokens by eliminating noise directly in logit space and a new sampling method that is temperature-invariant. |
| Outcome: | The proposed method outperforms existing methods with significant improvements in reasoning and creative writing tasks. |
Copied to clipboard
| Challenge: | Recent years have seen the paradigm shift of Named Entity Recognition (NER) systems from sequence labeling to span prediction. |
| Approach: | They experimentally implement 154 named entity recognition models on 11 datasets and show that span prediction can serve as a system combiner to re-recognize named entities from different systems’ outputs. |
| Outcome: | The proposed model can be used to re-recognize named entities from different systems’ outputs. |
Copied to clipboard
| Challenge: | Existing approaches to finding effective predictive signals from financial data are limited by their complexity and low signal-to-noise ratio. |
| Approach: | They propose a framework that combines code-level alpha representation with LLM-driven reasoning and evolutionary search. |
| Outcome: | The proposed framework combines code-level alpha representation with LLM-driven reasoning and evolutionary search. |
Copied to clipboard
| Challenge: | Existing text-to-image systems often produce visually plausible but semantically literal outputs. |
| Approach: | They propose a structured prompting framework inspired by Conceptual Metaphor Theory . they propose to identify source–target mappings, filter projectable source attributes and select a visual realization strategy in a reproducible reasoning workflow. |
| Outcome: | The proposed framework improves semantic alignment and controllability on metaphor prompts. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases. |
| Approach: | They propose a bilingual, multi-task benchmark for long context understanding that extends context windows and more sophisticated memory mechanisms to improve models' long context capabilities. |
| Outcome: | The proposed model outperforms open-source models but struggles on longer contexts. |
Copied to clipboard
| Challenge: | Existing pre-training objectives do not explicitly model relational facts in text . Experimental results show that ERICA can improve typical PLMs on several language understanding tasks, including relation extraction, entity typing and question answering. |
| Approach: | They propose a contrastive learning framework ERICA to obtain a deep understanding of entities and relations in text. |
| Outcome: | The proposed framework can improve PLMs on several language understanding tasks, especially under low-resource settings. |
Copied to clipboard
| Challenge: | Code LLMs lack reproducible data pipelines and training protocols for reproducible advancements in code intelligence. |
| Approach: | They propose a top-tier code LLM that releases model weights and inference code . reproducible data pipelines, rigorous experimental ablation results and training protocols are included . |
| Outcome: | The proposed model achieves comparable performance to leading models and serves as an "open cookbook" reproducible training data, rigorous experimental ablation results, and detailed training protocols are also included in the model. |
Copied to clipboard
| Challenge: | Neural machine translation suffers from slow translation speed due to the large search space . a trade-off has to be made between translation quality and speed, argues a new study . |
| Approach: | They apply cube pruning technique to speed up dynamic programming into neural machine translation to speed it up. |
| Outcome: | The proposed method can translate faster on GPUs and CPUs with better translation quality than naive beam search. |
Copied to clipboard
| Challenge: | Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers. |
| Approach: | They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc. |
| Outcome: | The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination. |
Copied to clipboard
| Challenge: | Existing methods to enhance reasoning capabilities of language models are expensive and often lack the ability to perform complex reasoning tasks. |
| Approach: | They propose a token-level multi-model collaboration strategy to enhance reasoning capabilities in language models by selecting the optimal tokens from the next token distributions. |
| Outcome: | The proposed method is superior to existing methods and will be released soon. |
Copied to clipboard
| Challenge: | Existing approaches to detect offensive content are expensive and require massive manual effort. |
| Approach: | They propose an approach capable of utilizing the bag-level labeled data for offensive language detection by an annotation-based model. |
| Outcome: | The proposed model can detect offensive language on both bag-level and sentence level. |
Copied to clipboard
| Challenge: | Existing methods for addressing item-level user interests are lacking in cross-domain generalization . RecBase model is domain-agnostic and can be used to enhance recommender systems' effectiveness . |
| Approach: | They propose a domain-agnostic foundational model pretrained with a recommendation-oriented objective that leverages a large-scale, heterogeneous, cross-domain corpus with unified textual representations and feature mappings to enhance cross- domain generalization. |
| Outcome: | The proposed model matches or surpasses baselines in zero-shot and cross-domain recommendation tasks on eight real-world datasets. |
Copied to clipboard
| Challenge: | Existing methods that rely on limited demos and out-of-demonstration (OOD) queries fail when faced with out- of-demotion queries. |
| Approach: | They propose a query-aware prompting method that elicits the inherent generalizability of large language models by query-based demo generation. |
| Outcome: | The proposed method outperforms state-of-the-art methods in the OOD setting and two public math benchmarks. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models struggle with 3D spatial reasoning as they fail to construct structured abstractions of the 3D environment depicted in video inputs. |
| Approach: | They propose a prompting method that induces MLLMs to generate 3D representations as reasoning traces for more accurate spatial question answering. |
| Outcome: | Extensive experiments on VSI-Bench and OST-Bech show that TRACE improves over prior prompting strategies across a diverse range of MLLM backbones. |
Copied to clipboard
| Challenge: | Existing models for word-level auto-completion (WLAC) do not meet the criterion of good auto-completes. |
| Approach: | They propose a measurable criterion to address the question: what kind of words are good auto-completions? they propose an approach to enhance WLAC performance by promoting adherence to the cri-terion. |
| Outcome: | The proposed approach outperforms the top-performing system submitted to the WLAC shared tasks in WMT2022 while using significantly smaller model sizes. |
Copied to clipboard
| Challenge: | Existing methods for grammatical error correction (GEC) have been developed. |
| Approach: | They propose a method which integrates the detection labels from a Seq2Edit model to construct a template as the input. |
| Outcome: | The proposed method can perform human-in-the-loop error correction tasks. |
Copied to clipboard
| Challenge: | a number of safety concerns hinder the deployment of open-domain dialog systems, such as offensive languages and toxic behaviors, such social bias is difficult to detect. |
| Approach: | They propose a Dial-Bias Framework for analyzing social bias in conversations . they introduce a Chinese social bias dialog dataset and conduct in-depth ablation studies . |
| Outcome: | The proposed framework is the first annotated Chinese social bias dialog dataset . the proposed framework also provides a fine-grained dialog bias measurement benchmark . |
Copied to clipboard
| Challenge: | Existing debt collection agents fail to tailor strategies to debtor personas, leading to ineffective collection. |
| Approach: | They present a commercial practice on debt collection agents that organizes debtor personas into a taxonomy and constructs a persona-aware conversation dataset. |
| Outcome: | The proposed agent increases recovery rate by 3.31% and collects additional 100K RMB after two months of testing. |
Copied to clipboard
| Challenge: | Significant concerns emerge when addressing cultural sensitivity and local values. |
| Approach: | They propose a localized Large Language Model (LLM) specifically for Arabic, a language imbued with unique cultural characteristics inadequately addressed by current mainstream models. |
| Outcome: | The proposed model sets the state-of-the-art standard for open Arabic LLMs across various benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to fine-tune large language models pose privacy risks . researchers have synthesized data with strong generation capabilities closed-source LLMs to alleviate this problem . |
| Approach: | They propose to combine general LLMs with genetic algorithm to produce relevant and diverse synthetic text under differential privacy constraints. |
| Outcome: | The proposed method significantly improves the performance of the model in downstream tasks while ensuring privacy. |
Copied to clipboard
| Challenge: | Existing methods to train a parser to perform zero-shot learning are limited by the lack of training data. |
| Approach: | They propose a decomposition-based method to unify the sentence structures of questions . their method can generalize to natural questions with novel text expressions . |
| Outcome: | The proposed method improves on synthetic data and on complex web questions with novel expressions. |
Copied to clipboard
| Challenge: | Using video vision language models, inference costs are often more expensive than finetuning. |
| Approach: | They investigate the optimal allocation of inference compute across three key scaling factors in video vision language models. |
| Outcome: | The proposed model configurations are based on three key scaling factors . the results can be applied to real-world tasks and tasks with fixed inference budgets. |
Copied to clipboard
| Challenge: | Open Information Extraction (OpenIE) is a key NLP task aimed at extracting structured information from unstructured text sources. |
| Approach: | They propose to categorize OpenIE into rule-based, neural, and pre-trained large language models and discuss each within a chronological framework. |
| Outcome: | The paper categorizes OpenIE approaches into rule-based, neural, and pre-trained large language models, discussing each within a chronological framework. |
Copied to clipboard
| Challenge: | Existing quantization techniques have been categorized as 'simple' and 'highly efficient' however, their configurations vary from each other and cannot be fairly compared . |
| Approach: | They propose a plug-and-play compression toolkit to explore the impact of quantization. |
| Outcome: | The proposed toolkit explores the impact of quantization on large language models. |
Copied to clipboard
| Challenge: | Semantic phrases (SP) are lexical combinations whose meanings or usages may not be fully derived from their individual components. |
| Approach: | They propose to consolidate existing multiword expression resources into a unified testbed to assess language models in semantic phrase processing tasks. |
| Outcome: | The evaluation suite covers idiomatic expressions, noun compounds, and verbal constructions. |
Copied to clipboard
| Challenge: | EmoDS can express emotions in both ways, but it is difficult to scale to large datasets. |
| Approach: | They propose an emotional dialog system that can express emotions in both ways . they use strong emotional words and neutral words to increase the intensity of emotions . |
| Outcome: | The proposed system performs better than baselines in BLEU, diversity and quality of emotional expression. |
Copied to clipboard
| Challenge: | Current multimodal benchmarks focus on facts within individual images, but neglect associative relations among multiple images. |
| Approach: | They propose a multi-image relational association task and a MMRA benchmark to evaluate LVLMs. |
| Outcome: | The proposed benchmarks show that entity-level multi-image perception tasks pose greater challenges than image-level tasks. |
Copied to clipboard
| Challenge: | Existing models learn user and item embeddings and generate reasons based on these embedds. |
| Approach: | They propose a concept-based explanation framework that leverages macro concepts to bridge the gap between the user/item embeddings and the recommendation reasons. |
| Outcome: | Extensive experiments on three datasets prove the proposed model is superior to existing models. |
Copied to clipboard
| Challenge: | Existing methods for expressive text-to-speech only implicitly learn prosody with masked token reconstruction tasks. |
| Approach: | They propose a cross-modal contrastive pre-training framework that learns from prosody variance of the same text token under different contexts. |
| Outcome: | The proposed framework can learn from prosody variance of a text token under different contexts. |
Copied to clipboard
| Challenge: | Existing models for NLP evaluations lack the ability to generate informative critiques in pointwise grading and pairwise comparison especially without references. |
| Approach: | They propose a method which can acquire pointwise grading critiques with pseudo references and revise these critiques via multi-path prompting to obtain informative evaluation data in different tasks and settings. |
| Outcome: | The proposed method outperforms all open-source models and even GPT-4 in system-level correlations of pointwise grading. |
Copied to clipboard
| Challenge: | Existing methods to extract product attribute value require multiple extractions to obtain all corresponding values. |
| Approach: | They propose an Efficient product Attribute Value Extraction approach using lightweight sparse-layer interaction. |
| Outcome: | The proposed method achieves significant efficiency gains with neutral or marginal loss in performance when the context is long and number of attributes is large. |
Copied to clipboard
| Challenge: | Existing knowledge graph embedding methods are limited in their expressiveness and lack structural information in the embeddable space. |
| Approach: | They propose to use a relation-aware network to learn query embedding . they first explore the Inception network to further increase interactions between head and relation embedders . |
| Outcome: | The proposed network improves performance on WN18RR and FB15k-237 datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive capabilities when leveraging in-context learning. |
| Approach: | They propose a method that discretizes uninformative tokens using a self-supervised pre-training technique. |
| Outcome: | The proposed method achieves state-of-the-art performance across classification tasks while requiring only 0.8% decrease in performance. |
Copied to clipboard
| Challenge: | Existing work integrates reinforcement learning with compiler feedback to enhance code generation quality but the long code generated by LLMs makes RL exploration ineffective. |
| Approach: | They propose a framework that integrates reinforcement learning and compiler feedback to enhance code generation quality. |
| Outcome: | The proposed framework outperforms state-of-the-art approaches in corresponding benchmarks and integrates reinforcement learning with compiler feedback to improve code generation quality. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have great potential to facilitate explainable diagnosis, but their effectiveness is often constrained by insufficient diagnostic expertise. |
| Approach: | They propose a unified LLM-based framework for faithful and explainable diagnosis that builds a high-quality diagnostic knowledge base through a record-driven explanation learning paradigm. |
| Outcome: | The proposed framework outperforms baselines on the DiReCT and JAMA benchmarks and improves the explanation completeness metric from 64.5% to 76.9% over the best existing methods. |
Copied to clipboard
| Challenge: | Existing evaluation methods rely on external evaluators, focusing on training and prompting strategies, but model-aware glass-box features are overlooked. |
| Approach: | They propose to use model-aware glass-box features to evaluate an LLM's output. |
| Outcome: | The proposed model-aware features are reliable quality indicators for self-evaluation on public benchmarks. |
Copied to clipboard
| Challenge: | Conventional wisdom in pruning Transformer-based language models is that it reduces model expressiveness, but new research shows pruning increases risk of overfitting when performed at the fine-tuning phase. |
| Approach: | They propose to reduce pruning risk under pretrain-and-finetune paradigm . they propose to use knowledge distillation to improve pruning performance . |
| Outcome: | The proposed method outperforms the leading competitors on the GLUE benchmark. |
Copied to clipboard
| Challenge: | Entity linking is a fundamental task in Natural Language Processing (NLP), connecting mentions within unstructured contexts to their corresponding entities in a Knowledge Base (KB). |
| Approach: | They propose a dual-encoder framework that can efficiently match mentions to two-encoding frameworks by a global-view. |
| Outcome: | The proposed framework achieves state-of-the-art on several entity linking benchmarks. |
Copied to clipboard
| Challenge: | Current frontier models sometimes generate false outputs or answers that are not substantiated by evidence. |
| Approach: | They propose Chinese SimpleQA, a Chinese benchmark to evaluate LLMs' factuality . they focus on Chinese language over 6 major topics with 99 diverse subtopics . |
| Outcome: | The Chinese SimpleQA benchmark evaluates the factuality ability of LLMs . the questions and answers are short and easy-to-evaluate . |
Copied to clipboard
| Challenge: | Existing MLLMs still struggle to achieve precise grounding in multi-image scenarios. |
| Approach: | They propose a Chain-of-Thought framework that integrates single-image grounding with multi-image comprehension to address this challenge. |
| Outcome: | The proposed model outperforms existing models in multi-image grounding tasks by 24.94% and surpasses larger 70B models. |
Copied to clipboard
| Challenge: | Existing benchmarks that assess Language Models (LMs) as Language Agents (LAs) for tool use focus on stateless, single-turn interactions or partial evaluations, overlooking the inherent stateful nature of interactions in multi-turn applications. |
| Approach: | They propose a multi-turn dialogue dataset with stateful tool interactions considering the whole life cycle of tool use across six key tasks in three stages . they also build VirtualMobile – an embodied virtual mobile evaluation environment to simulate API calls and assess the robustness of the created APIs. |
| Outcome: | The proposed dataset evaluates 13 open- and closed-source LLMs and provides detailed analysis at each stage. |
Copied to clipboard
| Challenge: | Existing applications of natural language processing (NLP) focus on patient-centered services, but the potential of NLP to benefit inexperienced doctors remains unexplored. |
| Approach: | They propose a human-AI cooperative framework to assist medical learners in practicing communication skills during patient consultations. |
| Outcome: | The proposed framework enables medical learners to practice communication skills during patient consultations while a coach agent provides immediate, structured feedback. |
Copied to clipboard
| Challenge: | Recent surge in jailbreaking attacks has revealed significant vulnerabilities in Large Language Models (LLMs) however, limited research into the underlying mechanisms that make LLMs vulnerable to such attacks has been conducted. |
| Approach: | They propose that LLMs' self-safeguarding capability is linked to specific activity patterns within their representation space. |
| Outcome: | The proposed models can be detected with a few pairs of contrastive queries, and the robustness can be manipulated by weakening or strengthening these patterns. |
Copied to clipboard
| Challenge: | Experimental results show that our method outperforms full-model tuning in few-shot abstractive summarization tasks. |
| Approach: | They propose a soft prompts architecture with prompt pre-training and prompt fine-tuning paradigm to support few-shot abstractive summarization. |
| Outcome: | The proposed model outperforms Prompt Tuning and Profix-Tuning on CNN/DailyMail and XSum datasets and outperfies Profix Tuning by a large margin. |
Copied to clipboard
| Challenge: | Existing summarization methods ignore the importance of summary structure, resulting in summaries that emphasize the most prominent information while omitting essential details from other sections. |
| Approach: | They propose a method that uses automatically extracted summary points to generate summaries. |
| Outcome: | The proposed methods improve quality and BERTScore of summaries and broaden the types of documents that can be effectively summarized. |
Copied to clipboard
| Challenge: | Existing word embedding methods learn semantic information at word level while neglecting meaningful inner structures of words like morphemes. |
| Approach: | They propose to use latent meanings of morphological compositions of words to train word embeddings. |
| Outcome: | The proposed models outperform baseline models on word similarity, syntactic analogy and text classification tasks. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have greatly expanded the scope of legal AI. |
| Approach: | They propose a method that generates questionnaires to help users refine queries . they leverage an iterative training process that collects valuable questionnaires . |
| Outcome: | The proposed method improves the completeness of queries and ensures the performance of domain-specific models in downstream legal tasks. |
Copied to clipboard
| Challenge: | Experimental results show that consistency regularization improves cross-lingual fine-tuning . pre-trained cross-linguistic models can transfer task-specific supervision from one language to the other . |
| Approach: | They propose to improve cross-lingual fine-tuning with consistency regularization . they use example consistency regularized to penalize prediction sensitivity to four types of data augmentations . |
| Outcome: | The proposed method improves cross-lingual fine-tuning across tasks . it can be generalized to other target languages without additional training . |
Copied to clipboard
| Challenge: | Recent methods based on pre-trained language models have shown strong supervised performance on commonsense reasoning. |
| Approach: | They propose to use a common framework to solve commonsense reasoning tasks using a dataset from NLI. |
| Outcome: | The proposed method achieves state-of-the-art unsupervised performance on two commonsense reasoning tasks. |
Copied to clipboard
| Challenge: | DA-Code is a code generation benchmark designed to assess LLMs on agent-based data science tasks. |
| Approach: | They propose a code generation benchmark specifically designed for LLMs on agent-based data science tasks. |
| Outcome: | The benchmark performs better than existing frameworks, but lacks accuracy . it is based on real-world data, and includes examples that cover a wide range of tasks . |
Copied to clipboard
| Challenge: | Pretrained sequence-to-sequence (seq2sequ) models have been widely used to solve extractive tasks, where parts of the input are extracted to form the desired output. |
| Approach: | They propose a simple fix to tokenization inconsistency that damages extractive nature of generative models by causing performance drop and hallucination. |
| Outcome: | The proposed model performs better in both in-domain and out-of-domain datasets with a notable average of +1.7 F1 gain when a BART model is trained on SQuAD and evaluated on 8 QA datasets. |
Copied to clipboard
| Challenge: | Existing methods to infer the missing links between entities are limited to the transductive setting . Query Adaptive Anchor Representation (QAAR) model is based on entity-independent features . |
| Approach: | They propose a query adaptive anchor representation model which extracts one opening subgraph and performs reasoning by one time for all candidate triples. |
| Outcome: | The proposed model outperforms state-of-the-art models in relation prediction task. |
Copied to clipboard
| Challenge: | Existing methods to update a multilingual model with new language pairs are expensive and time-consuming. |
| Approach: | They propose an entropy-based vocabulary substitution method that walks through new language pairs for incremental learning while remaining the size of the original vocabulary. |
| Outcome: | The proposed method achieves better performance and saves excess overhead in a multilingual machine translation task. |
Copied to clipboard
| Challenge: | Existing models for large language models lack the ability to calibrate their outputs towards human preference. |
| Approach: | They propose a multi-stage, gradient-free approach to calibrate an LLM-based evaluator toward human preference. |
| Outcome: | The proposed approach improves correlation with expert evaluation on multiple text quality evaluation datasets. |
Copied to clipboard
| Challenge: | Entity typing fails to assign an entity to the types beyond the predefined type set. |
| Approach: | They propose a generative entity typing paradigm that assigns types to entities . traditional classification-based approaches fail to assign entities to the types beyond the predefined set . they employ curriculum learning to train the model on heterogeneous data . |
| Outcome: | The proposed model outperforms the state-of-the-art model on heterogeneous training data. |
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a promising solution for large language models, but their deployment remains computationally expensive. |
| Approach: | They propose a framework which iteratively balances training data within a fixed computational budget and enables the transfer of knowledge from expensive teacher LLMs to smaller student models. |
| Outcome: | The proposed framework achieves state-of-the-art performance across diverse long-tailed datasets, enhancing both the efficiency and efficacy of the distilled models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have impressive capabilities but need for task-specific prompt engineering can hinder their generalization. |
| Approach: | They propose a lightweight and versatile retriever that automatically retrieves prompts for a given zero-shot task input. |
| Outcome: | The proposed model is universally applicable across tasks and models . it mitigates hallucination problem in chatGPT, and it improves even the strongest LLMs. |
Copied to clipboard
| Challenge: | Social media is an easy-to-access platform providing timely updates about societal trends and events. |
| Approach: | They propose a framework to extract epidemic-related events from social media posts to provide early warnings. |
| Outcome: | The proposed framework can detect epidemic events for three unseen epidemics of Monkeypox, Zika, and Dengue while existing models fail miserably. |
Copied to clipboard
| Challenge: | Large language model (LLM) agents have demonstrated strong problem-solving competence across domains like research and coding. |
| Approach: | They propose to use a tool repository to analyze the ability of large language model agents to solve complex problems. |
| Outcome: | The proposed model outperforms open-source and closed-source models in task completion rate and efficiency. |
Copied to clipboard
| Challenge: | Existing pre-trained models neglect to consider linguistic knowledge of texts . existing models neglect linguistic information, which is important for sentiment analysis . |
| Approach: | They propose a model that introduces word-level linguistic knowledge into pre-trained models to enhance sentiment analysis by querying SentiWordNet to acquire sentiment polarity. |
| Outcome: | The proposed model obtains state-of-the-art performance on a variety of sentiment analysis tasks. |
Copied to clipboard
| Challenge: | a recent study shows that agent research practices are far from standard, rigorous . lack of a standard evaluation protocol makes previous works not reproducible, authors say . |
| Approach: | They conduct an empirical study on the GAIA benchmark to investigate agent design choices . they find that lack of a standard evaluation protocol makes previous works not reproducible . |
| Outcome: | The proposed framework achieves state-of-the-art performance among open-source projects. |
Copied to clipboard
| Challenge: | Existing knowledge graph embedding models suffer from Z-paradox, a deficiency in expressiveness . Embedding-based models map each entity and relation into a vector or matrix . |
| Approach: | They propose a new knowledge graph embedding model that does not suffer from Z-paradox while preserves strong expressiveness to model various relation patterns with theoretical justification. |
| Outcome: | The proposed model outperforms existing models on link prediction tasks while maintaining strong expressiveness. |
Copied to clipboard
| Challenge: | Existing models for speech-to-speech translation suffer from distinct degradation in noisy environments and fail to translate visual speech. |
| Approach: | They propose a text-based audio-visual speech-to-speech translation model that integrates visual information with audio-only data to improve system robustness. |
| Outcome: | The proposed model outperforms models trained on audio-only corpus in two languages . it also improves with low-resource audio-visual data, compared with baselines . |
Copied to clipboard
| Challenge: | Recent studies have introduced legal theories into LLM workflows to improve their understanding of legal texts and reasoning accuracy. |
| Approach: | They evaluate an expert-annotated four-element knowledge base covering 155 criminal charges. |
| Outcome: | The proposed model can be used to analyze criminal charges and retrieve them in legal cases. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on simple synthesized queries that do not reflect real-world complexity, thereby offering limited perspectives in evaluating tool utilization. |
| Approach: | They propose a benchmark to evaluate LLMs’ ability in tool utilization within real-world scenarios. |
| Outcome: | The proposed benchmark improves LLMs’ ability in tool utilization within real-world scenarios and eliminates the restriction of pre-defined toolset. |
Copied to clipboard
| Challenge: | Prompt tuning is parameter-efficient but lags behind other state-of-the-art methods. |
| Approach: | They propose a parameter-efficient tuning method that only optimizes a soft prompt to adapt PTMs to downstream tasks. |
| Outcome: | The proposed method is parameter-efficient but lags behind other state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing automatic metrics are observed to correlate poorly with human evaluation. |
| Approach: | They propose to use OpenMEVA to evaluate open-ended story generation metrics. |
| Outcome: | The proposed test suite assesses the capabilities of open-ended story generation metrics on annotated stories and auto-constructed test examples. |
Copied to clipboard
| Challenge: | Existing methods for ERC lack human-like emotion reasoning and discrimination between similar emotions. |
| Approach: | They propose a multi-dimension curriculum with long CoT fine-tuning to clone human-like emotion reasoning for conversational emotion recognition. |
| Outcome: | The proposed model outperforms existing methods on three widely used datasets and shows that it is more intuitive and more accurate. |
Copied to clipboard
| Challenge: | Existing vision-and-language generation models cannot utilize pair-wise images and text through bi-directional generation due to the limitations of the model structure and pre-training objectives. |
| Approach: | They propose a framework which unifies vision-and-language generation as sequence generation problems. |
| Outcome: | The proposed framework achieves better performance than variants trained with uni-directional generation objectives or the variant without the commitment loss on image captioning and text-to-image generation datasets. |
Copied to clipboard
| Challenge: | Existing datasets for Chinese instruction tuning are not well-aligned with Chinese users’ interaction patterns. |
| Approach: | They propose to use Chinese instruction tuning datasets to improve instruction fine-tuning for Chinese users. |
| Outcome: | The proposed dataset shows that Chinese models achieve competitive performance in diverse benchmarks. |
Copied to clipboard
| Challenge: | Creating job requirements is a crucial step in the recruiting process, but it is difficult to specify the level of education, experience, relevant skills per the job description. |
| Approach: | They propose a conditional text generation task to generate job requirements based on job descriptions . they use a hierarchical decoder to label the job description with multiple skills . a skill knowledge graph is constructed to capture the global prior knowledge about skills based upon the model . |
| Outcome: | The proposed method is evaluated on real-world job posting data. |
Copied to clipboard
| Challenge: | Existing knowledge graphs are incomplete in tracking complex semantic relations of the target-oriented dialogue. |
| Approach: | They combine methods of knowledge retrieval and relationship prediction to construct a context-related dynamic KG and a metric to evaluate the tracked path automatically. |
| Outcome: | The proposed method can control the agent more logically and smoothly toward the complex target. |
Copied to clipboard
| Challenge: | Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been explored. |
| Approach: | They propose to use a benchmark to evaluate large language models' financial domain knowledge and practical abilities. |
| Outcome: | The proposed benchmark evaluates large language models' financial domain knowledge and practical abilities. |
Copied to clipboard
| Challenge: | Existing knowledge graphs focus on connecting intentions but lacks the ability to model the relationships between different intentions. |
| Approach: | They propose a framework to automatically generate an intention knowledge graph, capturing connections between user intentions. |
| Outcome: | The proposed model outperforms state-of-the-art methods and shows its utility. |
Copied to clipboard
| Challenge: | Teaching large language models to generate text with citations to evidence sources requires high-quality attribution data, which is costly and labor-intensive. |
| Approach: | They propose a framework for iteratively improving the attribution capability of large language models (LLMs) by attributing output to verifiable sources. |
| Outcome: | Experiments on three open-domain question-answering datasets show that START improves in aggregating information across multiple sources. |
Copied to clipboard
| Challenge: | Recent neural machine translation models have improved translation quality but they also introduce small perturbations like misspelling and paraphrasing. |
| Approach: | They propose a method that generates multi-granularity subword representations with reversible operations and deterministic segmentations. |
| Outcome: | The proposed method outperforms strong baselines on several translation tasks with a clear margin and exhibits good robustness in noisy, low-resource, and cross-domain datasets. |
Copied to clipboard
| Challenge: | Recent years have seen remarkable success in the use of deep neural networks on Chinese word segmentation (CWS) however, the performance of CWS systems has gradually reached a plateau with the rapid development of deep networks. |
| Approach: | They propose a fine-grained evaluation for existing Chinese word segmentation systems that allows us to diagnose the strengths and weaknesses of existing models. |
| Outcome: | The proposed model can diagnose strengths and weaknesses of existing models and alleviate negative transfer problem when doing multi-criteria learning. |
Copied to clipboard
| Challenge: | Existing studies on text style transfer neglect long style transfer at the discourse level. |
| Approach: | They propose a model that transfers text style into target styles with learnable style embeddings . they use a mask-and-fill framework to explicitly fuse style-specific keywords into generation . |
| Outcome: | The proposed model outperforms baselines in style transfer and content preservation. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to perceive the visual world with a deep concept structure cognition. |
| Approach: | They propose a concept-level benchmark to assess MLLMs’ hierarchical concept understanding and reasoning abilities. |
| Outcome: | The proposed model outperforms state-of-the-art models in concept structure reasoning evaluation. |
Copied to clipboard
| Challenge: | Existing preference-based methods for medical large vision-Language Models face limitations in medical settings . existing methods are limited by overfitting to superficial cues and pseudo convergence of the preference signal. |
| Approach: | They propose a framework that enables evidence-aware and adaptive preference learning for Med-LVLMs. |
| Outcome: | The proposed framework improves evidence-aware and adaptive preference learning for Med-LVLMs. |
Copied to clipboard
| Challenge: | Existing benchmarks rely on outcome-driven metrics such as profitability and look-ahead bias. |
| Approach: | They propose a diagnostic benchmark for instruction-grounded financial code generation under strict semantic and temporal constraints. |
| Outcome: | The proposed benchmarks show that the models fail under causal, structural, or functional constraints. |
Copied to clipboard
| Challenge: | Existing research on text-based mental health counseling is limited due to the lack of relevant corpora in Chinese language. |
| Approach: | They propose a Chinese dataset of psychological health support in the form of question and answer pair that is crawled from a mental health service platform and contains 22K questions and 56K long and wellstructured answers. |
| Outcome: | The proposed dataset contains 22K questions and 56K long and wellstructured answers. |
Copied to clipboard
| Challenge: | Synthesizing tool-use data through real-world simulations is effective for enhancing large language models (LLMs) however, training gains decay as synthetic data increases, and the model struggles to benefit from more synthetic data. |
| Approach: | They propose an iterative reinforced fine-tuning strategy to improve LLMs with external tools to augment their capabilities. |
| Outcome: | The proposed method achieves 13.11% better performance than the same-size base model and outperforms larger open-source and closed-source models. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed method is significantly better than the baselines including the one solely based on BERT. |
| Approach: | They propose a neural architecture which uses a network for error detection and a system for error correction based on BERT, with the latter connected to the other using what they call soft-masking technique. |
| Outcome: | The proposed method performs better than baselines including the one solely based on BERT, and is general and may be employed in other language detection-correction problems. |
Copied to clipboard
| Challenge: | Existing methods to retrieve evidences from corpus are difficult due to table-text discrepancy and data sparsity problem. |
| Approach: | They propose an optimized OpenQA Table-Text Retriever to retrieve tabular and textual evidences from tabular resources. |
| Outcome: | The proposed OpenQA Table-Text Retriever significantly outperforms existing methods on QA tasks. |
Copied to clipboard
| Challenge: | Pre-trained language models have improved dependency parsing accuracy in resource-rich languages . however, the accuracy drops sharply when the model is transferred to low-resource language . |
| Approach: | They propose a representation alignment and adversarial model to filter out useful knowledge from rich-resource language and ignore useless ones. |
| Outcome: | The proposed model outperforms baseline models on the benchmark datasets by 1.37 LAS and 1.34 UAS. |
Copied to clipboard
| Challenge: | Medical Multi-Modal Large Language Models (Med-MLLMs) are a promising new form of artificial general intelligence due to their ability to tackle complex tasks. |
| Approach: | They propose a new benchmark that comprehensively assesses medical multi-modal large language models in terms of distinct medical specialties and different diagnostic capacities. |
| Outcome: | The proposed model covers 15 medical specialties and different diagnostic capacities, and excludes overlap with existing VQA dataset. |
Copied to clipboard
| Challenge: | Existing evaluations emphasize final accuracy or coarse token counts, and lack automated tools to separate essential logic from structural redundancy. |
| Approach: | They propose a graph-driven framework that quantifies reasoning efficiency by converting free-form CoTs into directed dependency graphs and extracting the Shortest Effective Path needed to reach a correct solution. |
| Outcome: | Evaluating 21 LRMs, the proposed framework quantifies reasoning efficiency by converting free-form CoTs into directed dependency graphs and extracting the Shortest Effective Path (SEP) needed to reach a correct solution. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have expanded to more complex repository-level tasks. |
| Approach: | They propose a first approach to leveraging visual data to enhance the issue-resolving capabilities of Large Language Models (LLMs) they demonstrate the effectiveness of CodeV and provide valuable insights into leveraging visualization to resolve GitHub issues. |
| Outcome: | The proposed approach improves the issue-resolving capabilities of Large Language Models (LLMs) by using visual data. |
Copied to clipboard
| Challenge: | Existing methods for Sequence Labeling require high-quality annotations, but imperfect annotations are relatively easy to obtain from crowdsourcing (noisy labels) Existing approaches to learn a model without knowing the underlying ground truth label sequences in the target domain are expensive and time-consuming. |
| Approach: | They propose a framework Consensus Network that can be trained on annotations from multiple sources. |
| Outcome: | The proposed framework improves on learning with crowd annotations and unsupervised cross-domain model adaptation in two practical settings. |
Copied to clipboard
| Challenge: | Existing methods to assess the correctness of RAG models fail to capture the model’s internal state during answer generation. |
| Approach: | They propose a method to predict the correctness of RAG models by modeling the model’s uncertainty on quantified perturbations of input. |
| Outcome: | Extensive experiments across multiple large language models show that the proposed approach quantifies RAG robustness by aligning predictions with ground truth with a MSE 0.002 while offering flexibility for diverse qualitative metrics. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated impressive reasoning capabilities, yet there is ongoing debate about their capabilities and the potential data contamination problem. |
| Approach: | They propose to evaluate the reasoning capabilities of large language models in solving recent competition-level programming problems in Codeforces. |
| Outcome: | The proposed model has experienced a cliff-like decline in problems after September 2021, which shows the potential data contamination and the challenges for any existing LLM to solve unseen complex reasoning problems. |
Copied to clipboard
| Challenge: | Existing methods for cross-lingual entity alignment ignore useful pre-aligned links between two KGs. |
| Approach: | They propose a novel method that jointly learns embeddings in different KGs by propagating cross-KG information through pre-aligned seed alignments. |
| Outcome: | The proposed method achieves remarkable performance gains on three benchmark cross-lingual entity alignment datasets. |
Copied to clipboard
| Challenge: | Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs. |
| Approach: | They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. |
| Outcome: | The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities. |
Copied to clipboard
| Challenge: | Text-to-speech synthesis (TTS) has seen rapid progress in recent years, but still suffers from latencies. |
| Approach: | They propose a neural incremental TTS approach that synthesizes speech in an online fashion, playing a segment of audio while generating the next. |
| Outcome: | Experiments on English and Chinese TTS show that the proposed approach achieves similar speech naturalness compared to full sentence TTS, but with a constant (1-2 words) latency. |
Copied to clipboard
| Challenge: | Existing approaches to enhance neural machine translation (NMT) by using a TM have been reported to be effective. |
| Approach: | They propose a translation memory augmented neural machine translation model that is good at fitting data but more sensitive to fluctuations in training data. |
| Outcome: | The proposed model achieves consistent gains over conventional and existing models under two variance-preferable scenarios as well as the high resource scenario. |
Copied to clipboard
| Challenge: | a new benchmark for multilingual foundation models is being developed . brittleness of foundation models in the dimensions of semantics and multilinguality is a key limitation . |
| Approach: | They propose a benchmark for multilingual foundation models, SeaEval . they examine how well these models comprehend cultural practices, nuances, and values . |
| Outcome: | The proposed model can be used to evaluate multilingual and multicultural scenarios. |
Copied to clipboard
| Challenge: | a new task for context-situated pun generation uses a given context to generate puns . human evaluation shows that 69% of top retrieved pun words can be used to generate context-based puns. |
| Approach: | They propose a task where puns are generated based on contextual keywords and pun words. |
| Outcome: | The proposed system generates successful puns 31% of the time given a plausible tuple of context words and pun pairs. |
Copied to clipboard
| Challenge: | Existing efforts to automate conversational moderation have focused on banning harmful comments or deleting them, but such efforts can inadvertently push users towards echo chambers that exacerbate polarization. |
| Approach: | They propose a framework to assess models’ moderation capabilities independently of human intervention and propose 'conversational moderation' they propose to use language models as conversational moderators to provide specific feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
| Outcome: | The proposed framework assesses models’ moderation capabilities independently of human intervention and shows that appropriately prompted models provide specific and fair feedback on toxic behavior but struggle to influence users to increase their levels of respect and cooperation. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) often generate overconfident yet factually incorrect hallucinations. |
| Approach: | They propose a black-box-based framework that captures stubborn hallucinations by integrating internal geometric dynamics with output probability distributions. |
| Outcome: | The proposed framework outperforms white-box methods and reduces computational overhead by over 90%. |
Copied to clipboard
| Challenge: | Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs). |
| Approach: | They propose a hybrid-policy optimization approach that synergizes internal exploitation with external data to achieve stronger reasoning capabilities. |
| Outcome: | The proposed approach achieves state-of-the-art performance on six math reasoning benchmarks and superior performance on out-of distribution reasoning tasks. |
Copied to clipboard
| Challenge: | Existing studies on back translation (BT) focus on beam search or random sampling . a new method to generate synthetic data with a backward model is proposed to improve BT performance. |
| Approach: | They propose a method to generate synthetic data to trade off quality and importance factors . back translation (BT) is one of the most significant technologies in NMT research fields . |
| Outcome: | The proposed method outperforms the baseline methods on WMT14 DE-EN, EN-DE, and RU-EN benchmark tasks. |
Copied to clipboard
| Challenge: | Existing studies focus on overcoming catastrophic forgetting on original language pairs while lacking encouragement to learn new knowledge from incremental learning. |
| Approach: | They propose a knowledge transfer method that can adapt original MNMT models to diverse incremental language pairs by flexibly introducing knowledge from external models into original models, which encourages the models to learn new language pairs. |
| Outcome: | The proposed method outperforms baselines on multiple languages while maintaining performance on original language pairs. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases. |
| Approach: | They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process. |
| Outcome: | The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations. |
Copied to clipboard
| Challenge: | Existing approaches to textual robustness evaluation focus on slightly modifying the input data, which maintains the original meaning and results in a different prediction. |
| Approach: | They propose a multilingual robustness evaluation toolkit for NLP that integrates universal text transformations, task-specific transformations and adversarial attack. |
| Outcome: | The toolkit includes universal text transformation, task-specific transformation, adversarial attack, subpopulation, and their combinations to provide comprehensive robustness analyses. |
Copied to clipboard
| Challenge: | Existing methods to reward LLMs' outputs are not effective in mathematical reasoning scenarios and may lead to a decline in performance. |
| Approach: | They propose a process-based self-rewarding pipeline that integrates long-thought reasoning, step-wise LLM-as-a-Judge, and step- wise preference optimization within the existing paradigm. |
| Outcome: | The proposed model improves the performance of Large Language Models on multiple mathematical reasoning benchmarks and shows that it can surpass human capabilities. |
Copied to clipboard
| Challenge: | Recent multilingual pre-trained models have been demonstrated effective in many cross-lingual tasks. |
| Approach: | They propose a framework that leverages code-switched data with multi-view learning to fine-tune XLM-R. |
| Outcome: | The proposed model achieves state-of-the-art on zero-shot cross-lingual sentiment classification and dialogue state tracking tasks. |
Copied to clipboard
| Challenge: | Label-specific topics are widely used for supporting personality psychology, aspectlevel sentiment analysis, and crossdomain sentiment classification. |
| Approach: | They propose a supervised topic model based on the Siamese network which trades off label-specific word distributions with document-specific label distributions in a uniform framework. |
| Outcome: | The proposed model can trade off label-specific word distributions with document-specific label distributions in a uniform framework. |
Copied to clipboard
| Challenge: | a good translation should implicitly mirror user traits rather than translate the original content semantically. |
| Approach: | They propose a framework that captures user traits from historical inputs . they propose 'user-driven' NMT to model user behavior under a zero-shot learning fashion . |
| Outcome: | The proposed framework can capture user traits from historical inputs under zero-shot learning fashion. |
Copied to clipboard
| Challenge: | Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. |
| Approach: | They propose to conduct efficient text-video Retrieval with a salient-and-correlated AdaPter . they propose a low-rank modulation module to refine per-image features from frozen CLIP backbone . |
| Outcome: | Experiments on four TVR datasets show that the proposed method performs better than other methods. |
Copied to clipboard
| Challenge: | Existing generation models struggle to maintain a coherent event sequence throughout the generated text. |
| Approach: | They propose a long text generation model which can represent prefix sentences at sentence level and discourse level in the decoding process. |
| Outcome: | The proposed model can generate more coherent texts than state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing approaches often fail to leverage the linguistic intelligence of Large Language Models (LLMs) Existing models lack the ability to follow text instructions for controllable Text-to-Speech (TTS). |
| Approach: | They propose a framework where an LLM acts as a conductor, understanding user instructions and generating a textual plan - explicit vocal features. |
| Outcome: | The proposed model outperforms open- and closed-source models in speech synthesis and achieves zero-shot cross-lingual generalization. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have experienced rapid development in recent years, but there is a notable lack of effective and specialized multimodal evaluation datasets in the financial domain. |
| Approach: | They introduce FinMME, a multimodal large language model with 11,000 financial research samples and 20 annotators. |
| Outcome: | The proposed model performs better than state-of-the-art models, highlighting its challenging nature. |
Copied to clipboard
| Challenge: | Current document image parsing solutions rely on specialized models or generate content autoregressively. |
| Approach: | They propose a multimodal document image parsing model that integrates specialized models with autogeneous content generation. |
| Outcome: | The proposed model achieves state-of-the-art performance across diverse page-level and element-level settings while ensuring superior efficiency. |
Copied to clipboard
| Challenge: | Distant supervision is an important paradigm for automatically extracting relations . but the examples collected can be noisy and pose significant challenge for labeling . |
| Approach: | They propose a method to predict whether two entities participate in a relation at a given time spot. |
| Outcome: | The proposed model performs better in WIKI-TIME and NYT-10 datasets compared with the best existing models . the proposed model is based on a dataset with a valid period of a certain relation of two entities in the knowledge base . |
Copied to clipboard
| Challenge: | a new adaptive merging method is proposed to improve fine-tuning performance . traditional methods often encounter task interference when merging full fine-uning models . |
| Approach: | They propose an adaptive merging method that directly measures model parameters using the Frobenius norm . |
| Outcome: | The proposed method outperforms baseline methods in various fine-tuning scenarios. |
Copied to clipboard
| Challenge: | Semantic compositionality (SC) is defined as the phenomenon that the meaning of a complex linguistic unit can be composed of the meanings of its constituents. |
| Approach: | They propose to incorporate sememes into SC models and employ them in learning multiword expressions. |
| Outcome: | The proposed models achieve significant performance boost compared to baseline methods without sememe knowledge. |
Copied to clipboard
| Challenge: | Large language models have demonstrated considerable capabilities across various tasks . however, they often fall short of the performance achieved by domain-specific state-of-the-art models . |
| Approach: | They propose a tuning-free method to augment domain-specific abilities of Large language models . they leverage insights from the response preference of expert models to augment LLMs . |
| Outcome: | The proposed method outperforms the expert model on 4 ScienceWorld tasks. |
Copied to clipboard
| Challenge: | PLATO-2 is a high-quality open-domain chatbot that can generate one-to-many mappings and improve response quality. |
| Approach: | They propose a curriculum learning process to build a high-quality open-domain chatbot . they use a coarse-grained generation model and latent variables to train a generative model . |
| Outcome: | The proposed model improves on Chinese and English data and can generate diverse responses and select the best response. |
Copied to clipboard
| Challenge: | Existing language models are pre-trained and distilled on general corpus like Wikipedia, which has gaps with the news domain and may be suboptimal for news intelligence. |
| Approach: | They propose a method to distill existing language models on Wikipedia to enable efficient news intelligence. |
| Outcome: | The proposed model can be used to build and test a news intelligence application on Wikipedia and Wikipedia. |
Copied to clipboard
| Challenge: | Existing federated learning frameworks require substantial data and computational resources to develop large language models. |
| Approach: | They propose a method that distributes a quantized version of the model’s parameters during training and combine it with a popular fine-tuning method to significantly reduce communication costs. |
| Outcome: | The proposed method enables accurate estimations for parameter updates while preventing clients from accessing a model whose performance is comparable to the centrally hosted one. |
Copied to clipboard
| Challenge: | Existing work in multilingual pretraining relies on the shared vocabulary and bilingual contexts to encourage the correlation across languages. |
| Approach: | They propose to plug a cross-attention module into a Transformer encoder to explicitly build the interdependence between languages. |
| Outcome: | The proposed model outperforms existing models on XTREME and English-to-French translation datasets. |
Copied to clipboard
| Challenge: | Existing studies measure the superiority of DA methods in terms of their performance on a specific test set, but some do not exhibit consistent improvements across translation tasks. |
| Approach: | They propose to evaluate DA methods from two perspectives to determine their generalization ability . they find that DA method's test performance does not exhibit consistent improvements across translation tasks . |
| Outcome: | The proposed methods do not exhibit consistent improvements across translation tasks. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) improves training efficiency by updating only a small portion of the weights in Large Language Models. |
| Approach: | They propose a rotation-aware scheme to fine-tune rotated outlier-free LLMs for effective weight-activation quantization. |
| Outcome: | The proposed method improves low-bit LoRA convergence and post-training quantization robustness. |
Copied to clipboard
| Challenge: | Existing RP-LLMs employ only a single role with numerous dialogues, but Crab enables dynamic configuration of desired roles, thereby enhancing related flexibility and adaptability. |
| Approach: | They propose a Configurable Role-Playing LLM with Assessing Benchmark that combines a Role dataset curation, persona-emodying Llm construction, and comprehensive benchmark creation for RP dialogue generation. |
| Outcome: | The proposed model outperforms existing LLMs in performing fine-grained evaluations of RP while keeping dialogue per role minimal. |
Copied to clipboard
| Challenge: | Existing models for speech-driven SQL parsing are based on a cascaded approach, resulting in data scarcity and inconsistent performance. |
| Approach: | They propose a direct generalizable speech-to-SQL parsing model which avoids error compounding across cascaded systems. |
| Outcome: | The proposed model avoids error compounding and achieves state-of-the-art results by 4.7% improvement over baseline. |
Copied to clipboard
| Challenge: | Recent text-to-image models require multiple passes of prompt engineering by humans to produce satisfactory results for real-world applications. |
| Approach: | They propose a deep generative model to generate high-quality prompts from raw descriptions using visual feedback. |
| Outcome: | The proposed model produces high-quality prompts from simple raw descriptions . it can be integrated to a cloud-native AI platform to provide better image generation service in the cloud. |
Copied to clipboard
| Challenge: | Event extraction (EE) is a critical task in natural language processing, yet deploying a practical EE system remains challenging. |
| Approach: | They propose to leverage the superior extraction capability of LLMs and instruction-following ability of LRMs to construct a robust and highly available EE system. |
| Outcome: | The proposed method can identify and correct errors in SLMs predictions based on automatically generated feedback information and improve performance. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks but their performance in complex logical reasoning tasks remains unsatisfactory. |
| Approach: | They propose a propositional logic prompting method which generates expanded logical information descriptions and utilizes them as an additional augmentation to original contexts. |
| Outcome: | Extensive experiments show that Logic-of-Thought boosts the performance of various prompting methods with a striking margin across five logical reasoning tasks. |
Copied to clipboard
| Challenge: | Recent advances in pre-trained language models have made it possible to generate human-like text. |
| Approach: | They propose to integrate an open-ended text adventure game in Chinese, named KuiLeiXi, where players interact with the AI until the plot goals are reached. |
| Outcome: | The proposed game lacks incentives and relies on players to explore on their own. |
Copied to clipboard
| Challenge: | Pre-trained language models process text as a sequence of characters, ignoring more coarse granularity, e.g., words. |
| Approach: | They propose a new pre-training paradigm for Chinese that incorporates word representations along with characters and can model a sentence in a multi-granular manner. |
| Outcome: | The proposed model can bring an average increase of 1.5% under the 12-layer setting, which achieves new state-of-the-art among base-size models on the CLUE benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for prompt injection have focused on optimizing the suffix, overlooking the role of the prompt. |
| Approach: | They propose a method that incorporates an efficient optimization algorithm and two semantics-guided prompt organization strategies to optimize the suffix sequence for universal goal hijacking. |
| Outcome: | The proposed method can generate a fixed suffix that can concatenate to arbitrary user prompts for universal goal hijacking. |
Copied to clipboard
| Challenge: | Existing studies overlook the multi-turn instruction following ability of large language models (LLMs) Extensive experiments show that Parrot improves current LLMs by up to 7.2% in multi- turn instruction following. |
| Approach: | They propose a method for collecting multi-turn instructions that feature human-like queries, such as anaphora and ellipsis, and a context-aware preference optimization strategy to further enhance LLMs for complex queries. |
| Outcome: | The proposed method improves existing LLMs by up to 7.2% in multi-turn instruction following. |
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) exhibit significant limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. |
| Approach: | They propose a benchmark that provides a fine-grained evaluation of MLLMs’ perception and reasoning capabilities. |
| Outcome: | The proposed benchmark shows that existing MLLMs exhibit limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. |
Copied to clipboard
| Challenge: | Extreme multi-label text classification (EMTC) involves predicting multiple labels from a vast pool of candidates based on a user’s textual query. |
| Approach: | They propose a Quantized and Efficient Learning with Sampling Technique that uses a hash sampling module to reduce the data volume to one-fourth of its original size. |
| Outcome: | Extensive experiments show that QUEST outperforms existing methods while requiring fewer computational resources. |
Copied to clipboard
| Challenge: | Existing relation extraction methods focus on extracting intra-sentence relations for single entities. |
| Approach: | They propose a relation extraction dataset from Wikipedia and Wikidata with three features . document-level relation extraction is a task to identify relational facts between entities . |
| Outcome: | The proposed dataset is the largest human-annotated dataset for document-level RE from plain text. |
Copied to clipboard
| Challenge: | Recent video generation models struggle to synthesize complex dynamics with a coherent chain of consequences. |
| Approach: | They propose a framework that injects visual reasoning signals from multimodal models into video generation. |
| Outcome: | a new framework that leverages multimodal models to generate sparse keyframes significantly improves quality of generated videos. |
Copied to clipboard
| Challenge: | Extensive experiments on a human-labeled golden set showed our tuned PaLM2-XS model achieved 85.56% good ratio. |
| Approach: | They propose a two-stage tuning approach to acquire the dedicated Large Language Model for the feature, followed by a reinforcement learning approach for targeted refinement. |
| Outcome: | The proposed model achieves 85.56% good quality on Rewrite and proofread tasks on human-labeled golden sets. |