Papers by Feng Xu
Copied to clipboard
| Challenge: | Recent pre-trained language models have achieved remarkable zero-shot performance . we propose a self-learning framework that utilizes unlabeled data of target languages . |
| Approach: | They propose a self-learning framework that utilizes unlabeled data of target languages to select silver labels for cross-lingual transfer tasks. |
| Outcome: | The proposed framework outperforms baseline models on two cross-lingual tasks by 10 F1 on average and 2.5 accuracy on natural language inference (NLI). |
Copied to clipboard
| Challenge: | Unsupervised neural machine translation (NMT) is a new approach for machine translation . the model uses only one shared encoder to map pairs of sentences from different languages to a shared-latent space . |
| Approach: | They propose an unsupervised approach which trains the model without labeling data . they propose two independent encoders but share some partial weights to extract high-level representations of input sentences. |
| Outcome: | The proposed approach achieves significant improvements on English-German, English-French and Chinese-to-English translation tasks. |
Copied to clipboard
| Challenge: | Existing top-k attention methods struggle to strike a balance between efficiency and accuracy. |
| Approach: | They propose a top-k attention approach that integrates low-overhead techniques into the Top-k Attention process to achieve 7.2 speedup compared to vanilla full attention. |
| Outcome: | The proposed approach achieves 7.2 speedup compared to current top-k attention methods while maintaining model accuracy. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are gaining popularity due to their lack of knowledge hallucination and lack of a coherent model. |
| Approach: | They propose a self-supervised quantized representation method to compress KG structural and semantic knowledge into discrete codes that align the format of language sentences. |
| Outcome: | The proposed framework outperforms existing unsupervised methods producing more distinguishable codes on KG link prediction and triple classification tasks. |
Copied to clipboard
| Challenge: | Recent advances in large generative models have catalyzed a paradigm shift in content generation to Personalized Generation (PGen). |
| Approach: | They propose a multi-level taxonomy that systematically formalizes PGen's key components, core objectives, and abstract workflows. |
| Outcome: | The proposed taxonomy bridging PGen research across multiple modalities highlights open challenges and promising directions for future exploration. |
Copied to clipboard
| Challenge: | Existing LLM-based agents struggle with low diversity and suboptimal code generation. |
| Approach: | They propose an approach that iteratively expands tree nodes through an introspective process that meticulously analyzes solutions and results from parent and sibling nodes. |
| Outcome: | The proposed approach shows a 4% improvement in performance compared to the strong open-source AutoML agents. |
Copied to clipboard
| Challenge: | Discourse structure analysis is an important research topic in natural language processing. |
| Approach: | They propose to construct a macro discourse structure framework and annotate 147 Newswire articles. |
| Outcome: | The proposed framework can lay the foundation for further analysis of macro discourse structure. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have attracted considerable attention from academic and industrial communities due to their outstanding performance in various natural language processing tasks. |
| Approach: | They propose a Contrastive Learning Framework for Human Alignment to evaluate the noise within the data and dynamically adjust the training process. |
| Outcome: | The proposed framework surpasses other algorithms in terms of reward model scores, automatic evaluations, and human assessments on the widely used dataset "Helpful and Harmless" |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have propelled the development of Conversational Recommendation Agents (CRAs). |
| Approach: | They propose a multi-turn preference optimization paradigm that leverages Expectation Confirmation Theory to explicitly model the evolution of user satisfaction throughout multi-turned dialogues. |
| Outcome: | The proposed paradigm eliminates the significant sampling overhead of existing MTPO methods while ensuring the optimization process drives meaningful improvements. |
Copied to clipboard
| Challenge: | Current studies focus on single-language or single-document tasks for news summarization . lack of a benchmark inhibits researchers from adequately studying this invaluable problem. |
| Approach: | They propose a novel task that unifies Multi-lingual, Cross-lingual and Multi-document Summarization into one task. |
| Outcome: | The proposed task encapsulates the real-world requirements all-in-one and is validated by extensive analysis. |
Copied to clipboard
| Challenge: | Continual learning (CL) for large language models (LLMs) aims to enable sequential knowledge acquisition without catastrophic forgetting. |
| Approach: | They propose a framework that aligns replay schedules with a model-centric notion of time. |
| Outcome: | Experiments on three benchmarks show that FOREVER consistently mitigates catastrophic forgetting. |
Copied to clipboard
| Challenge: | Existing methods to learn from unlabeled data are difficult for zero-shot text classification tasks. |
| Approach: | They propose a self-training based method to efficiently leverage unlabeled data. |
| Outcome: | The proposed method significantly outperforms existing methods in zero-shot text classification tasks on benchmarks and a real-world e-commerce dataset. |
Copied to clipboard
| Challenge: | TexSmart supports fine-grained named entity recognition (NER) Large-scale fine-granular entity types are expected to provide richer semantic information for downstream NLP applications. |
| Approach: | They introduce TexSmart, a text understanding system that supports fine-grained named entity recognition (NER) and enhanced semantic analysis functionalities. |
| Outcome: | The proposed system supports fine-grained named entity recognition (NER) and enhanced semantic analysis functions. |
Copied to clipboard
| Challenge: | Emotion recognition in conversation (ERC) is essential for dialogue systems to identify the emotions expressed by speakers. |
| Approach: | They propose a method that incorporates both belief and desire to accurately identify emotions by extracting emotion-eliciting events from utterances and construct graphs that represent beliefs and desires in conversations. |
| Outcome: | The proposed model outperforms existing models on four popular ERC datasets and validates its performance with multiple state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing models exhibit blind tool-use reasoning patterns, which significantly increases inference overhead and degrades model performance. |
| Approach: | They propose an MLLM that performs adaptive tool-use by determining whether a visual problem truly requires tools. |
| Outcome: | The proposed model outperforms existing methods in visual reasoning tasks. |
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information. |
| Approach: | They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information . |
| Outcome: | The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks. |
Copied to clipboard
| Challenge: | Existing intent classification models rely on a pre-defined intent set and supervised labels, which is limited in some practical scenarios. |
| Approach: | They propose to extend an IND intent classifier to an open-world intent set including IND and OOD intents. |
| Outcome: | The proposed task can classify IND and OOD intents while discovering new unlabeled OOD types incrementally. |
Copied to clipboard
| Challenge: | Existing studies have only considered language models as knowledge bases in a static setting . memorizing conflicting information is still challenging for LMs and hinders memorization of other unrelated one-to-one relationships. |
| Approach: | They propose two requirements for treating language models as temporal knowledge bases . they propose a dataset which is aimed at probing temporally-scoped knowledge . |
| Outcome: | The proposed model can store conflicting information and use stored knowledge for temporal knowledge queries. |
Copied to clipboard
| Challenge: | Existing methods for idea generation either trivially prompt LLMs or expose LLM to extensive literature without indicating useful information. |
| Approach: | They propose a chain-of-ideas agent that organizes literature in a chains structure . they propose evaluating idea-generation methods from different perspectives . |
| Outcome: | The proposed agent outperforms existing methods and matches human quality in idea generation. |
Copied to clipboard
| Challenge: | Recent research shows that training with CAD may lead models to overly focus on modified features while ignoring other important contextual information. |
| Approach: | They propose to use contrastive learning to promote global feature alignment and learning counterfactual clues to improve model performance. |
| Outcome: | The proposed method outperforms the state-of-the-art on out-of distribution (OOD) datasets. |
Copied to clipboard
| Challenge: | Recent advances have enabled Large Language Models (LLMs) to potentially exhibit reasoning capabilities, but complex logical reasoning remains a challenge. |
| Approach: | They propose a novel language model that internalizes and emulates the reasoning processes of logical solvers and avoids parsing errors by learning strict adherence to solver syntax and grammar. |
| Outcome: | The proposed model outperforms state-of-the-art solver-augmented language models and few-shot prompting methods on public deductive reasoning benchmarks. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to memorize, recall, and reason in sustained interactions. |
| Approach: | They propose a multimodal real-world conversation benchmark for evaluating open-ended abilities of multimodal large language models. |
| Outcome: | The proposed benchmarks show that the models perform better in open-ended conversations. |
Copied to clipboard
| Challenge: | Existing approaches to generate SQL-to-text using seq2seq models do not capture graph-structured information in SQL query. |
| Approach: | They propose a graph-to-sequence model to encode global structure information into node embeddings. |
| Outcome: | The proposed model outperforms the Seq2Seq and Tree2Sq baselines on the WikiSQL and Stackoverflow datasets. |
Copied to clipboard
| Challenge: | Existing keyphrase extraction models incorrectly determine a keyphrase as a phrase but output other candidates as keyphrases because they contain the same word. |
| Approach: | They propose a new approach that detects both implicit and explicit centrality within a heterogeneous graph as the importance score of each candidate keyphrase. |
| Outcome: | The proposed approach outperforms state-of-the-art keyphrase extraction models on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing work on cross-domain text classification relies on domain-invariant features or task-agnostic features. |
| Approach: | They propose a two-stage framework for cross-domain text classification that leverages or reuses rich labeled data from the source domain and unlabeled data in the target domain. |
| Outcome: | The proposed framework achieves state-of-the-art on a public cross-domain text classification benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to retrieval-augmented generation ignore valuable structure that is crucial for document organization. |
| Approach: | They propose a framework that explicitly incorporates structural information throughout the RAG process. |
| Outcome: | The proposed framework incorporates structural information throughout the RAG process. |
Copied to clipboard
| Challenge: | Existing defenses against forgery are inadequate for healthcare. |
| Approach: | They propose a large-scale benchmark for pre-hoc, evidence-grounded medical forgery detection using a doctor inspection guideline and gold edit locations. |
| Outcome: | Experiments show that the proposed solution can detect and explain medical scans with high fidelity and accuracy. |
Copied to clipboard
| Challenge: | Existing studies on the use of LLMs for estimating user intents are either too far from real human thought processes or require labeled samples. |
| Approach: | They propose a deliberative agent framework that leverages human thought process to build high-level domain knowledge and a tree-structured knowledge base to store refined experience and data. |
| Outcome: | The proposed framework is able to build high-level domain knowledge and efficiently store it across multiple steps. |
Copied to clipboard
| Challenge: | Existing multi-agent systems have shown strong potential for machine translation (MT) but their performance in multidomain translation remains unsatisfactory due to cross-domain word ambiguity . |
| Approach: | They propose a multi-agent collaborative disambiguation framework for MDT that leverages the collaborative capabilities of LLMs for disambiguations. |
| Outcome: | The proposed framework improves translation performance across multiple domains and improves disambiguation accuracy. |
Copied to clipboard
| Challenge: | Recent advances in machine translation have focused on a single pre-trained decoder . encoder-decoder architectures have received relatively little attention in NMT . |
| Approach: | They propose a method that leverages LLMs as MT encoders and pairs them with lightweight decoders to develop universal translation models. |
| Outcome: | The proposed method matches or surpasses baselines in terms of translation quality but achieves 75% reduction in memory footprint of the KV cache. |
Copied to clipboard
| Challenge: | Existing methods for model merging struggle to maintain performance gains as the number of merged models increases. |
| Approach: | They propose a Reparameterized Heavy-Tailed method to extend the merged model’s coverage and enhance performance. |
| Outcome: | The proposed method extends the merged model’s coverage and enhances performance on 19 benchmarks, including knowledge-intensive and general-purpose tasks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to the development of LLM-based autonomous agents. |
| Approach: | They propose a Reinforcement Learning-based Human-Agent Collaboration method which trains a policy model to determine the most opportune stages for human intervention within the task-solving process. |
| Outcome: | The proposed method improves human-agent collaboration significantly through well-planned, limited human intervention. |
Copied to clipboard
| Challenge: | Existing studies focus only on textual quality and numerical accuracy for headline generation. |
| Approach: | They propose a framework for using rationales of key elements of Topic, Entities, and Numerical reasoning in news articles to enhance LLMs' ability to generate topic-aligned texts with precise numerical accuracy. |
| Outcome: | The proposed framework improves the ability of large language models to generate high-quality texts with precise numerical accuracy. |
Copied to clipboard
| Challenge: | Large-scale industrial ranking systems operate under stringent real-time performance requirements. |
| Approach: | They propose a client-side framework that determines whether a user’s query is complete at each typing . this method leverages client-based typing behavior for real-time early prediction . |
| Outcome: | The proposed framework achieves offline precision/recall/accuracy of 0.7936/0.8196/0.7742 and decreases online response time by 640.5193.65 milliseconds. |
Copied to clipboard
| Challenge: | Effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. |
| Approach: | They propose a multi-dimensional benchmark for evaluating LLMs’ alignment in Chinese with 8 main categories, 683 real-scenario rooted queries and corresponding human verified references. |
| Outcome: | The benchmark uses a human-in-the-loop data curation pipeline, 683 real-scenario rooted queries and human verified references. |
Copied to clipboard
| Challenge: | MMDialog is a dataset of 1.08 million real-world dialogues with 1.53 million unique images across 4,184 topics. |
| Approach: | They propose to use a curated set of 1.08 million dialogues with 1.53 million unique images to generalize the open domain. |
| Outcome: | The proposed system can predict responses to multi-modal content with state-of-the-art techniques and measure their performance. |
Copied to clipboard
| Challenge: | RAAMove is a comprehensive multi-domain corpus dedicated to the annotation of move structures in Research Article (RA) abstracts. |
| Approach: | They propose a multi-domain corpus dedicated to the annotation of move structures in RA abstracts. |
| Outcome: | The proposed corpus is based on a human-annotated dataset and a BERT-based model to verify its effectiveness. |
Copied to clipboard
| Challenge: | Existing methods decompose the COQE task into multiple subtasks and solve them in a pipeline manner, but ignore the intrinsic connection between subtask and the error propagation among stages. |
| Approach: | They propose a unified generative model that solves COQE in one shot by concatenating all the comparative tuples into a target output sequence. |
| Outcome: | The proposed model significantly outperforms the SOTA method on multiple benchmarks and ablation experiments. |
Copied to clipboard
| Challenge: | Information retrieval (IR) is an indispensable technique for locating relevant resources from vast amounts of data. |
| Approach: | They propose a framework that facilitates information refinement through synergy between RMs and LLMs. |
| Outcome: | The proposed framework improves the performance of large-scale retrieval benchmarks on web searches and low-resource retrieval tasks. |
Copied to clipboard
| Challenge: | Existing systems require users to manually select models or employ rigid routing rules that fail to capture the continuous spectrum of query complexity. |
| Approach: | They propose a quality-constrained intelligent prompt routing framework that automatically selects optimal models based on predicted response quality and user-specified tolerance levels. |
| Outcome: | The proposed framework achieves 43.9% cost reduction while maintaining quality parity with strongest model in the Claude family and processes requests with sub-150ms latency. |
Copied to clipboard
| Challenge: | Existing evaluations of large audio-language models focus on answer accuracy and robustness to acoustic perturbations, but they assume that inputs remain semantically answerable. |
| Approach: | They propose a repair-aware evaluation setting that explicitly distinguishes between answerable and unanswerable audio inputs. |
| Outcome: | The proposed evaluation setting distinguishes between answerable and unanswerable audio inputs. |
Copied to clipboard
| Challenge: | Recent sequential modeling approaches focus on extracting information from textual sources while ignoring rich information from other modalities such as image and web layout. |
| Approach: | They propose a novel MUltimodal Structural Transformer that integrates multiple modalities for web information extraction. |
| Outcome: | The proposed model outperforms existing methods on WebSRC and Common Crawl benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to detect contamination of public benchmarks are too superficial to reflect deeper forms of contamination. |
| Approach: | They propose generalization-based approaches to unmask a cross-lingual form of contamination that inflates LLMs’ performance while evading current detection methods. |
| Outcome: | The proposed model outperforms existing detection methods while avoiding contamination of public benchmarks in the pre-training data. |
Copied to clipboard
| Challenge: | Spoken Language Understanding (SLU) is a task-oriented dialogue system . open-source toolkit provides a unified, modularized, and extensible toolkit for SLU . |
| Approach: | They introduce an open-source toolkit to provide a unified toolkit for spoken language understanding. |
| Outcome: | The proposed toolkit unifies 10 models for both single-intent and multi-intention scenarios. |
Copied to clipboard
| Challenge: | Discourse analysis is becoming increasingly important in the field of natural language processing. |
| Approach: | They propose to annotate macro discourse information and additional discourse information to make annotation more objective and accurate. |
| Outcome: | The results show that the annotations are more objective and accurate than the previous ones. |
Copied to clipboard
| Challenge: | Existing slot filling models memorize inherent patterns of entities and contexts from training data. |
| Approach: | They propose a perturbed semantic structure awareness transferring method for slot filling models . they use two MLM-based training strategies to learn contextual semantic structure and word distribution . |
| Outcome: | The proposed method outperforms existing methods and gains strong generalization while preventing model from memorizing inherent patterns of entities and contexts. |
Copied to clipboard
| Challenge: | Existing evaluation methods for human-machine interactions are static and can be misleading. |
| Approach: | They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement. |
| Outcome: | The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment. |
Copied to clipboard
| Challenge: | Existing approaches to generate narrative-driven recommendation are based on large language models (LLMs) but the RAG paradigm is inherently ill-suited for such special queries. |
| Approach: | They propose a novel retrieve-rank paradigm that generatively retrieves structurally adaptive and semantically aligned candidates, ensuring both extensive candidate coverage and high-quality information. |
| Outcome: | The proposed paradigm outperforms the existing paradigm and the existing one under real-world scenarios. |
Copied to clipboard
| Challenge: | Existing approaches to multimodal sentiment analysis treat entire modality as an independent unit for feature enhancement or denoising, which often suppresses redundant noise at the cost of weakening critical information. |
| Approach: | They propose a ModaLity-aware noise dynAmic editiNg framework that performs modality-awful block partitioning by dividing features of each modality into multiple blocks. |
| Outcome: | Experiments on five models and four datasets show that MoLAN+ achieves the state-of-the-art performance. |
Copied to clipboard
| Challenge: | Recent code large language models have demonstrated impressive performance on code-related tasks. |
| Approach: | They propose a paradigm that learns from expert battles to address these limitations . they create an arena where leading LLMs challenge each other with evaluations . |
| Outcome: | The proposed model improves on existing models by leveraging expert battles . it achieves state-of-the-art performance even without relying on proprietary models . |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) are coarse, single-dimensional metrics and do not explicitly assess fine-grained legal reasoning. |
| Approach: | They propose a Practical Law Benchmark to evaluate large language models in real-world legal practice scenarios. |
| Outcome: | The proposed model is based on 850 questions and 13 scenarios with expert-designed evaluation rubrics. |
Copied to clipboard
| Challenge: | Continual learning (CL) is crucial for large language models without costly retraining. |
| Approach: | They propose a framework for recurrent knowledge identification and fusion that enables dynamic estimation of parameter importance distributions to enhance knowledge transfer. |
| Outcome: | The proposed framework mitigates catastrophic forgetting and enhances knowledge transfer. |
Copied to clipboard
| Challenge: | Existing black-box jailbreak methods often rely on model feedback . existing methods may be intercepted by content moderators during the search process . |
| Approach: | They propose a method that guides malicious prompt construction by local training a mirror model of the target black-box model through benign data distillation. |
| Outcome: | The proposed method achieves a 92% attack success rate and 80% stealth rate on a subset of AdvBench. |
Copied to clipboard
| Challenge: | Recent work on neural machine translation (NMT) has demonstrated impressive performance improvements and became the de-facto standard. |
| Approach: | They propose a dynamic curriculum learning method to reorder training samples in training using a Transformer-based system. |
| Outcome: | The proposed method outperforms baselines on three low-resource machine translation benchmarks and different sized data of WMT’16 En-De. |
Copied to clipboard
| Challenge: | Scientific discovery evolution does not occur ex nihilo but is characterized by structural deepening and reconfiguration of existing functionalities. |
| Approach: | They propose a framework for hypothesis generation based on evolutionary narratives . they extract structured P-M-L-F quadruples from citation networks and introduce a mechanism to assess their semantic compatibility. |
| Outcome: | The proposed framework reduces logical disconnects by evaluating its semantic compatibility. |
Copied to clipboard
| Challenge: | Current music information retrieval systems struggle to meet linguistic diversity challenges . current systems struggle with text queries in non-English languages . |
| Approach: | They propose a music information retrieval system that supports both ABC notation and MIDI . CLaMP 2 includes a multilingual text encoder and a multiple-modal music encoder . |
| Outcome: | The proposed system achieves state-of-the-art results in multilingual semantic search and music classification across modalities. |
Copied to clipboard
| Challenge: | Large Vision-Language Models (LVLMs) suffer from multimodal hallucinations . however, the generated hallucines could influence the models’ subsequent generation . |
| Approach: | They propose a framework to evaluate LVLMs' behaviors when encountering generated hallucinations and a method to revise the output distribution of LVLs with the one derived from the residual visual input. |
| Outcome: | The proposed framework reduces the performance of open-source LVLMs by 31%, indicating that they are prone to accept the generated hallucinations and make false claims that they would not have supported without distractions. |
Copied to clipboard
| Challenge: | Existing Plan-and-Solve prompting methods are difficult to implement for complex questions. |
| Approach: | They propose a plan-and-solve prompting method based on Question Decomposition Meaning Representation (QDMR) it allows LLM to generate a QDMR graph to represent problem-solving logic . |
| Outcome: | The proposed method can represent and execute the problem-solving logic of complex questions more accurately than existing methods. |
Copied to clipboard
| Challenge: | Existing methods struggle to control fine-grained reasoning strategies due to conceptual entanglement in LRMs’ hidden states. |
| Approach: | They propose to decompose strategy-entangled hidden states into a disentangled feature space by using Sparse Autoencoders to identify the few strategy-specific features from the vast pool of SAE features. |
| Outcome: | The proposed method outperforms existing methods by 15% in control effectiveness. |
Copied to clipboard
| Challenge: | Existing benchmark-based evaluations cannot accurately reflect the performance of real-world applications. |
| Approach: | They propose a reliable strategy for domains to choose more robust LLMs for real-world applications. |
| Outcome: | The proposed strategy addresses the challenges faced by domains to choose more robust LLMs for real-world applications. |
Copied to clipboard
| Challenge: | Existing benchmarks fail to assess large language models’ format-following proficiency adequately. |
| Approach: | They propose a benchmark to evaluate large language models' ability to follow complex, domain-specific formats. |
| Outcome: | The proposed framework evaluates large language models' ability to follow complex, domain-specific formats across open-source and closed-source models. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing methods for inference use heuristics to determine which positions to unmask and which tokens to commit . MEDAL is an inference-time scaling framework that integrates Monte Carlo Tree SEarch initialization for Diffusion Language Model inference. |
| Approach: | They propose a framework that integrates Monte Carlo Tree SEarch initialization for Diffusion Language Model inference. |
| Outcome: | The proposed framework achieves 22.0% improvement over existing inference strategies across multiple benchmarks. |
Copied to clipboard
| Challenge: | Prior Diversify-then-Verify pipelines generate interpretations and then retrieve evidence . ambiguous queries require RAG to disambiguate into interpretations that can be answered from corpus . |
| Approach: | They propose a novel approach that unifies diversification with verification by integrating retriever relevance and generator answerability feedback early. |
| Outcome: | The proposed approach improves grounding-aware F1 by 23% over baselines across multiple LLMs. |
Copied to clipboard
| Challenge: | Existing methods suffer from key information loss and difficulty in adjusting the length of compressed sequences based on documentation lengths. |
| Approach: | They propose two strategies for compressing tool documentation into concise and precise summary sequences for tool-using language models. |
| Outcome: | The proposed approach achieves comparable performance to the upper-bound baseline under 16x compression ratio. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on specific application scenarios, emphasizing task completion but failing to dissect the underlying skills that drive these outcomes. |
| Approach: | They propose a Massive Multitask Agent Understanding benchmark that evaluates LLMs across five domains and offline tasks. |
| Outcome: | The Massive Multitask Agent Understanding (MMAU) benchmark evaluates models across five domains including Tool-use, Directed Acyclic Graph (DAG) QA, Data Science and Machine Learning coding, Contest-level programming and Mathematics. |
Copied to clipboard
| Challenge: | Unlike existing MoE approaches that rely on fixed TopK Routing, our dynamic expert selection framework dynamically allocates experts based on the confidence level in expert selection for each input. |
| Approach: | They propose a dynamic expert selection framework that dynamically allocates experts based on the confidence level in expert selection for each input. |
| Outcome: | The proposed method achieves an average improvement of 0.7% with less than 90% activated parameters and outperforms dense models in QA and machine translation tasks. |
Copied to clipboard
| Challenge: | Existing data on MBTI personality detection are based on self-reported labels and fail to capture the full range of population personality traits. |
| Approach: | They construct a manually annotated MBTI personality detection dataset with soft labels under the guidance of psychologists and use them to identify the task. |
| Outcome: | The MBTIBench is the first manually annotated MBti personality detection dataset with soft labels under the guidance of psychologists. |
Copied to clipboard
| Challenge: | Existing methods for product attribute value extraction focus on extracting values for a set of known attributes with sufficient training data. |
| Approach: | They propose a prompt tuning approach to extract attributes from product information using mixed prompts. |
| Outcome: | The proposed approach improves on two product benchmarks and shows parameter-efficient training and avoids model overfitting. |
Copied to clipboard
| Challenge: | Existing methods for integrating internal and external knowledge lack effective control mechanisms for generating hallucinations and dealing with outdated knowledge. |
| Approach: | They propose a framework that decouples, identifies, and purposefully optimizes parameter subspaces related to adherence and robustness. |
| Outcome: | The proposed framework decouples, identifies, and purposefully optimizes parameter subspaces related to adherence and robustness. |
Copied to clipboard
| Challenge: | Existing approaches to memory management rely on final task performance as the primary reward, resulting in severe reward sparsity and ineffective credit assignment. |
| Approach: | They propose a framework for fine-grained feedback alignment using a Chunk-level step reward and Evidence-Anchored Reward Attribution to redistribute global rewards based on memory items utilized as evidence in reasoning. |
| Outcome: | The proposed framework outperforms baselines and supports generalization across different model configurations and backbones. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have enabled agentic systems trained with reinforcement learning over multi-turn interaction, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token and memory costs. |
| Approach: | They propose a framework that represents the accumulated observation-action history as a compact rendered image. |
| Outcome: | The proposed framework preserves over 95% of text-based agent performance while significantly reducing token consumption (>50%), yielding consistent token and memory efficiency. |
Copied to clipboard
| Challenge: | Existing RL methods rely on unstructured self-sampling to fit scalar rewards, resulting in inefficient rollouts. |
| Approach: | They propose a structured template-guided RL framework that augments policy optimization with explicit template guidance. |
| Outcome: | Experiments show that TemplateRL outperforms GRPO and GRPI by 99% on AIME and 41% on AMC with superior stability on weak models and remarkable cross-domain generalization. |
Copied to clipboard
| Challenge: | Existing work in vision language cross-modal reasoning uses binary or multi-choice classification based on source image and textual query. |
| Approach: | They propose a task where a textual premise is the background presumption on each source image. |
| Outcome: | The proposed task is based on a dataset of 15,360 movie screenshots and human-curated premise templates from 6 pre-defined categories. |
Copied to clipboard
| Challenge: | Decompilation is the process of converting compiled code back into a high-level programming language for analysis when source code is unavailable. |
| Approach: | They propose two methods to improve decompilation performance without fine-tuning and fine-grained alignment enhancement to achieve further improvements. |
| Outcome: | The proposed methods achieved a Re-Executability performance improvement of approximately 3.90% on the Decompile-Eval benchmark, establishing a new state-of-the-art performance of 52.41%. |
Copied to clipboard
| Challenge: | Autoregressive (AR) models excel at generating temporally coherent audio by producing tokens sequentially, yet they often falter in faithfully following complex textual prompts. |
| Approach: | They propose a lightweight auxiliary model trained with a GAE-inspired objective to predict final instruction-following quality from partial generations. |
| Outcome: | The proposed model achieves 10 points improvement in CLAP score over baseline AR models while maintaining computational parity with best-of-N decoding. |
Copied to clipboard
| Challenge: | Recent studies have incorporated reward models to guide response selection or decoding, aiming to obtain higher-quality data. |
| Approach: | They propose a Hierarchical Sampling framework for self-taught reasoners that allocates a fixed sampling budget to problem boundary-level problems and then reallocates the remaining budget toward high-utility problems during a re-sampling phase. |
| Outcome: | The proposed framework outperforms baseline models without additional sampling budgets across multiple reasoning benchmarks and backbone LLMs. |
Copied to clipboard
| Challenge: | Existing research results on explicit sentiment analysis are limited . implicit sentiment analysis is a process of analyzing text based on whether it contains explicit sentiment words. |
| Approach: | They propose a model that integrates external knowledge and contextual features . they use a knowledge graph to supplement implicit sentiment expression . |
| Outcome: | The proposed model can achieve better results on the SMP2019 implicit sentiment analysis dataset. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated exceptional performance across a broad spectrum of cross-lingual Natural Language Processing (NLP) tasks. |
| Approach: | They propose to use Word-level Cross-lingual Structure to prove that the word-level embedding on the hidden layers isomorphic between languages. |
| Outcome: | The proposed method significantly improves on two representative LLM foundations, LLaMA2 and BLOOM. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable generality, often solving tasks with a single carefully engineered prompt. |
| Approach: | They propose to cast automatic workflow generation as Bayesian inference over a posterior distribution on workflows and instantiate BayesFlow as Bayer-based workflow generation framework. |
| Outcome: | The proposed framework improves accuracy by 9 percentage points over baselines and 65 percentage points on pool-wide benchmarks. |
Copied to clipboard
| Challenge: | Existing explanations for user reviews often fail to meet user-centric aspects, reducing their usefulness to users. |
| Approach: | They propose a paradigm that refines initial explanations generated by existing models during the inference stage to enhance their quality in multiple aspects. |
| Outcome: | The proposed model improves explanations generated by existing models during the inference stage to enhance their quality in multiple aspects. |
Copied to clipboard
| Challenge: | Existing frameworks for learning and deployment of neural dialogue models have been used for online chit-chat scenarios. |
| Approach: | They propose a hierarchical inductive transfer framework to learn and deploy dialogue skills continually and efficiently. |
| Outcome: | The proposed framework achieves comparable performance under deployment-friendly model capacity. |
Copied to clipboard
| Challenge: | Large-scale high-quality training data is important for improving the performance of models. |
| Approach: | They propose a framework that motivates the model to automatically generate rationales on existing datasets and improves the performance of reasoning through reinforcement learning. |
| Outcome: | The proposed model outperforms InstructGPT on multiple reasoning datasets and outperformed InstructGPT on other datasets. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for text-to-audio-video (T2AV) generation are largely designed for human-recorded videos or single-speaker settings. |
| Approach: | They propose a failure-driven diagnostic benchmark for multi-talker dialogue-centric audio-video generation. |
| Outcome: | The benchmark evaluates multi-speaker dialogue generation at four levels: audio-visual signal fidelity, temporal attribute consistency, social interaction, and cinematic expression. |
Copied to clipboard
| Challenge: | Existing studies have demonstrated that direct preference optimization (DPO) can be effective in generalizing large language models, but its effectiveness in video domain remains limited. |
| Approach: | They propose a framework that utilizes detailed video captions as a proxy of video content to enable language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. |
| Outcome: | The proposed framework shows that it can be used to align language models with video content and improves performance on open-ended video QA tasks. |
Copied to clipboard
| Challenge: | Despite its versatility, CLIP-based applications often suffer from misunderstandings regarding user intent, leading to discrepancies between the required number of objects and the actual outputs. |
| Approach: | They empirically evaluate CLIP’s understanding of quantity from text, image, and cross-modal perspectives by carefully designing different experimental settings and datasets. |
| Outcome: | The proposed model has shown significant success in various downstream tasks, including editing, generation, and quality evaluation. |
Copied to clipboard
| Challenge: | Existing methods to enhance length extrapolation of large language models have been developed, but a systematic survey is lacking. |
| Approach: | They propose to examine the effects of positional encoding on length extrapolation. |
| Outcome: | The proposed methods improve the extrapolation of large language models, but they are still lacking a systematic survey. |
Copied to clipboard
| Challenge: | Existing Key-value Memory Neural Networks are effective for shallow reasoning over documents . but extending them to Knowledge Based Question Answering is not trivial . |
| Approach: | They propose a mechanism to enable conventional KV-MemNNs models to perform interpretable reasoning for complex questions. |
| Outcome: | The proposed solution provides better reasoning abilities on complex questions and achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing prompt tuning methods only introduce prompts at the input layer, limiting performance and leaving large room for improvement. |
| Approach: | They propose a method that involves tuning a small set of soft prompts for pre-trained language models. |
| Outcome: | The proposed method outperforms state-of-the-art methods with pre-trained models on the SuperGLUE benchmark. |
Copied to clipboard
| Challenge: | Existing approaches to multi-turn Text-to-SQL tasks rely on unstable APIs or expensive fine-tuning. |
| Approach: | They propose a training-free framework that leverages small-scale LRMs through in-context learning to enable accurate context-dependent parsing. |
| Outcome: | The proposed framework outperforms in-context learning baselines at the 4B scale and surpasses state-of-the-art models at the 8B and 14B scales. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced natural language processing, demonstrating exceptional reasoning, tool usage, and memory capabilities. |
| Approach: | They propose a competition-based benchmark framework specifically designed to assess LLMs within multi-agent environments. |
| Outcome: | The proposed framework enhances the LLMs’ abilities in navigating complex social and cognitive dimensions by over threefold between the strongest and weakest LLM models. |
Copied to clipboard
| Challenge: | Metaphors are pervasive in communication, making them crucial for natural language processing. |
| Approach: | They propose a multicultural multimodal metaphor dataset designed for cross-cultural studies of metaphor in Chinese and English. |
| Outcome: | The proposed model improves metaphor comprehension across cultural backgrounds and cultural domains. |
Copied to clipboard
| Challenge: | Existing models for LLM role-playing lack high-quality datasets with explicit reasoning traces and reliable reward signals aligned with human preferences. |
| Approach: | They propose a unified framework for cognitive-level persona simulation that strictly distinguishes characters’ first-person thinking processes from LLMs’ third-person reasoning. |
| Outcome: | The proposed framework outperforms the Qwen3-32B baseline model and achieves a 30.26% and 14.97% performance on the minimax benchmarks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for large language models are labor-intensive and lack efficiency. |
| Approach: | They propose a framework dedicated to assessing long-text generation that includes in-depth human-curated meta-questions spanning various domains . they use a set of proxy-quests with pre-annotated answers to assess the content's quality by incorporating the generated texts as contextual background. |
| Outcome: | The proposed framework assesses the quality of long-text content by matching it with references through human evaluation or automated metrics. |
Copied to clipboard
| Challenge: | Recent advances in vision-language-action models prioritize robotic action mastery . however, models trained on visual-text pairs struggle to interpret multimodal data . |
| Approach: | They propose a framework that integrates multimodal data after initial control mastery and a Mixture-of-Experts architecture to minimize task interference. |
| Outcome: | The proposed framework surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks and achieves six times higher performance on visual question-answering datasets. |
Copied to clipboard
| Challenge: | Experimental results show that the proposed model consistently outperforms the traditional RNNSearch and the newly emerged state-of-the-art Transformer on English-German and Chinese-English translation tasks. |
| Approach: | They propose an approach for applying GANs to NMT by building a conditional sequence generative adversarial net with two adversarials. |
| Outcome: | The proposed model outperforms the existing RNNSearch and Transformer on English-German and Chinese-English translation tasks. |
Copied to clipboard
| Challenge: | Existing approaches to mitigate inference inefficiency and optimization difficulty are fragmented and constrained by inherent trade-offs. |
| Approach: | They propose a framework that reconceptualizes discrete reasoning steps as a continuous probabilistic flow, quantifying the contribution of each step toward the ground-truth answer. |
| Outcome: | The proposed framework achieves a superior balance between inference efficiency and reasoning performance on challenging benchmarks. |
Copied to clipboard
| Challenge: | Unified Multimodal Models have achieved remarkable success in cross-modal comprehension, but a gap persists in their ability to translate internal knowledge into faithful and controllable synthesis. |
| Approach: | They propose a self-improvement framework that partitions a single UMM into three collaborative roles: Proposer, Solver, and Judge. |
| Outcome: | The proposed framework improves on TIIF, DPG, CompBench and UniCycle benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in large language models have sparked interest in creating autonomous agents. |
| Approach: | They propose a framework that jointly optimizes both task-planning and self-reflective evolution capabilities in language agents. |
| Outcome: | The proposed framework improves task planning and self-reflective evolution capabilities in language agents. |
Copied to clipboard
| Challenge: | Existing LLM-based recommenders lack explicit modeling of geographic signals . without explicit modeling geographic signals, recommenders struggle to capture core mobility patterns . |
| Approach: | They propose a framework that utilizes geography as a decision variable within the reasoning process. |
| Outcome: | The proposed framework achieves over 10% relative gains in hit rate over strongest LLM-based baselines and improves cross-city transfer. |
Copied to clipboard
| Challenge: | Mixed Boolean-Arithmetic (MBA) expressions are difficult to simplify because of interleaving bitwise and arithmical operations. |
| Approach: | They propose a method to learn and reduce MBA expressions using a string to string method . they propose to use a dataset to train the method to reduce MBA rules . |
| Outcome: | The proposed method outperforms all other tools in terms of accuracy, solving time, and performance overhead. |
Copied to clipboard
| Challenge: | Existing approaches to generate training data with pre-trained language models have been found effective in various scenarios. |
| Approach: | They propose an unsupervised zero-shot learning method that generates a dataset from scratch and trains a tiny task model under supervision of the synthesized dataset. |
| Outcome: | The proposed method is annotated-free and efficient, but can provide useful insights from the perspective of data-free model-agnostic knowledge distillation and unreferenced text generation evaluation. |
Copied to clipboard
| Challenge: | Existing models for dialogue generation lack the flexibility to handle such freedoms. |
| Approach: | They propose to take into account dialogue history and future conversation to implicitly reconstruct the scenario knowledge. |
| Outcome: | The proposed approach outperforms state-of-the-art models on diversity and relevance and expresses scenario-specific knowledge. |
Copied to clipboard
| Challenge: | Recent vision-language models (VLMs) have shown impressive capabilities as general visual assistants, but there are two challenges to their performance: (1) lacking task diversity in pretraining and visual instruction tuning; (2) annotation error and bias in GPT-4 synthesized instruction tuning data. |
| Approach: | They propose a two-stage instruction tuning framework that fine tunes VLMs firstly and further tuned on GPT-4 synthesized data. |
| Outcome: | The proposed framework outperforms the traditional single-stage visual instruction tuning framework and achieves state-of-the-art performance across a wide range of multi-modal evaluation benchmarks. |
Copied to clipboard
| Challenge: | Detecting disfluency can be difficult because of the flexible nature of reparandum structure and the lack of a nested structure. |
| Approach: | They propose a semi-supervised approach which extracts hidden features from self-attention without any Recurrent Neural Network (RNN) or Convolutional Neural Net (CNN). |
| Outcome: | The proposed approach improves over baselines by using unlabelled data . identifying and removing non-fluent factors would help to improve spontaneous speech quality . |
Copied to clipboard
| Challenge: | Existing automatic question generation methods focus on encoding passage and answer to generate question. |
| Approach: | They propose an automatic question generation approach which integrates question generation with its dual problem, question answering, into a unified primal-dual framework. |
| Outcome: | The proposed approach outperforms existing methods on SQuAD and HotpotQA benchmarks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities in environmental perception, reasoning-based decision-making, and simulating complex human behaviors, particularly in interactive role-playing contexts. |
| Approach: | They propose a framework to assess LLMs' proficiency in portraying advanced human behaviors through murder mystery games using eight intricately crafted scripts. |
| Outcome: | The framework evaluates LLMs' performance in portraying advanced human behaviors through murder mystery games. |
Copied to clipboard
| Challenge: | Parameter-Efficient Fine-Tuning (PEFT) is an alternative to Full-Parameter Fine-tuning, but its effectiveness on complex tasks such as reasoning and instruction-following remains unclear. |
| Approach: | They propose to use PEFT to reduce the number of trainable parameters while freezing the weights of LLMs. |
| Outcome: | The proposed methods perform well on standard tasks, but weaknesses on complex and adversarial settings call for new directions beyond current paradigms. |
Copied to clipboard
| Challenge: | Current methods for Continual Dialogue State Tracking (DST) struggle with catastrophic forgetting and knowledge transfer between tasks. |
| Approach: | They propose a framework for task skill localization and consolidation that enables effective knowledge transfer without relying on memory replay. |
| Outcome: | The proposed framework shows a 7.6% increase in Avg. JGA and 11% rise in BWT metrics over existing state-of-the-art methods. |
Copied to clipboard
| Challenge: | Visual reasoning is a multi-step and compositional problem that requires intensive text-vision interactions. |
| Approach: | They propose a visual reasoning model that uses a feature-wise linear modulation technique to enable textual/visual pipelines to mutually control each other. |
| Outcome: | The proposed model outperforms existing models on visual reasoning benchmarks CLEVR and NLVR . it can generate a textual answer to a visual question answering problem with images . |
Copied to clipboard
| Challenge: | Existing methods to extract product attribute value require multiple extractions to obtain all corresponding values. |
| Approach: | They propose an Efficient product Attribute Value Extraction approach using lightweight sparse-layer interaction. |
| Outcome: | The proposed method achieves significant efficiency gains with neutral or marginal loss in performance when the context is long and number of attributes is large. |
Copied to clipboard
| Challenge: | Existing methods to detect and correct spelling errors in Chinese take external input or just heuristic rules. |
| Approach: | They propose to incorporate phonological and visual similarity knowledge into Chinese language models by using a specialized graph convolutional network. |
| Outcome: | The proposed method outperforms existing models on three human-annotated datasets. |
Copied to clipboard
| Challenge: | Existing prompting methods for Large Language Models (LLMs) suffer from excessive token usage and limited generalisability across diverse reasoning tasks. |
| Approach: | They propose an Adaptive Causal Prompting with Sketch-of-Thought framework that leverages structural causal models to infer the causal effect of a query on its answer. |
| Outcome: | The proposed framework outperforms existing prompting baselines in terms of accuracy, robustness, and computational efficiency. |
Copied to clipboard
| Challenge: | Existing models that generate generic aspects do not provide personalized informative recommendations. |
| Approach: | They propose a model that integrates aspect category as another input dimension to facilitate memorizing fine-grained aspect terms. |
| Outcome: | The proposed model outperforms baseline model on restaurant review datasets in the restaurant domain. |
Copied to clipboard
| Challenge: | Large language models have significantly advanced Multilingual Machine Translation (MMT) yet scaling to many languages while maintaining robust performance across directions remains challenging. |
| Approach: | They propose a strategy to reduce the number of translations in one direction . they propose auxiliary parallel sentences to promote cross-lingual transfer . |
| Outcome: | The proposed model performs on par with or better than substantially larger baselines. |
Copied to clipboard
| Challenge: | Using a multi-modal multi-granularity tokenizer, we analyze ancient Chinese scripts . a large proportion of the characters in ancient Chinese are rare or undeciphered . |
| Approach: | They propose a multi-modal multi-granularity tokenizer specifically designed for ancient Chinese scripts. |
| Outcome: | The proposed tokenizer improves on the part-of-speech tagging task on the Chu bamboo slip script. |
Copied to clipboard
| Challenge: | Graph2Tree model encodes graph-structured input and decodes tree-structures output. |
| Approach: | They propose a novel Graph-to-Tree Neural Network consisting of a graph encoder and a hierarchical tree decoder that encodes an augmented graph-structured input and decodes a tree-structure-output. |
| Outcome: | The proposed model outperforms or matches the performance of other state-of-the-art models on two problems, neural semantic parsing and math word problem. |
Copied to clipboard
| Challenge: | Existing methods for extracting chemical procedures from literature are insufficient and low-quality due to the inherent ambiguity of chemical language and the high cost of human annotation. |
| Approach: | They propose a fully fine-tuned large language model (LLM) as a chemical executor to convert between unstructured experimental procedures and structured action sequences. |
| Outcome: | The proposed model outperforms the baseline model on R2D and D2A tasks by 10%. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation abstracts the meaning of sentences into a single-rooted, acyclic and directed graph. |
| Approach: | They propose to use a metric to evaluate concept alignment and relation alignment to improve Chinese AMR parsing evaluation methods. |
| Outcome: | The proposed method is more robust and compatible with concept alignment and relation alignment and more robust in evaluating arcs. |
Copied to clipboard
| Challenge: | Existing training-based model editing methods struggle to incorporate new knowledge while preserving unrelated general knowledge. |
| Approach: | They propose a framework that uses geometric relationships to differentiate between neurons associated with new knowledge updates and those related to general knowledge perturbations. |
| Outcome: | The proposed framework avoids updating neurons with directions approximately orthogonal to existing knowledge, thus preserving the model’s generalization ability. |
Copied to clipboard
| Challenge: | Existing studies assessing the spatial abilities of VLMs lack a solid theoretical foundation and lack measurable data. |
| Approach: | They propose a psychometric framework defining five basic spatial abilities in Visual Language Models. |
| Outcome: | The proposed framework defines five basic spatial abilities in Visual Language Models (VLMs) it provides a comprehensive evaluation benchmark and methodological perspective for embodied AI development . |
Copied to clipboard
| Challenge: | Prior work on out-of-domain detection requires in-domain task labels and is limited to supervised classification scenarios. |
| Approach: | They propose a method to construct out-of-domain detectors efficiently using pre-trained transformers. |
| Outcome: | The proposed method greatly improves out-of-domain detection ability in a more general scenario. |
Copied to clipboard
| Challenge: | Autonomous agents powered by large language models (LLMs) have attracted significant research interest, but there are few standards for developing specialized models for agent tasks. |
| Approach: | They propose a series of large action models with dense and mixture-of-expert architectures that unifies, augments, and synthesizes diverse datasets to enhance agent generalizability and performance. |
| Outcome: | The proposed models outperform GPT-4, Claude-3, and many other models in terms of tool use and outperformed GPT-based models on multiple agent ability benchmarks. |
Copied to clipboard
| Challenge: | Teaching large language models to generate text with citations to evidence sources requires high-quality attribution data, which is costly and labor-intensive. |
| Approach: | They propose a framework for iteratively improving the attribution capability of large language models (LLMs) by attributing output to verifiable sources. |
| Outcome: | Experiments on three open-domain question-answering datasets show that START improves in aggregating information across multiple sources. |
Copied to clipboard
| Challenge: | Recent advances in NLP are driven by a variety of Large Language Models (LLMs), such as GPT-3 (175B) and PaLM (540B). |
| Approach: | They propose a taxonomy that categorizes the methods into four groups and summarizes the metrics for evaluating the generation quality. |
| Outcome: | The proposed taxonomy categorizes the generation methods into four groups and summarizes the metrics for evaluating the quality. |
Copied to clipboard
| Challenge: | Existing scaling methods for extending context window rely on empirical approaches and lack understanding of the internal distribution within RoPE resulting in suboptimal performance. |
| Approach: | They propose to optimize the context window extending task from the view of rotary angle distribution by minimizing disturbance between rotary angles to maintain consistency with the pre-training phase. |
| Outcome: | The proposed approach reduces by up to 72% of the distributional disturbance when extending LLaMA2’s context window to 8k, and reduces it by up 32% when extending to 16k. |
Copied to clipboard
| Challenge: | Recent model merging-based methods struggle to effectively manage the trade-off between learning new knowledge and preventing catastrophic forgetting. |
| Approach: | They propose a model merging framework that utilizes learning and forgetting signals from the training trajectory to dynamically monitor the model’s training status. |
| Outcome: | The proposed framework achieves significant performance improvements over existing state-of-the-art methods on three CL benchmarks with various model sizes (from 770M to 13B). |
Copied to clipboard
| Challenge: | Existing studies have shown the effectiveness of sequence-to-sequence (Seq2Seque) on mathematics solving. |
| Approach: | They propose a graph-to-sequence neural network which can learn hierarchical information of graphs inputs to solve mathematical problems and speculate answers. |
| Outcome: | The proposed neural network outperforms other neural networks in hidden information learning and mathematics resolving. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual knowledge graph (KG) alignment rely on entity embeddings derived from monolingual KG structural information. |
| Approach: | They propose a topic entity graph to represent entities with contextual information in KGs. |
| Outcome: | The proposed model outperforms state-of-the-art methods by a large margin. |
Copied to clipboard
| Challenge: | In-context Learning (ICL) is a new paradigm for large language model evaluation. |
| Approach: | They propose an open-source toolkit for ICL and LLM evaluation. |
| Outcome: | The proposed framework is highly flexible and flexible and can be easily combined with other tools to suit users' needs. |
Copied to clipboard
| Challenge: | Prior studies have focused on strengthening multimodal reasoning by improving representation alignment or increasing computation, but these methods do not characterize the differences in visual demands across tasks. |
| Approach: | They propose an entropy-driven task-adaptive visual attention allocation framework that uses visual attention entropic as a control signal to dynamically allocate attention according to task demands. |
| Outcome: | The proposed framework achieves consistent performance gains across diverse reasoning tasks, datasets, and models, providing a clear direction toward more reliable multimodal reasoning. |
Copied to clipboard
| Challenge: | Existing RP-LLMs employ only a single role with numerous dialogues, but Crab enables dynamic configuration of desired roles, thereby enhancing related flexibility and adaptability. |
| Approach: | They propose a Configurable Role-Playing LLM with Assessing Benchmark that combines a Role dataset curation, persona-emodying Llm construction, and comprehensive benchmark creation for RP dialogue generation. |
| Outcome: | The proposed model outperforms existing LLMs in performing fine-grained evaluations of RP while keeping dialogue per role minimal. |
Copied to clipboard
| Challenge: | Existing retrieval-augmented strategies for large language models fail to capture dynamic reasoning required to resolve execution failures. |
| Approach: | They propose a framework that implements Experience Replay to transform transient rectification steps into persistent knowledge. |
| Outcome: | The proposed framework improves model accuracy by 8.45% on complex tasks while reducing token consumption by 28.65% and interaction turns by 25.82%. |
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) exhibit significant limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. |
| Approach: | They propose a benchmark that provides a fine-grained evaluation of MLLMs’ perception and reasoning capabilities. |
| Outcome: | The proposed benchmark shows that existing MLLMs exhibit limitations when extracting essential information and reasoned properties from diagrams and performing complex reasoning based on these visual inputs. |