Papers by Ping Li
Copied to clipboard
| Challenge: | Empirical studies show that our approach gains approximately an improvement of 1 BLEU score on most benchmarks over the Transformer baseline. |
| Approach: | They propose to extract several semantic kernels from a source sentence to capture global semantic information. |
| Outcome: | Empirical results show that the proposed approach improves 1 BLEU score on benchmarks . it is also 1.7 times faster than previous works on average at inference time . |
Copied to clipboard
| Challenge: | Existing pruning methods require inefficient retraining for billion-scale LLMs or rely on heuristicically designed metrics to determine pruning masks, leading to performance degradation. |
| Approach: | They propose a convex optimization model that induces sparsity in large language models by leveraging FISTA. |
| Outcome: | The proposed method can remove 50% of model parameters while retaining 98.6% and 95.6% of the zero-shot performance. |
Copied to clipboard
| Challenge: | Different Open Information Extraction (OIE) tasks require different types of information. |
| Approach: | They propose to adapt an OIE Graph to different OIE tasks with simple rules . they implement an end-to-end OIA generator and make it open-accessible . |
| Outcome: | The proposed system achieves new SOTA performance on three popular OIE tasks. |
Copied to clipboard
| Challenge: | Recent work on distantly supervised (DS) ultra-fine entity typing has received significant attention . however, DS data is noisy and often suffers from missing or wrong labeling issues resulting in low precision and low recall. |
| Approach: | They propose a noise model to estimate unknown labeling noise distribution over input contexts and noisy type labels and a model to train on denoised data. |
| Outcome: | The proposed model outperforms baseline methods on the Ultra-Fine entity typing dataset and OntoNotes dataset. |
Copied to clipboard
| Challenge: | Existing top-k attention methods struggle to strike a balance between efficiency and accuracy. |
| Approach: | They propose a top-k attention approach that integrates low-overhead techniques into the Top-k Attention process to achieve 7.2 speedup compared to vanilla full attention. |
| Outcome: | The proposed approach achieves 7.2 speedup compared to current top-k attention methods while maintaining model accuracy. |
Copied to clipboard
| Challenge: | Existing multilingual video corpus moment retrieval methods are based on a two-stream structure. |
| Approach: | They propose a multilingual video corpus moment retrieval task that uses a two-stream structure to generate a query-visual similarity and a subtitle stream exploits the query-subtitle similarity. |
| Outcome: | The proposed method improves accuracy on a large-scale video corpus moment retrieval dataset. |
Copied to clipboard
| Challenge: | Paraphrase generation is an important natural language generation task . however, the effectiveness of paraphrase generation can be limited due to the limited data available. |
| Approach: | They propose a weakly supervised approach to paraphrase generation that leverages reinforcement learning for effective model training with data selection. |
| Outcome: | The proposed model improves the state-of-the-art performance on four weakly supervised paraphrase generation tasks. |
Copied to clipboard
| Challenge: | Existing representation learning methods such as Word2vec represent word embeddings in the semantic space. |
| Approach: | They propose an efficient method for searching vectors via a non-metric matching function: inner product. |
| Outcome: | Experiments on data representations learned for different machine learning tasks show the proposed method outperforms existing methods. |
Copied to clipboard
| Challenge: | Existing document understanding models focus on entity categories while ignoring the extraction of entity boundaries. |
| Approach: | They propose a hypergraph attention document semantic entity recognition framework which uses hypergraph focus to focus on entity boundaries and entity categories at the same time. |
| Outcome: | The proposed framework can improve the performance of existing models on FUNSD, CORD, XFUND and SROIE. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit notable deficiencies in temporal reasoning . phrasing changes can lead LLMs to produce inconsistent outputs . |
| Approach: | They investigate the mechanistic interpretability of temporal ordering within event temporal reasoning . they identify a sparse subset of attention heads that are causally responsible for reasoning outcomes . |
| Outcome: | The proposed model outperforms other models in a variety of tasks and is validated by intervention-based experiments. |
Copied to clipboard
| Challenge: | Existing methods for image-to-text generation store all knowledge within parameters, thus requiring computational-expensive fine-tuning. |
| Approach: | They propose a Retrieval-augmented Visual Language Model that stores all the knowledge within parameters and can be used to retrieve it from the external database. |
| Outcome: | The proposed model significantly boosts performance for image-to-text generation tasks with 4x less parameters compared with baseline methods. |
Copied to clipboard
| Challenge: | Existing evaluation methods for text style transfer are unsatisfactory. |
| Approach: | They propose to use a graph-based method to extract attribute content from sentences . they propose an efficient regularization to leverage attribute-dependent content as guiding signals. |
| Outcome: | The proposed method is based on a YELP and IMDB dataset and it is able to detect errors in the human evaluation. |
Copied to clipboard
| Challenge: | Existing methods for decoding large language models generate one token per step, causing high inference latency. |
| Approach: | They propose a method that integrates retrieved exact patterns with logit-driven future cues. |
| Outcome: | Experiments on Spec-Bench, HumanEval, and MGSM-ZH show that RACER outperforms training-free methods and accelerates inference. |
Copied to clipboard
| Challenge: | Long-context Document Visual Question Answering (DocVQA) methods struggle with visual semantics or handling finite context windows. |
| Approach: | They propose a new approach to longcontext document visual question answering that transforms retrieval into adaptive evidence chain construction using a Bi-Layered Graph. |
| Outcome: | The proposed approach achieves an average accuracy improvement of 14.07% on M5BookVQA and exhibits robust generalization with a 13.38% gain across four established benchmarks. |
Copied to clipboard
| Challenge: | Emotion classification is an important task with applications in education, virtual reality, and robotics. |
| Approach: | They propose to use token embeddings to generate a "semantic-anchor graph" using semantic anchors, sentences can be projected onto them to form a graph . |
| Outcome: | Empirically, the proposed system can generate meaningful semantic anchors and discriminative graph patterns for different emotion. |
Copied to clipboard
| Challenge: | Existing LLMs often rely on complex prompting or extensive fine-tuning to introduce new capabilities while preserving strong generalizability. |
| Approach: | They propose a large-scale pre-training corpus to enhance LLM agents' capabilities . they use 103B agent-specific data encompassing 76,537 APIs . |
| Outcome: | The proposed training corpus outperforms open-source LLMs and commercial LLM agents on three agent benchmarks. |
Copied to clipboard
| Challenge: | Textual backdoor attacks are vulnerable to backdoors and can be used to infect models trained on poisoned data. |
| Approach: | They propose an efficient attribution-based pipeline to defend against two insertion-based poisoning attacks, BadNL and InSent. |
| Outcome: | The proposed method can generalize sufficiently well in two common attack scenarios, which consistently improves previous methods. |
Copied to clipboard
| Challenge: | a helpful review is largely concerned with the metadata of its target product . a selector learns from both the key-value product metadata and one of its reviews to take an action . |
| Approach: | They propose a framework that uses product metadata to assess helpfulness of free-text reviews . they use two real-world datasets from amazon.com and Yelp.com to test the framework . |
| Outcome: | The proposed framework can achieve state-of-the-art performance with substantial improvements . it uses two real-world datasets from Amazon.com and Yelp.com . |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved remarkable success across diverse domains, but their potential as effective language teachers remains inadequately assessed. |
| Approach: | They propose a framework to evaluate Chinese language teachers' pedagogical competence against international standards. |
| Outcome: | The proposed framework evaluates 13 latest multilingual and Chinese LLMs against international standards for Chinese language teachers. |
Copied to clipboard
| Challenge: | Existing metrics for video captioning are based on text-based comparisons with ground-truth references. |
| Approach: | They propose a reference-free benchmark that assesses video captions based on their utility . they will release the benchmark to facilitate reproducible research . |
| Outcome: | The proposed benchmark improves on human-verified, fine-grained questions . it correlates significantly better with human judgments than existing metrics . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive results in Machine Translation by following instructions, even without training on parallel data. |
| Approach: | They propose a Translate After LEarNing Textbook approach which aims to enhance LLMs’ ability to translate low-resource languages by learning from a textbook. |
| Outcome: | The proposed approach improves translation performance by 14.8% using 112 low-resource languages from FLORES-200 with two LLMs: ChatGPT and BLOOMZ. |
Copied to clipboard
| Challenge: | Scaling LLM-based agents to long-horizon deep research is constrained by context-noise trade-off . solving a single query may require hundreds of interactions with noisy environments . |
| Approach: | They propose a factorized memory architecture that decouples the cognitive state into a Fluid Working Context for immediate reasoning and a persistent Knowledge Graph for long-term retention. |
| Outcome: | The Cognitive Scaffold outperforms baselines on Xbench-DeepSearch, BrowseComp-ZH, and GAIA . it achieves 74.7% Avg@3 and 87.0% Pass@3 on xbench, browseComp, and 88.3% Pass@3. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
Copied to clipboard
| Challenge: | Recent advances in disentanglement work on coarse levels in the disenanglement of closely related properties, such as syntax and semantics in human languages. |
| Approach: | They propose a deep decomposable model based on VAE to disentangle syntax and semantics by using total correlation penalties on KL divergences. |
| Outcome: | The proposed model significantly improves the disentanglement quality between syntactic and semantic representations for semantic similarity tasks and syntaktic similarity task. |
Copied to clipboard
| Challenge: | In-Context Learning (ICL) is a key method in prompt engineering, but its long retrieved contexts and limited token throughput will slow reasoning speeds. |
| Approach: | They propose a method that leverages the overlap between context and model output to generate drafts from the context. |
| Outcome: | The proposed method achieves the highest mean speedup on Vicuna-7B, Llama2-7B-Chat, and Llma3-8B-Instruct tasks. |
Copied to clipboard
| Challenge: | retrieval-augmented generation (RAG) is a powerful tool for NLP applications . but it is challenging to encode large knowledge bases as compact offline structures . |
| Approach: | They propose a coarse-to-fine hierarchical graph inference method that uses random walks to retrieve information from a corpus of documents. |
| Outcome: | The proposed method reduces offline indexing costs and accelerates retrieval. |
Copied to clipboard
| Challenge: | a recent study shows that late-interaction methods trade off retrieval accuracy and efficiency by exploiting cross-modal interactions only in the late stage. |
| Approach: | They propose an inflating and shrinking approach to exploit cross-modal interactions . they inflate code inputs and shrink code outputs to exploit interactions progressively . |
| Outcome: | The proposed method exploits cross-modal interactions in the late stage to achieve retrieval speed. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) require substantial computational resources during deployment. |
| Approach: | They propose a method to identify outlier tokens and exclude them from quantization . they find that the method can deliver a 6.4 times reduction in memory usage and a 2.5 times increase in throughput . |
| Outcome: | The proposed method delivers a 6.4 times reduction in memory usage and a 2.5 times increase in throughput under 2-bit quantization. |
Copied to clipboard
| Challenge: | Long-video understanding is bottlenecked by the high cost of processing massive visual tokens. |
| Approach: | They propose a decoupled framework for query-guided visual token pruning . their method reduces visual tokens by 90% and accelerates inference by 98% . |
| Outcome: | The proposed framework reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average. |
Copied to clipboard
| Challenge: | Recent studies have focused on neural information retrieval (IR) models. |
| Approach: | They propose a framework which can benefit from both labeled and more abundant unlabeled data . they propose supervised retrieval over several strong baselines for IR . |
| Outcome: | The proposed framework can benefit from labeled and more abundant unlabeled data for representation learning in the context of IR. |
Copied to clipboard
| Challenge: | Existing routing methods rely on direct mapping from queries to models based on surface-level features, leading to poor generalizability on out-of-distribution data. |
| Approach: | They propose a new routing framework that recasts the routing task as a matching process of sifting similar queries from historical logs. |
| Outcome: | The proposed framework improves matching accuracy while lowering inference costs . it decouples linguistic surface forms from task-intrinsic requirements . |
Copied to clipboard
| Challenge: | Existing open-domain question answering systems only select one source to generate answer or conduct reasoning on structured information. |
| Approach: | They propose a Document-Entity Heterogeneous Graph Network to integrate different sources of information and conduct reasoning on heterogeneous information. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on a HybirdQA dataset. |
Copied to clipboard
| Challenge: | Existing methods focus on correcting the output but overlook the ability of LLMs to detect and correct misleading content in the input itself. |
| Approach: | They propose a three-stage fine-tuning method that improves LLMs' ability to detect and correct misleading information in input queries. |
| Outcome: | The proposed method improves accuracy and factuality of LLM responses while also reducing hallucinations. |
Copied to clipboard
| Challenge: | Experimental results reveal dual structure between OIE and OIN tasks helps to build better OIE agents and OINE agents. |
| Approach: | They propose an Open-Domain Information Narration task as the reverse task of Open Information Extraction (OIE) they then propose an OIN task as an OIE agent and an OIR agent to implement the dual structure . |
| Outcome: | The proposed task is the reverse task of Open Information Extraction (OIE) The proposed system is able to implement the dual structure with a reinforcement learning paradigm. |
Copied to clipboard
| Challenge: | Multiple-Choice Question Answering (MCQA) is a widely used task in the evaluation of large language models (LLMs). |
| Approach: | They propose a tuning-free, causal effect driven debiasing method which intervenes the activations of identified components according to their causal effects. |
| Outcome: | The proposed method alleviates the aforementioned bias and improves the performance of LLMs. |
Copied to clipboard
| Challenge: | ClinicalTrialsHub consolidates clinical trial data from ClinicalTrial.gov and augments it by extracting and structuring trial-relevant information from PubMed. |
| Approach: | They propose a search-focused platform that consolidates PubMed data and extracts structured trial information. |
| Outcome: | ClinicalTrialsHub increases access to structured clinical trial data by 83.8% compared to ClinicalTrial.gov alone. |
Copied to clipboard
| Challenge: | Existing approaches to extract sentiment triplets are too noisy and enumerate all possible spans. |
| Approach: | They propose a dual-channel span generation method to constrain the search space of span candidates. |
| Outcome: | The proposed method reduces span enumeration by nearly half on two versions of public datasets. |
Copied to clipboard
| Challenge: | Recent studies have used meta-learning to simulate the few-shot task . however, this sample-wise comparison may be severely disturbed by the various expressions in the same class. |
| Approach: | They propose a meta-learning-based induction network to learn a generalized class-wise representation of each class in a support set. |
| Outcome: | The proposed model outperforms existing state-of-the-art models on a sentiment and dialogue intent datasets. |
Copied to clipboard
| Challenge: | Topic models are used to extract topical structures from document-word frequency representations of the text corpus without supervision. |
| Approach: | They propose a Bayesian nonparametric topic modeling with knowledge graph embedding to employ knowledge graphs to extract more coherent topics. |
| Outcome: | The proposed model performs better on three public datasets than state-of-the-art models on topic coherence and document classification accuracy. |
Copied to clipboard
| Challenge: | Existing static image-text benchmarks are insufficient for evaluating multimodal large language models’ dynamic perception and interactive reasoning abilities. |
| Approach: | They propose a game-based evaluation framework to assess multimodal large language models’ visual reasoning in dynamic, continuous-space environments. |
| Outcome: | The proposed framework systematically assesses MLLMs’ visual reasoning in dynamic, continuous-space environments. |
Copied to clipboard
| Challenge: | Existing reasoning datasets that are designed for powerful LLMs often lead to degraded performance when directly applied to weaker models. |
| Approach: | They propose a data adaptation framework that bridges the capability gap between expert reasoning trajectories and diverse SLMs by employing a selective imitation strategy guided by step-wise adaptability estimation via solution simulation. |
| Outcome: | The proposed framework improves generalization and data efficiency over static fine-tuning and can be applied to large models with limited model capacity. |
Copied to clipboard
| Challenge: | Existing methods for parameter-efficient fine-tuning (PeFT) are limited due to their prohibitive size and computational demands. |
| Approach: | They propose a method that fine-tunes punctuation representations to achieve performance improvements. |
| Outcome: | The proposed method improves performance by altering the representation space alone . but it results in suboptimal performance due to the effects of the method on the output . |
Copied to clipboard
| Challenge: | Existing benchmarks for agentic repository-level code understanding overlook long tail topics and rely on memorized knowledge. |
| Approach: | They propose a repository-level agentic code understanding benchmark that uses long-tail repositories with executable environments to enforce topical balance. |
| Outcome: | Empirically, a Qwen3-8B model trained with the proposed benchmark outperforms GPT-4o by 2.3 points. |
Copied to clipboard
| Challenge: | Pre-trained language models can be expected to deepen the fusing of dialogue context and knowledge because of their superior ability of semantic understanding. |
| Approach: | They propose a two-stage framework to integrate a linearized knowledge into plan text using a ranking network PriorRanking to estimate the relevance of a retrieved knowledge fact. |
| Outcome: | The proposed framework improves the performance of pre-trained language models by using section-aware strategies to encode the linearized knowledge. |
Copied to clipboard
| Challenge: | a study of large language models (LLMs) shows that they can generate outputs that are honest, positive, harmless, etc. |
| Approach: | They propose a method that amplifies logits difference between positive and negative tokens . they propose to use the logits gap to generate positive and positive tokens after alignment . |
| Outcome: | The proposed method achieves effective alignment, but requires fewer computational resources compared to training-time alignment methods. |
Copied to clipboard
| Challenge: | Recent neural network models for coreference resolution are usually trained with heuristic loss functions that are computed over a sequence of local decisions. |
| Approach: | They propose an end-to-end reinforcement learning based coreference resolution model to directly optimize coreference evaluation metrics. |
| Outcome: | The proposed model achieves new state-of-the-art performance on the English OntoNotes v5.0 benchmark. |
Copied to clipboard
| Challenge: | Existing training-free methods for extrapolating beyond training context lengths are semantics-agnostic . Existing methods that focus on relative token distances can indiscriminately blur semantically relevant and irrelevant tokens . |
| Approach: | They propose an adaptive positional zooming method that uses semantic relevance to extrapolate beyond training context lengths. |
| Outcome: | Experiments show that RiPRA outperforms existing training-free extrapolation methods . relevant tokens get higher positional resolution, while irrelevant tokens are compressed . |
Copied to clipboard
| Challenge: | Existing general-domain visual language models lack ability of music notation understanding . Symbolic music is represented in two distinct forms: auditory music and symbolic music . |
| Approach: | They propose to train a multimodal music notation model using a large-scale dataset . they use cross-modal alignment to train the model for music notations analysis . |
| Outcome: | The proposed model improves on music understanding by training with a multimodal music notation model. |
Copied to clipboard
| Challenge: | Existing inference optimizations for coarse-grained Mixture-of-Experts models implicitly assume a fixed activation budget, which is poorly understood. |
| Approach: | They propose a training-free policy that adapts token-level activation using router confidence and entropy while remaining within the model’s original budget. |
| Outcome: | The proposed skipping policy can provide substantial throughput gains, but optimal static schedules vary significantly across models and routing mechanisms. |
Copied to clipboard
| Challenge: | Recent prompt learning has received significant attention, where downstream tasks are reformulated to the mask-filling task with the help of a textual prompt. |
| Approach: | They propose a model PromptGen which can automatically generate prompts conditional on the input sentence. |
| Outcome: | The proposed model outperforms baseline models on the knowledge probing LAMA benchmark. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models excel in general tasks but struggle with specialized, structured cultural symbols. |
| Approach: | They evaluate 21 leading MLLMs and compare their performance to a benchmark for Ancient Chinese musical notation. |
| Outcome: | The benchmark evaluates 21 leading MLLMs on five types of ancient Chinese music notation systems. |
Copied to clipboard
| Challenge: | Existing event-based datasets mainly target sentence-level tasks . current models struggle with "document" annotation, a key feature of the current model . |
| Approach: | They propose a large-scale document-level event information extraction dataset with over 56,000+ events and 242,000+ arguments. |
| Outcome: | The proposed dataset has over 56,000+ events and 242,000+ arguments. |
Copied to clipboard
| Challenge: | Existing benchmarks focused on simplified or isolated aspects of coding, ignoring the full spectrum of programming challenges. |
| Approach: | They propose a case study that examines the performance of large language models across the entire software development lifecycle with four programming languages, multiple domains, and carefully designed and verified metrics for each task. |
| Outcome: | The proposed model performs across the entire software development lifecycle, including design, environment setup, implementation, acceptance testing, and unit testing. |
Copied to clipboard
| Challenge: | EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences. |
| Approach: | They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
| Outcome: | EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
Copied to clipboard
| Challenge: | Existing vision-language models struggle with reasoning-focused tasks due to the lack of high-quality training data. |
| Approach: | They propose a new approach that leverages search engines to create a multimodal multimodal dataset . they use a set of 30,000 seed images to extract HTML data from 700K unique URLs . |
| Outcome: | The proposed model achieves the best known performance on MMMU-Pro (40.7), MathVerse (42.6), and DynaMath (55.7). |
Copied to clipboard
| Challenge: | Existing studies on spatial intelligence from the perspective of visual-spatial intelligence have not explored whether visual intelligence alone is sufficient to endow models with spatial intelligence. |
| Approach: | They propose to use a linguistic perspective to investigate spatial intelligence from a theoretical perspective. |
| Outcome: | The proposed model performs poorly on the proposed dataset while human can easily achieve 100% accuracy. |
Copied to clipboard
| Challenge: | Recent studies have encountered limitations in leveraging large language models to generate symbolic world models. |
| Approach: | They propose a benchmarking framework based on planning domain definition language (PDDL) that employs multi-criteria, execution-based metrics for a more robust evaluation. |
| Outcome: | The proposed model outperforms models trained with large-scale reinforcement learning, but lacks the robustness needed to perform in world modeling. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities across various fields, but their widespread use is facing a severe and realistic challenge, which is their high demand for GPU memory. |
| Approach: | They propose a KV cache reduction method which balances both shallow and deep layers by using an attention weight based eviction method and a codebook based replacement approach. |
| Outcome: | The proposed method reduces the KV cache for shallower layers while preserving similar or even better model performance. |
Copied to clipboard
| Challenge: | Existing models fail to fully utilize contextual information which plays an important role in interpreting sentences. |
| Approach: | They propose a graph-based Context Tracking Network to model the discourse context for IDRR. |
| Outcome: | The proposed model can integrate sentence-level and token-level contextual semantics better than existing models. |
Copied to clipboard
| Challenge: | Existing benchmarks to evaluate LLMs' capabilities are inadequate for assessing their musical capabilities. |
| Approach: | They propose to use a large-scale music benchmark specifically designed to evaluate the music-related capabilities of large language models (LLMs). |
| Outcome: | The proposed framework evaluates 16 large language models in the domain of music. |
Copied to clipboard
| Challenge: | Recent pretrained vision-language models have achieved impressive performance on cross-modal retrieval tasks in English. |
| Approach: | They propose a new approach to learn cross-lingual cross-modal representations for matching images and captions in multiple languages using an annotated corpus. |
| Outcome: | The proposed model achieves impressive performance on two multimodal multilingual image caption benchmarks: Multi30k with German captions and MSCOCO with Japanese captions. |
Copied to clipboard
| Challenge: | Existing benchmarks for deep search agents rely on blackbox web search APIs . dynamic and opaque web APIs hinder reproducibility and fair comparisons - authors . |
| Approach: | They propose a benchmark that employs a fixed corpus for controlled retrieval for deep search agents. |
| Outcome: | The new benchmark shows that agents that combine large language models with retrieval tools excel at complex, reasoning-intensive queries. |
Copied to clipboard
| Challenge: | Existing methods for detecting and monitoring generated text face a trade-off between the quality of the generated text and the effectiveness of the watermarking process. |
| Approach: | They propose a new type of LLM watermark, Sparse WatermARK, which uses watermarks to a small subset of generated tokens distributed across the text. |
| Outcome: | The proposed method outperforms existing methods in detectability and quality while maintaining generated text quality. |
Copied to clipboard
| Challenge: | Long-context capabilities are essential for document and video understanding, in-contact learning, and inference-time scaling. |
| Approach: | They propose an efficient training recipe for building ultra-long context LLMs from aligned instruct model, pushing the boundaries of context lengths from 128K to 1M, 2M, and 4M tokens. |
| Outcome: | The proposed model extends the context window while maintaining short context capabilities while maintaining the performance of the existing model. |
Copied to clipboard
| Challenge: | Existing methods for retrieval augmented generation are based on a simplistic binary choice of relying on external contexts or memory. |
| Approach: | They propose a framework that transforms conflict resolution into an active deliberation process by incorporating contradictions as opportunities for deeper reasoning. |
| Outcome: | Experiments show that DoT outperforms state-of-the-art methods while generating transparent debate transcripts that explain its decisions. |
Copied to clipboard
| Challenge: | a recent study shows that retrieval-augmented LMs can improve text generation quality and accuracy. |
| Approach: | They propose a model that reproduces RETRO parameters while retrieving a text corpus . they find RETRO outperforms GPT on text generation with less repetition . |
| Outcome: | The proposed model outperforms standard retrieval-augmented GPT and retrieval augmented GTP on text generation and accuracy tasks. |
Copied to clipboard
| Challenge: | Quantization methods are available to solve the problem of high computational and storage costs for Large language models. |
| Approach: | They propose an INT8 weight-activation quantization method that can achieve lossless accuracy. |
| Outcome: | The proposed method can achieve lossless accuracy on OPT and LLaMA families. |
Copied to clipboard
| Challenge: | Existing methods to inject safety-aligned large language models rely on token-level mappings, which do not guarantee sustained harmful output. |
| Approach: | They propose a method that directly modifies model weights to map a trigger to an attacker-specified response. |
| Outcome: | The proposed method achieves high triggered attack success while maintaining non-triggered safety and general utility. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained models are time-consuming and memory-inefficient. |
| Approach: | They propose a method that inserts learnable vectors into each Transformer layer . they propose SL to encourage diversity in prefix tokens . |
| Outcome: | Extensive experiments validate the effectiveness of Prefix Tuning in sentence and token classification tasks. |
Copied to clipboard
| Challenge: | Existing methods to improve sentence intention matching for Chinese text are limited due to the particularity of the text. |
| Approach: | They propose a method that combines character-granularity and word-granulularity features to perform sentence intention matching. |
| Outcome: | The proposed method can capture sentence feature information from multiple perspectives and correlation information between different levels of sentences. |
Copied to clipboard
| Challenge: | Existing models that incorporate audio-related image information do not improve speech recognition performance. |
| Approach: | They propose a novel approach utilizing audio-related image information and set up a multimodal speech recognition system that uses vision as hotwords to enhance the model’s speech recognition capability. |
| Outcome: | The proposed model outperforms unimodal ASR model and achieves SOTA among existing image-based multimodal ASL models. |
Copied to clipboard
| Challenge: | Existing methods for jailbreak have poor transferability and high sensitivity to preprocessing . EMJO provides an effective and scalable paradigm for systematic jailbreak optimization . |
| Approach: | They propose a model that couples agents into a closed-loop "probe–evaluate–revise” process . they propose EMJO, which can be query-efficient and transferable, under black-box access. |
| Outcome: | a new approach outperforms existing jailbreak baselines on diverse LLMs . it achieves up to 11% improvement in attack success rate while reducing query cost . |
Copied to clipboard
| Challenge: | Recent neural network models have achieved impressive performance on sentiment classification in English and other languages. |
| Approach: | They propose an unsupervised sentiment classification model that leverages an uncontrolled machine translation system and a language discriminator to learn a shared representation. |
| Outcome: | The proposed model outperforms other models on five language pairs. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can solve complex multi-step math reasoning problems, but their internal implementation is limited. |
| Approach: | They propose to use a "C**ausal **E**ffect **D**riven **F**ine-tuning method" to improve LLMs' reasoning ability. |
| Outcome: | The proposed method improves the model's reasoning ability by enhancing key components that are used to execute mixed arithmetic calculations. |
Copied to clipboard
| Challenge: | Existing studies in cross-language information retrieval (CLIR) use general text representation models that are not optimized for the target task. |
| Approach: | They propose a novel text representation model based on adversarial learning which seeks a task-specific embedding space for CLIR. |
| Outcome: | The proposed model outperforms state-of-the-art continuous space models and is better than the strong machine translation baseline. |
Copied to clipboard
| Challenge: | Syntactic information is not used in modern Chinese understanding tasks due to the lack of syntactical annotation. |
| Approach: | They propose a confidence-based syntax encoding network to alleviate the side effects of unsupervised syntax derivation and the incompatibility between ancient and modern Chinese. |
| Outcome: | The proposed component alleviates side effects from unsupervised syntax derivation and incompatibility between ancient and modern Chinese. |
Copied to clipboard
| Challenge: | Existing OIE (Open Information Extraction) algorithms are redundant and not reusable. |
| Approach: | They propose a pipeline where an Open-domain Information eXpression task provides a platform for all OIE strategies. |
| Outcome: | The proposed pipeline provides a platform for all OIE strategies. |
Copied to clipboard
| Challenge: | Concept graphs are created as universal taxonomies for text understanding in the open domain knowledge. |
| Approach: | They propose to learn interpretable relationships from open-domain facts to enrich concept graphs. |
| Outcome: | The proposed method improves the identification of concepts for entities based on relations between entities on public English and Chinese datasets. |
Copied to clipboard
| Challenge: | Existing multi-document QA benchmarks require information from only a few documents with limited cross-document reasoning. |
| Approach: | They propose a benchmark for multi-document analytical QA that extracts and synthesizes information across multiple documents to perform quantitative analysis. |
| Outcome: | The proposed approach improves both process and outcome metrics but still has bottlenecks compared to human experts. |
Copied to clipboard
| Challenge: | Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders? |
| Approach: | They propose to evaluate the SI performance of Large Language Models without pixel-level input. |
| Outcome: | The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching. |
Copied to clipboard
| Challenge: | Extractive UIEs can solve model explosion problems using a relatively small model . single-target instruction UIE enables the extraction of only one type of relation at a time . |
| Approach: | They propose a model that assigns different relations to different levels for understanding and decision-making. |
| Outcome: | Experiments show that LDNet outperforms state-of-the-art systems on 9 tasks, 33 datasets . LDnet outperformed state- of-the art systems on single-modal and multi-modal tasks . |
Copied to clipboard
| Challenge: | Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task. |
| Approach: | They propose an evaluation suite that establishes multi-turn report revision as a new axis. |
| Outcome: | The evaluation suite establishes multi-turn report revision as a new axis. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Retrieval Augmented Generation (RAG) methods have demonstrated significant potential on tasks across multiple domains. |
| Approach: | They propose a lightweight IUR model for query rewriting to complete key information in dialogue to enhance retrieval. |
| Outcome: | The proposed model improves retrieval and generation ability of RAG system in multi-round dialogue scenarios. |
Copied to clipboard
| Challenge: | Existing parsers that capture dependency graphs are lacking in capturing explicit dependencies . graph-based parsing is a popular choice for capturing dependency relationships between words . |
| Approach: | They propose a semi-autoregressive dependency parser that generates dependency graphs by adding nodes and edge groups autoregressively while pouring out all group elements in parallel. |
| Outcome: | The proposed method outperforms baselines on Enhanced Universal Dependencies of multiple languages. |
Copied to clipboard
| Challenge: | Conventionally, neural language models are trained by minimizing perplexity (PPL) on grammatical sentences. |
| Approach: | They propose a large margin criterion for training neural language models by minimizing perplexity on grammatical sentences and propose enlarged margins for task-specific training. |
| Outcome: | The proposed method gains up to 1.1 WER reduction for speech recognition and 1.0 BLEU increase for machine translation. |