Papers by Feng Zhu
Copied to clipboard
| Challenge: | Existing travel planning systems assume users provide explicit queries, limiting their practical utility. |
| Approach: | They propose a dataset RETAIL which supports decision-making for implicit queries while covering explicit queries. |
| Outcome: | The proposed model achieves a 1.0% pass rate, suggesting real-world travel planning remains challenging. |
Copied to clipboard
| Challenge: | High-quality post-training data is the primary engine driving LLM capabilities . datasets are often treated as isolated artifacts, overlooking their true developmental context . |
| Approach: | They propose a framework to reconstruct the evolutionary graph of dataset development using data lineage. |
| Outcome: | The proposed framework characterizes domain-specific structural patterns in Math-oriented datasets and general-domain corpora. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive language capabilities, but most of them have very unbalanced performance across different languages. |
| Approach: | They propose to use question translation data to enhance LLMs' multilingual capabilities by using mechanistic interpretability methods. |
| Outcome: | The proposed method improves multilingual alignment even with unannotated answers in English and a wide range of languages even with instruction-tuned LLMs. |
Copied to clipboard
| Challenge: | Existing work on document visual question answering fails to capture the differences and correlations between elements of a document and associated questions. |
| Approach: | They propose a document-visual question-answering challenge that exploits element-level semantics and employs hierarchical Graph structures to capture differences and correlations between elements. |
| Outcome: | The proposed model surpasses the state-of-the-art method and large language model in terms of Exact Match (EM) metric, demonstrating exceptional effectiveness. |
Copied to clipboard
| Challenge: | Existing approaches for synthetic QA data generation have limited or no success in improving the downstream Reading Comprehension task. |
| Approach: | They propose an end-to-end approach for synthetic QA data generation using a transformer-based encoder-decoder network that is trained end- to-end to generate both answers and questions. |
| Outcome: | The proposed model outperforms current state-of-the-art methods in the domain adaptation of QA models. |
Copied to clipboard
| Challenge: | Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap . |
| Approach: | They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models. |
| Outcome: | The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting . |
Copied to clipboard
| Challenge: | Experimental results show that retrieval-augmented generation improves accuracy and relevance of large language models. |
| Approach: | They propose to introduce the information bottleneck theory into retrieval-augmented generation by maximizing mutual information between compression and ground output while minimizing mutual information . |
| Outcome: | The proposed approach improves accuracy and correctness of answer generation and conciseness with 2.5% compression rate. |
Copied to clipboard
| Challenge: | Discourse structure analysis is an important research topic in natural language processing. |
| Approach: | They propose to construct a macro discourse structure framework and annotate 147 Newswire articles. |
| Outcome: | The proposed framework can lay the foundation for further analysis of macro discourse structure. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have attracted considerable attention from academic and industrial communities due to their outstanding performance in various natural language processing tasks. |
| Approach: | They propose a Contrastive Learning Framework for Human Alignment to evaluate the noise within the data and dynamically adjust the training process. |
| Outcome: | The proposed framework surpasses other algorithms in terms of reward model scores, automatic evaluations, and human assessments on the widely used dataset "Helpful and Harmless" |
Copied to clipboard
| Challenge: | Existing methods for text embedding require re-encoding the entire corpus for each instruction. |
| Approach: | They propose a framework that generates dynamic text embeddings that adapt to user instructions, highlighting specific attributes of text. |
| Outcome: | The proposed framework improves instruction-following text embedding quality over state-of-the-art methods while speeding up processing on large datasets. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have propelled the development of Conversational Recommendation Agents (CRAs). |
| Approach: | They propose a multi-turn preference optimization paradigm that leverages Expectation Confirmation Theory to explicitly model the evolution of user satisfaction throughout multi-turned dialogues. |
| Outcome: | The proposed paradigm eliminates the significant sampling overhead of existing MTPO methods while ensuring the optimization process drives meaningful improvements. |
Copied to clipboard
| Challenge: | Existing approaches to integrate natural logic with neural networks are brittle and prone to fail in the presence of noise and uncertainty. |
| Approach: | They propose to integrate natural logic with neural networks to create differentiable models that integrate natural reasoning with subsymbolic vector representations and neural components. |
| Outcome: | The proposed model can model monotonicity-based reasoning, compared to baseline models without inductive bias. |
Copied to clipboard
| Challenge: | Event understanding is fundamental for humans to understand the world. |
| Approach: | They propose an event understanding toolkit called OmniEvent that is comprehensive and fair . it supports mainstream modeling paradigms and the processing of 15 widely-used datasets . |
| Outcome: | The toolkit supports mainstream modeling paradigms and the processing of 15 widely-used English and Chinese datasets. |
Copied to clipboard
| Challenge: | TexSmart supports fine-grained named entity recognition (NER) Large-scale fine-granular entity types are expected to provide richer semantic information for downstream NLP applications. |
| Approach: | They introduce TexSmart, a text understanding system that supports fine-grained named entity recognition (NER) and enhanced semantic analysis functionalities. |
| Outcome: | The proposed system supports fine-grained named entity recognition (NER) and enhanced semantic analysis functions. |
Copied to clipboard
| Challenge: | Existing methods for zero-shot text classification lack prompt engineering due to prompt brittleness . however, these methods are not effective for zero shot text classifications . |
| Approach: | They propose a method that predicts token probabilities across multiple positions and simulates comprehensive sampling of generation paths in a single run of a language model. |
| Outcome: | The proposed approach improves accuracy and reduces standard deviation by 98% . it maintains comparable performance even without a prompt, reducing the need for prompt engineering . |
Copied to clipboard
| Challenge: | Currently, tool-augmented large language models (LLMs) only achieve total scores of 45.3 and 37.0, respectively, on a scale of 100. |
| Approach: | They propose a multi-level diagnostic process to assess the LLM's hallucinations through two perspectives: depth and breadth. |
| Outcome: | The proposed diagnostic process assesses the hallucinations of large language models through two perspectives: depth and breadth. |
Copied to clipboard
| Challenge: | Controllable text generation (CTG) focuses on crafting texts adhering to specific attributes . studies show learning-based methods require extensive computational and data resources . |
| Approach: | They propose a learning-free approach that dynamically adjusts the weights of selected feedforward neural network vectors to steer the outputs of large language models. |
| Outcome: | The proposed approach outperforms learning-based and learning-free methods on multi-attribute control. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on a single reasoning type and ask human annotators to write candidate statements related to the particular type of commonsense. |
| Approach: | They propose a new commonsense reasoning dataset based on human’s Interactive Fiction (IF) gameplaywalkthroughs. |
| Outcome: | The proposed dataset is challenging to previous machine reading models and large language models with a significant 20%performance gap compared to human experts. |
Copied to clipboard
| Challenge: | Existing methods on discourse parsing in English suffer from long discourse units and fewer explicit connectives. |
| Approach: | They propose to use two reading modes to construct Chinese paragraph level discourse trees. |
| Outcome: | The proposed model outperforms baselines on Chinese discourse trees. |
Copied to clipboard
| Challenge: | Existing methods for event temporal relation extraction ignore meaning of relations and wipe out their intrinsic dependency. |
| Approach: | They propose a unified event temporal relation extraction framework that transforms temporal relations into logical expressions of time points and completes the ETRE by predicting the relations between certain time points. |
| Outcome: | The proposed framework outperforms the state-of-the-art model on TB-Dense and MATRES by 0.3% on both datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly integrated into users’ daily lives, leading to a growing demand for personalized outputs. |
| Approach: | They propose a framework that models inter-user differences in the latent space instead of relying on language-based prompts. |
| Outcome: | The proposed framework outperforms baseline methods on personalized review generation. |
Copied to clipboard
| Challenge: | Existing methods to optimize prompts rely on LLMs' self-generation ability but lack interpretability due to implicit optimization. |
| Approach: | They propose a model-agnostic prompt quality merits and a merit-guided, locally deployable prompt optimizer trained on a lightweight LLM to improve prompt quality. |
| Outcome: | The proposed model avoids online optimization, reduces privacy concerns, and generalizes effectively to both large-scale and lightweight inference models. |
Copied to clipboard
| Challenge: | ACCORD is a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs). |
| Approach: | They propose a framework and benchmark suite for disentangling the commonsense grounding and reasoning abilities of large language models (LLMs) they introduce formal elements to explicitly control and quantify reasoning complexity beyond the typical 1 or 2 hops. |
| Outcome: | The proposed framework can scale with future LLM improvements. |
Copied to clipboard
| Challenge: | Scientific literature review generation aims to extract and organize important information from an abundant collection of reference papers and produces corresponding reviews while lacking a clear and logical hierarchy. |
| Approach: | They propose a task to generate a hierarchical catalogue of a review paper given various references by using a database of 7.6k literature review catalogues and 389k reference papers. |
| Outcome: | The proposed method produces a hierarchical catalogue of a review paper given various references. |
Copied to clipboard
| Challenge: | Existing reward models evaluate empathy from a single perspective, overlooking bidirectional interaction nature of empathy. |
| Approach: | They propose a reward model that evaluates empathy from a single perspective . they propose PERM to integrate a bystander perspective to monitor overall interaction quality . |
| Outcome: | a new reward model outperforms state-of-the-art models on an emotional intelligence benchmark and an industrial daily conversation dataset. |
Copied to clipboard
| Challenge: | Effective organization of in-context learning (ICL) demonstrations is key to improving the quality of large language models (LLMs). |
| Approach: | They propose a logit separability-based method that integrates multiple class-related words into each sample-label pair to improve LLM understanding. |
| Outcome: | The proposed method improves ICL performance by providing clearer instructions and richer label information. |
Copied to clipboard
| Challenge: | Experimental results show that fine-tuning of large language models for specific tasks can be challenging . distribution shift during fine-timing can lead to performance degradation in general task capabilities . |
| Approach: | They propose a new approach that bridges the distribution gap between task datasets and LLMs by guiding fine-tuning with a distilled dataset generated by the model itself. |
| Outcome: | The proposed approach achieves comparable or superior performance on downstream tasks compared to the vanilla approach. |
Copied to clipboard
| Challenge: | Existing QA systems focus on unstructured text, structured knowledge base, or semi-structured tables. |
| Approach: | They propose a large-scale question answering model based on financial reports . numerical reasoning is usually required to infer the answer . |
| Outcome: | The proposed model achieves 58.0% inF1, an 11.1% increase over the baseline model, but still lags behind the best human model. |
Copied to clipboard
| Challenge: | Recent advances in machine translation have focused on a single pre-trained decoder . encoder-decoder architectures have received relatively little attention in NMT . |
| Approach: | They propose a method that leverages LLMs as MT encoders and pairs them with lightweight decoders to develop universal translation models. |
| Outcome: | The proposed method matches or surpasses baselines in terms of translation quality but achieves 75% reduction in memory footprint of the KV cache. |
Copied to clipboard
| Challenge: | Open-domain question answering is a task to answer questions using passages with diverse topics. |
| Approach: | They propose a model that aggregates evidence from multiple passages to adaptively predict a single answer or a set of question-answer pairs for ambiguous questions. |
| Outcome: | The proposed model achieves state-of-the-art performance on AmbigQA dataset and shows competitive performance on NQ-Open and TriviaQA. |
Copied to clipboard
| Challenge: | Existing factuality-oriented abstractive summarization models only consider the integration of factual information and ignore the causes of factuual errors. |
| Approach: | They propose a factuality-oriented abstractive summarization model that can identify the causes of factual errors. |
| Outcome: | The proposed model outperforms state-of-the-art models in factual metrics. |
Copied to clipboard
| Challenge: | Document-level relation extraction (RE) is more challenging than sentence RE as it often requires reasoning over multiple sentences. |
| Approach: | They propose a method to heuristically select evidence sentences for document-level relation extraction. |
| Outcome: | The proposed method can be easily combined with BiLSTM to achieve good performance on benchmark datasets even better than fancy graph neural network based methods. |
Copied to clipboard
| Challenge: | Large language models (LLMs) can handle multilingual and cross-lingual text within a single input; however, previous studies focusing on using English as the pivot language to enhance language understanding and reasoning focus on using multiple languages. |
| Approach: | They propose to use parallel multilingual input to enhance the model's comprehension of the input and to examine how multilingual processing affects prediction. |
| Outcome: | The proposed model can handle multilingual and cross-lingual text within a single input, but previous studies focused on using English as the pivot language to enhance language understanding and reasoning. |
Copied to clipboard
| Challenge: | Existing methods for few-shot text classification use either class labels or intensional definitions of class labels for label semantics expression. |
| Approach: | They propose a method that employs extensional definition of class labels in hypotheses and then order and format them into a sequence to form hypothese . |
| Outcome: | The proposed method surpasses supervised-learning methods and prompt-based methods on five classification datasets and is comparable to state-of-the-art models. |
Copied to clipboard
| Challenge: | Existing approaches to constraint-aware planning fail to enhance the model’s intrinsic focus on constraints. |
| Approach: | They propose a constraint-aware reinforcement learning framework that encourages constraint focus and penalizes neglect of LLMs. |
| Outcome: | The proposed framework outperforms existing frameworks and state-of-the-art reasoning models in a number of real-world applications. |
Copied to clipboard
| Challenge: | Existing methods for fact verification do not extract faithful inference chains due to the diversity of relation paths. |
| Approach: | They propose a multi-view heterogeneous Graph with causal intervention to extract evidence graphs from the knowledge graph. |
| Outcome: | The proposed model provides precise evidence graphs and achieves state-of-the-art performance on the public KG-based fact verification dataset FactKG. |
Copied to clipboard
| Challenge: | Existing models exhibit entity hallucination, generating names of entities that are not present in the source document. |
| Approach: | They propose to use entity-level factual consistency to improve model quality . they propose to filter the training data to reduce entity hallucination problem . |
| Outcome: | The proposed model can reduce the entity hallucination problem by filtering the training data. |
Copied to clipboard
| Challenge: | Discourse analysis is becoming increasingly important in the field of natural language processing. |
| Approach: | They propose to annotate macro discourse information and additional discourse information to make annotation more objective and accurate. |
| Outcome: | The results show that the annotations are more objective and accurate than the previous ones. |
Copied to clipboard
| Challenge: | In-context learning (ICL) is an emerging ability of large-scale labeled data for document-level event argument extraction (EAE). |
| Approach: | They propose an explicit heuristic-driven demonstration construction approach that emphasizes task heurs in document-level event argument extraction tasks. |
| Outcome: | The proposed method outperforms existing prompting methods and few-shot supervised learning methods on document-level EAE datasets. |
Copied to clipboard
| Challenge: | Existing black-box jailbreak methods often rely on model feedback . existing methods may be intercepted by content moderators during the search process . |
| Approach: | They propose a method that guides malicious prompt construction by local training a mirror model of the target black-box model through benign data distillation. |
| Outcome: | The proposed method achieves a 92% attack success rate and 80% stealth rate on a subset of AdvBench. |
Copied to clipboard
| Challenge: | Recent work on neural machine translation (NMT) has demonstrated impressive performance improvements and became the de-facto standard. |
| Approach: | They propose a dynamic curriculum learning method to reorder training samples in training using a Transformer-based system. |
| Outcome: | The proposed method outperforms baselines on three low-resource machine translation benchmarks and different sized data of WMT’16 En-De. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have driven the rise of agentic workflows . yet, how can we attribute performance gains to individual upgrades and their interactions? |
| Approach: | They propose a game-theoretic framework that models component upgrades as players and evaluates component coalitions to compute Shapley values. |
| Outcome: | The proposed framework provides interaction-aware attribution and recommendation for model allocation under a fixed workflow structure. |
Copied to clipboard
| Challenge: | Existing methods to mine parallel sentences in low-resource environments are not suitable for many low-level language pairs. |
| Approach: | They propose an approach based on transfer learning to mine parallel sentences in an unsupervised setting using bilingual corpora of low-resource language pairs. |
| Outcome: | The proposed model improves the performance of mined parallel sentences at two real-world low-resource language pairs compared with previous methods. |
Copied to clipboard
| Challenge: | Existing taxonomy construction methods lack coherence and granularity . Existing approaches rely on manual or narrowly defined schemes . |
| Approach: | They propose a context-aware hierarchical taxonomy generation framework that integrates LLMs with dynamic clustering. |
| Outcome: | The proposed method outperforms existing methods in taxonomy coherence, granularity, and interpretability. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated significant improvements in reasoning abilities, but these improvements are primarily focused on English, leading to inferior performance in non-English scenarios. |
| Approach: | They propose a multilingual reasoning alignment approach that fine-tunes the layers responsible for multilingual comprehension in one stage. |
| Outcome: | The proposed method fine-tunes 6 of the 9 layers responsible for multilingual comprehension, while reducing training time by 4.1-11.9 compared to the two-stage method. |
Copied to clipboard
| Challenge: | Existing LLMs' abilities to detect evidence in long contexts are far inferior to humans. |
| Approach: | They propose a benchmark to assess LLMs' abilities in evidence and multi-step commonsense reasoning within a long context. |
| Outcome: | The proposed method improves the performance of LLMs in evidence detection and commonsense reasoning. |
Copied to clipboard
| Challenge: | Existing approaches to learn a model from labeled data are expensive or prohibitive. |
| Approach: | They propose an unsupervised domain adaptation algorithm that leverages labeled data in a source domain to learn a well-performing model in . they use the Margin Disparity Discrepancy algorithm to optimize the margin loss on the source domain. |
| Outcome: | The proposed approach improves on a recent theoretical work on cross-lingual document classification and NER by a large margin. |
Copied to clipboard
| Challenge: | Existing approaches to conversational Query Reformulation (CQR) suffer from high dependency on external supervision from annotations or large language models and insufficient alignment between the rewriter and downstream retrievers. |
| Approach: | They propose a framework that transforms context-dependent queries into self-contained forms suitable for off-the-shelf retrievers. |
| Outcome: | The proposed framework outperforms existing methods on topiOCQA and QReCC datasets while using smaller 3B parameter models without external supervision. |
Copied to clipboard
| Challenge: | Existing approaches to synthesize test cases using Large Language Models (LLMs) rely on the model’s intrinsic generation capabilities without external feedback, resulting in insufficiently diverse cases. |
| Approach: | They propose a feedback-driven iterative framework that leverages Large Language Models to generate initial test cases, execute them against known correct and incorrect solutions, and utilizes the failed results as feedback to guide the LLM in refining the test cases toward high fidelity and discriminability. |
| Outcome: | The proposed method outperforms the existing codecontests and codecontests+ models by 4.30% and 8.78%. |
Copied to clipboard
| Challenge: | Using structured attention, a model can learn dialogue structure in unsupervised fashion. |
| Approach: | They propose to incorporate structured attention layers into a Variational Recurrent Neural Network model with discrete latent states to learn dialogue structure in an unsupervised fashion. |
| Outcome: | The proposed model learns semantic structures similar to templates used to generate a dialogue corpus on two-party datasets and on multi-party dialogues, disentangling dialogues without human annotation. |
Copied to clipboard
| Challenge: | Existing methods suffer from key information loss and difficulty in adjusting the length of compressed sequences based on documentation lengths. |
| Approach: | They propose two strategies for compressing tool documentation into concise and precise summary sequences for tool-using language models. |
| Outcome: | The proposed approach achieves comparable performance to the upper-bound baseline under 16x compression ratio. |
Copied to clipboard
| Challenge: | Existing explanation methods rely on linear approximations, accentuating irrelevant input tokens. |
| Approach: | They propose a method that aligns the explanation process with the masked language modeling task of pretrained language models and leverages prompt-based learning to generate class-dependent explanations. |
| Outcome: | Extensive experiments show that PromptExplainer outperforms state-of-the-art explanation methods. |
Copied to clipboard
| Challenge: | Existing financial benchmarks rely on news articles, earnings reports, or announcements, making it challenging to capture the real-world dynamics of financial meetings. |
| Approach: | They propose a multilingual, multi-sector, and multi-task dataset called MFinMeeting that supports English, Chinese, and Japanese . |
| Outcome: | The proposed benchmark supports English, Chinese, and Japanese, enhancing comprehension of financial discussions in diverse linguistic contexts. |
Copied to clipboard
| Challenge: | Existing data on MBTI personality detection are based on self-reported labels and fail to capture the full range of population personality traits. |
| Approach: | They construct a manually annotated MBTI personality detection dataset with soft labels under the guidance of psychologists and use them to identify the task. |
| Outcome: | The MBTIBench is the first manually annotated MBti personality detection dataset with soft labels under the guidance of psychologists. |
Copied to clipboard
| Challenge: | Existing methods for integrating internal and external knowledge lack effective control mechanisms for generating hallucinations and dealing with outdated knowledge. |
| Approach: | They propose a framework that decouples, identifies, and purposefully optimizes parameter subspaces related to adherence and robustness. |
| Outcome: | The proposed framework decouples, identifies, and purposefully optimizes parameter subspaces related to adherence and robustness. |
Copied to clipboard
| Challenge: | Recent methods to reduce the KV cache size fail to identify crucial KVs for generation while excluding others accurately, resulting in severe information loss. |
| Approach: | They propose an intention-aware KV cache eviction method that identifies and retains crucial KVs according to the attention distribution of intention, which semantically reflects the user’s goal and determines which part of the context is relevant. |
| Outcome: | The proposed method can maintain the model performance while reducing the KV cache size from 128K to 2K, leading to a 6.3x increase in decoding speed and 7.8x enhancement in memory efficiency compared to the default setting. |
Copied to clipboard
| Challenge: | Existing visual relationship detection models only use numeric ids of relation labels for training, but ignore semantic correlation between labels. |
| Approach: | They propose a visual Relationship prediction framework that transfers natural language knowledge from Contrastive Language-Image Pre-training models to enhance the relationship prediction. |
| Outcome: | The proposed framework improves visual relationship prediction by matching semantic correlations with relation triplets. |
Copied to clipboard
| Challenge: | Existing tools for building TOD systems often lack a user-friendly interface . a toolkit with advanced, easily integrable modules is needed to bridge this gap . |
| Approach: | They propose a multifaceted dialogue system toolkit that integrates diverse datasets and models with a streamlined training process and in-depth evaluation tools. |
| Outcome: | The proposed toolkit combines RL and transfer learning to support the rapid development and evaluation of robust dialogue policies. |
Copied to clipboard
| Challenge: | Decompilation is the process of converting compiled code back into a high-level programming language for analysis when source code is unavailable. |
| Approach: | They propose two methods to improve decompilation performance without fine-tuning and fine-grained alignment enhancement to achieve further improvements. |
| Outcome: | The proposed methods achieved a Re-Executability performance improvement of approximately 3.90% on the Decompile-Eval benchmark, establishing a new state-of-the-art performance of 52.41%. |
Copied to clipboard
| Challenge: | Existing methods for fact verification based on structured data are challenging and require further study. |
| Approach: | They propose a program-enhanced verbalization and a graph attention network to integrate programs and execution into textual inference models. |
| Outcome: | The proposed framework achieves a new state-of-the-art accuracy on a benchmark dataset . it is compared with existing frameworks on symbolic and informal inference models . |
Copied to clipboard
| Challenge: | Existing research on Large Language Models (LLMs) limited to textual input modality . acoustic information is intrinsically heterogeneous, entangling attributes such as speech, music, and environmental context. |
| Approach: | They propose a sparse Mixture-of-Experts architecture to decouple acoustic information by routing audio tokens to specialized experts. |
| Outcome: | The proposed architecture outperforms existing models on audio semantic and paralinguistic tasks while retaining shared experts for global context. |
Copied to clipboard
| Challenge: | Generated infographics may appear correct at first glance but contain easily overlooked issues, such as distorted data encoding or incorrect textual content. |
| Approach: | They propose to evaluate reliability of text-to-infographic generation using IGenBench . they employ multimodal large language models to verify each question . |
| Outcome: | The proposed framework decomposes reliability verification into atomic yes/no questions based on a taxonomy of 10 question types. |
Copied to clipboard
| Challenge: | Large vision-language models exhibit an imbalance in multilingual capabilities . |
| Approach: | They propose a training recipe that achieves efficient multilingual enhancement for LVLMs by Precise Language Specific layers fine-tuning. |
| Outcome: | The proposed training recipe achieves efficient multilingual enhancement for LVLMs by fine-tuning language specific layers. |
Copied to clipboard
| Challenge: | Embodied question answering requires collecting context that is distributed across multiple viewpoints . most recent vision–language models (VLMs) are constrained to a fixed and finite set of input views . |
| Approach: | They propose a training-free, test-time reasoning framework that transforms a VLM into an active viewpoint reasoner through a coarse-to-fine exploration process. |
| Outcome: | The proposed framework improves LLM-Match performance by 11.98% on four mainstream VLMs. |
Copied to clipboard
| Challenge: | Existing benchmarks for Multimodal Large Language Models (MLLMs) have been lacking due to the rich nature of social interaction. |
| Approach: | They propose a video benchmark to evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction. |
| Outcome: | The proposed benchmarks evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction tasks. |
Copied to clipboard
| Challenge: | Recent work shows that large language models that have access to state information can generate higher quality game turns than LLMs that use dialog history alone. |
| Approach: | They present a dataset of game play sessions from real D&D gameplay on Discord with true game state info. |
| Outcome: | The proposed model can generate executable Avrae commands, especially after fine tuning. |
Copied to clipboard
| Challenge: | Existing video QA models lack the capacity for deep video understanding and flexible multistep reasoning. |
| Approach: | They propose a video question answering model which performs dynamic multistep reasoning between questions and videos. |
| Outcome: | The proposed model improves on three widely used video QA datasets and displays better interpretability by backtracing along with the attention mechanisms to the video scene graphs. |
Copied to clipboard
| Challenge: | Existing methods focus on distantly-labeled rationales, ignoring the potential important non-rationale words and not distinguishing the importance of different rationale words. |
| Approach: | They propose two novel auxiliary loss functions to make better use of distantly-labeled rationales, which encourage models to maintain their focus on important words beyond labeled rationals (PINs) and alleviate redundant training on non-helpful rationale (NoIRs). |
| Outcome: | The proposed methods outperform existing methods on two representative classification tasks while maintaining the ability to spread focus to other unlabeled important words. |
Copied to clipboard
| Challenge: | Existing NDR models suffer from large performance drop on hypothetical questions, e.g., “what the annualized rate of return would be if the revenue in 2020 was doubled”. |
| Approach: | They propose a learning to imagine module which can be seamlessly incorporated into NDR models to perform the imagination of unseen counterfactual. |
| Outcome: | The proposed model can perform the imagination of unseen counterfactuals on hypothetical questions. |
Copied to clipboard
| Challenge: | Compared with sentence-level topic structure, paragraph-level topics can grasp and understand the context of a document from a higher level. |
| Approach: | They propose a hierarchical paragraph-level topic structure representation with three layers to guide corpus construction. |
| Outcome: | The proposed method achieves the largest Chinese paragraph-level topic structure corpus, achieving high quality. |
Copied to clipboard
| Challenge: | Existing length control methods involve fine-tuning the parameters of LLMs, which is inefficient and suboptimal for practical use. |
| Approach: | They propose an iterative sampling framework that regulates LLMs to generate length-constrained text without modifying the underlying parameters. |
| Outcome: | The proposed method achieves 100% success rates on Llama3.1 tasks with minimal additional computational overhead. |
Copied to clipboard
| Challenge: | Chinese Spelling Correction (CSC) is a task to detect and correct misspelled characters in Chinese texts. |
| Approach: | They propose a general detector-corrector multi-task framework which exploits the visual and phonological features of the misspelled characters and minimizes their misleading impact on the context. |
| Outcome: | The proposed framework outperforms the state-of-the-art methods on Chinese Spelling Correction tasks. |
Copied to clipboard
| Challenge: | Existing TEC-VQA benchmarks focus on high-resource languages like English and Chinese . existing benchmarks have a "visual-textual misalignment" problem resulting in unreliable evaluation results . |
| Approach: | They propose a benchmark that expands multilingual QA pairs in non-text-centric datasets through translation to eight languages, including Standard Chinese, Korean, and six minority languages. |
| Outcome: | The proposed benchmarks are contamination-free and more challenging . they include eight languages including Chinese, Korean, and six minority languages . |
Copied to clipboard
| Challenge: | Discourse parsing is a challenging task and plays a critical role in discourse analysis. |
| Approach: | They propose a macro discourse structure presentation schema to present the macro level discourse structure analysis. |
| Outcome: | The proposed corpus is based on two tasks of macro discourse structure analysis, including structure identification and nuclearity recognition. |
Copied to clipboard
| Challenge: | Document-level Relation Extraction (DocRE) aims to extract relations among entity pairs in documents. |
| Approach: | They propose a logic constraint framework that uses bidirectional constraints to model rules by beta contribtion and reconstruct rule consistency loss by bidirectional constraint. |
| Outcome: | The proposed framework outperforms existing models in relation extraction performance and logical consistency. |
Copied to clipboard
| Challenge: | Existing video captioning benchmarks and models produce generic captions for videos that lack specific identification of individuals, locations, or organizations. |
| Approach: | They propose a task of directly summarizing news videos into captions that are entity-aware . they validate the effectiveness of their approach across three video captioning models . |
| Outcome: | The proposed approach is effective across three video captioning models. |
Copied to clipboard
| Challenge: | Recent advances in vision-language-action models prioritize robotic action mastery . however, models trained on visual-text pairs struggle to interpret multimodal data . |
| Approach: | They propose a framework that integrates multimodal data after initial control mastery and a Mixture-of-Experts architecture to minimize task interference. |
| Outcome: | The proposed framework surpasses state-of-the-art vision-language-action (VLA) methods on multimodal understanding benchmarks and achieves six times higher performance on visual question-answering datasets. |
Copied to clipboard
| Challenge: | Existing input-centric solutions fail to reverse this intrinsic mechanism of information loss. |
| Approach: | They propose a Variational Information Flow framework that leverages a probabilistic perspective to model visual saliency relevant to the question-answer pair as a latent distribution. |
| Outcome: | The proposed framework improves general VQA, fine-grained perception and visual grounding. |
Copied to clipboard
| Challenge: | Existing MRC models may overuse name information to make predictions, causing name bias . |
| Approach: | They propose a Causal Interventional paradigm for MRC to mitigate name bias by analyzing pre-trained knowledge and context representations. |
| Outcome: | The proposed model is robust to names and performs competitively on the original SQuAD. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) training typically improves single-sample success rates but limited exploration of diverse reasoning trajectories. |
| Approach: | They propose a training paradigm that interleaves conventional RL with inverse reinforcement learning (IRL) they propose 'Steering Probability Squeezing' to enhance exploration without external supervision . |
| Outcome: | The proposed training paradigm improves Pass@k and improves exploration of diverse reasoning trajectories without external supervision. |
Copied to clipboard
| Challenge: | Unsupervised clustering aims at discovering the semantic categories of data according to some distance measured in the representation space, but different categories overlap with each other at the beginning of the learning process. |
| Approach: | They propose a framework to leverage contrastive learning to promote better separation between different categories by optimizing a clustering objective defined in the representation space. |
| Outcome: | The proposed framework improves state-of-the-art accuracy and normalized mutual information on short text clustering and combines top-down and bottom-up instance discrimination to achieve better distances. |
Copied to clipboard
| Challenge: | Existing efforts focus on activation within ongoing dialogues, while overlooking a key real-world bottleneck. |
| Approach: | They propose a conversation starter generation system that generates personalized starters to guide users into conversation without explicit user intent. |
| Outcome: | The proposed system improves user active days by +1.84 and click-through rate by +94.25 and has been deployed in production. |
Copied to clipboard
| Challenge: | Existing approaches to learn dialogue discourse parsing with related tasks require additional annotation, thus limiting their generality. |
| Approach: | They propose a multitasking framework that integrates dialogue discourse parsing with addressee recognition to reflect relation-based structure of dialogue. |
| Outcome: | The proposed framework outperforms baselines on the Molweni and STAC datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) exhibit severe hallucinations, which undermine reliability of automated scientific document understanding systems. |
| Approach: | They propose a framework for mitigating scientific measurement hallucinations through enhanced reasoning and targeted optimization. |
| Outcome: | The proposed framework significantly reduces hallucination rates and improves overall accuracy on the MeasEval benchmark. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have a significant disadvantage for low-resource languages . VEEF-Multi-LLM-8B excels in multilingual instruction-following tasks . |
| Approach: | They propose a low-resource multilingual large language model that expands the vocabulary for multilingual support. |
| Outcome: | The proposed model outperforms existing models in multilingual instruction-following tasks, but lags behind English-centric models in some tasks. |
Copied to clipboard
| Challenge: | Recent studies show that about 30% of summaries generated by neural text summarization suffer from fact fabrication. |
| Approach: | They propose an automatic evaluation metric to measure factual consistency and a learning algorithm that maximizes the metric during model training. |
| Outcome: | The proposed method improves factual consistency and overall quality of summarization models. |
Copied to clipboard
| Challenge: | Existing zero-shot detectors fail when applied to specialized content due to domain shift . DivScore outperforms state-of-the-art detectors in specialized domains . |
| Approach: | They propose a zero-shot detection framework that uses normalized entropy-based scoring and domain knowledge distillation to identify LLM-generated text in specialized domains. |
| Outcome: | The proposed framework outperforms state-of-the-art detectors on medical and legal datasets with 14.4% higher AUROC and 64.0% higher recall. |
Copied to clipboard
| Challenge: | Generalized Category Discovery (GCD) is a crucial task that aims to recognize both known and novel categories from a set of unlabeled data. |
| Approach: | They propose a framework that introduces Large Language Models into the training loop to generate category names without human effort. |
| Outcome: | The proposed framework outperforms SOTA models on three benchmark datasets and generates accurate category names for the discovered clusters. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they reflect and amplify social biases. |
| Approach: | They propose a Mandarin-specific evaluation framework to examine social identity biases in Chinese LLMs using Mandarin-based prompts. |
| Outcome: | The proposed framework compares ingroup (“We”) and outgroup (“They”) framings across 240 social groups salient in the Chinese context, using a two-tiered measurement framework that assesses both sentiment and toxicity. |
Copied to clipboard
| Challenge: | Existing studies only focus on spatial relations extraction as a classification task . spatial information is one kind of critical information for natural language understanding . |
| Approach: | They propose a hybrid model that generates null-role relations and extracts non-null-rol . they propose varying kinds of schemes to represent spatial relation . |
| Outcome: | The proposed model outperforms the baselines on the spatial relation extraction task on SpaceEval. |
Copied to clipboard
| Challenge: | Existing approaches to self-detection only retrospectively evaluate LLM-generated answers, leading to over-trust in incorrectly generated answers. |
| Approach: | They propose a self-detection paradigm that considers the comprehensive answer space beyond LLM-generated answers to mitigate the over-trust in LLM generated incorrect answers. |
| Outcome: | The proposed framework can be integrated with existing approaches for superior self-detection. |
Copied to clipboard
| Challenge: | Conventional topic models suffer different problems in different settings. |
| Approach: | They propose a novel way to model word-pairs named biterms in the whole corpus and a Graph Convolutional Networks (GCNs) with residual connections to extract transitive features from biterm. |
| Outcome: | The proposed model generates more coherent topics compared with previous methods. |
Copied to clipboard
| Challenge: | Social media platforms provide an ideal testbed for large language models that exhibit human-like behavior. |
| Approach: | They propose an LLM-based social **Bot that enhances human-like generative capabilities through an adversarial learning framework. |
| Outcome: | The proposed framework generates human-like content aligned with diverse user profiles . it exhibits strong social responsiveness, more accurately modeling opinion dynamics . |
Copied to clipboard
| Challenge: | Existing knowledge injection benchmarks for large language models lack standardized testing grounds. |
| Approach: | They propose a knowledge injection benchmark that leverages recently-added and expert-curated facts from Wikipedia’s “Did You Know...” entries. |
| Outcome: | The proposed framework improves reliability accuracy by 29.1%. |
Copied to clipboard
| Challenge: | Large language models have significantly advanced Multilingual Machine Translation (MMT) yet scaling to many languages while maintaining robust performance across directions remains challenging. |
| Approach: | They propose a strategy to reduce the number of translations in one direction . they propose auxiliary parallel sentences to promote cross-lingual transfer . |
| Outcome: | The proposed model performs on par with or better than substantially larger baselines. |
Copied to clipboard
| Challenge: | Existing methods for extracting chemical procedures from literature are insufficient and low-quality due to the inherent ambiguity of chemical language and the high cost of human annotation. |
| Approach: | They propose a fully fine-tuned large language model (LLM) as a chemical executor to convert between unstructured experimental procedures and structured action sequences. |
| Outcome: | The proposed model outperforms the baseline model on R2D and D2A tasks by 10%. |
Copied to clipboard
| Challenge: | Autonomous agents powered by large language models (LLMs) have attracted significant research interest, but there are few standards for developing specialized models for agent tasks. |
| Approach: | They propose a series of large action models with dense and mixture-of-expert architectures that unifies, augments, and synthesizes diverse datasets to enhance agent generalizability and performance. |
| Outcome: | The proposed models outperform GPT-4, Claude-3, and many other models in terms of tool use and outperformed GPT-based models on multiple agent ability benchmarks. |
Copied to clipboard
| Challenge: | a neural network model for natural language inference (NLI) is proposed. |
| Approach: | They propose a neuro-symbolic natural logic framework based on reinforcement learning with introspective revision that rewards specific reasoning paths through policy gradients. |
| Outcome: | The proposed model shows superior capability in monotonicity inference, generalization, and interpretability compared with previous models on the existing datasets. |
Copied to clipboard
| Challenge: | Prior work has explored step-level supervision using Shannon-entropy-based uncertainty signals, which conflate inherent state complexity with agent confidence. |
| Approach: | They propose a hierarchical group-based RL framework that leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a mechanism of trajectory-aware reward and trajectory-independent penalty. |
| Outcome: | Experiments on ALFWorld, WebShop, and Search-Augmented QA show that STAPO achieves state-of-the-art performance while substantially alleviating trajectory neglect. |
Copied to clipboard
| Challenge: | Document-level relation extraction is a challenging task as it requires reasoning across multiple sentences. |
| Approach: | They propose to use a recommend-revise scheme to reduce the workload of annotators by providing them with candidate relation instances from distant supervision to supplement and remove relational facts. |
| Outcome: | The proposed dataset is the first large-scale and human-annotated dataset for relation extraction. |
Copied to clipboard
| Challenge: | Existing methods for named entity disambiguation are weak for short text . performance of existing methods drops dramatically for short texts . |
| Approach: | They propose a knowledge-enhanced method for named entity disambiguation . they use factual knowledge graph and conceptual knowledge graph to provide additional knowledge . |
| Outcome: | The proposed method achieves significant improvement on a large manually annotated short-text dataset and the state-of-the-art on three standard datasets. |
Copied to clipboard
| Challenge: | Existing approaches to multilingual machine translation suffer from performance degradation, resulting in a single model being inferior to separately trained bilingual models on resource-rich languages. |
| Approach: | They propose a transformer-based model with a small parameter overhead for multilingual machine translation that outperforms strong multilingual baselines on 64 of 66 language directions. |
| Outcome: | The proposed model outperforms strong multilingual baselines on 64 of 66 language directions, 42 of which have above 0.5 BLEU improvement. |
Copied to clipboard
| Challenge: | Existing methods of implicit discourse relation recognition (IDRR) focus on three aspects: enhancing discourse units representation, enhancing semantic interaction, and joint learning with other tasks. |
| Approach: | They propose a joint model to recognize the relation label and generate the target sentence containing the meaning of relations simultaneously. |
| Outcome: | The proposed model achieves the best performance against several state-of-the-art systems on Chinese and English datasets. |
Copied to clipboard
| Challenge: | Generative AI has demonstrated unprecedented creativity in the field of computer vision, yet such phenomena have not been observed in the realm of literary creation. |
| Approach: | They propose a framework for unleashing the creativity of large language models (LLMs) they assign LLMs to different roles involved in real-world scenario, they write . |
| Outcome: | The proposed framework outperforms baselines in terms of coherence, relevance, interestingness and overall quality on automatically generated screenplays. |
Copied to clipboard
| Challenge: | Existing methods for encoding layout information rely on millions of learnable parameters . polar coordinates provide superior choice for layout modeling, study finds . |
| Approach: | They propose to model layout attention with Gaussian biases by feeding polar coordinates into 2-D Gausssian kernels. |
| Outcome: | The proposed model improves on three widely used benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to training multi-turn attackers to probe model safety vulnerabilities rely on turn-level optimization, which is insufficient for learning long-term attack strategies. |
| Approach: | They propose a multi-turn reinforcement learning problem that optimizes the harmfulness of the final-turn response as the outcome reward. |
| Outcome: | The proposed approach improves attack success rates across multiple models and benchmarks, highlighting the effectiveness of the proposed approach. |