Papers by Min Sun
Copied to clipboard
| Challenge: | Iterative non-autoregressive (NAR) models have recently demonstrated impressive performance in varied generation tasks, surpassing the autoregressive Transformer. |
| Approach: | They propose a strategy to conduct efficient refinements without performance declines by using two simple metrics to identify potential problems existing in current refinement processes. |
| Outcome: | The proposed model outperforms the autoregressive Transformer by around one BLEU on average. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) is widely adopted in Large Language Models, but is flat and has limitations such as a significant burden on one retriever and constant granularity limits the ceiling of retrieval performance. |
| Approach: | They propose a progressive retrieval paradigm with coarse-to-fine granularity for RAG, termed FunnelRAG, so as to balance effectiveness and efficiency. |
| Outcome: | The proposed paradigm achieves comparable retrieval performance while the time overhead is reduced by nearly 40%. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on specific predefined model abilities, such as world knowledge, reasoning, etc., making it difficult for users to determine which LLM best suits their particular needs. |
| Approach: | They propose to evaluate large language models from a user-centric perspective and use real-world use cases to identify their effectiveness under distinct intents. |
| Outcome: | The proposed benchmarks achieve a correlation between human preference and the user-reported scenarios and human intents. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are limited by knowledge cutoff and can generate factual hallucinations when handling time-sensitive news. |
| Approach: | They propose a two-stage zero-shot fake news detection framework that uses a hierarchical salience and saliency-calibrated minimum margin of relevance algorithm to extract core entities accurately. |
| Outcome: | The proposed framework outperforms existing zero-shot baselines and even most few-shot methods on two public datasets. |
Copied to clipboard
| Challenge: | Existing models with reasoning capabilities suffer from a severe length collapse in open-ended writing . |
| Approach: | They propose a framework that embeds a dynamic plan-write-reflect cycle into the generation process and train a model with interleaved reasoning traces. |
| Outcome: | The proposed framework achieves state-of-the-art performance on long-form benchmarks compared to other models on the same dataset. |
Copied to clipboard
| Challenge: | Motivational interviewing (MI) is a directive, client-centered counseling approach for eliciting clients' motivation for behavioral change. |
| Approach: | They propose a multi-LLM agent framework for controllable MI dialogue generation . therapist and client agents generate MI-coded utterances guided by MI codes . |
| Outcome: | The proposed framework can generate fluent dialogues with minimal intervention time and a high level of evaluation. |
Copied to clipboard
| Challenge: | Image-to-text tasks such as captioning and controllable image descriptions have received extensive attention for decades. |
| Approach: | They propose a new perspective for image-to-text to generate spatial descriptions by combining two objects in an image. |
| Outcome: | The proposed model is awe-inspiring and human-like, and the proposed end-to-end architecture is the better choice for their integration. |
Copied to clipboard
| Challenge: | Reinforcement learning fine-tuning methods suffer from inefficient exploration and slow convergence . supervised fine- tuning methods have limited performance ceiling and less solid theoretical foundation . |
| Approach: | They propose a Guess-Think-Answer framework that combines supervised and supervised learning in a unified training paradigm. |
| Outcome: | The proposed framework outperforms both standalone SFT and RL training models on three text classification benchmarks. |
Copied to clipboard
| Challenge: | Document assistant chatbots are empowered with extensive capabilities by Large Language Models (LLMs) however, they suffer from hallucinations that are difficult to verify in the context of given documents. |
| Approach: | They propose a document assistant chatbot with reliable attribution that enables users to seek relevant information from given documents. |
| Outcome: | The proposed system generates answers with detailed inline citations, which can be attributed to the original document paragraphs, facilitating verification of factual consistency of the generated text. |
Copied to clipboard
| Challenge: | Large language models have been widely adopted in natural language processing, yet they produce unreliable content. |
| Approach: | They propose to model the attribution task as preference learning and introduce an automatic preference optimization framework that synthesizes attribution preference data. |
| Outcome: | The proposed method achieves state-of-the-art citation F1 with higher answer quality than existing methods. |
Copied to clipboard
| Challenge: | Argumentation mining (AM) aims to detect arguments and their inherent relations from textual compositions. |
| Approach: | They propose a method to model the inter-relationships among three subtasks within a generative framework. |
| Outcome: | The proposed method achieves state-of-the-art performance on two AM benchmarks. |
Copied to clipboard
| Challenge: | Severe acoustic degradation results in unreliable ASR outputs . et al., 2024b): critical concerns regarding reliability and fairness of ASR . |
| Approach: | They propose a multimodal framework that reframes ASR as semantics-guided speech reconstruction. |
| Outcome: | The proposed framework achieves an average reduction in WER while also attaining 98.71% BERTScore and 96.7% USE over advanced baselines. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting can dramatically improve the multi-step reasoning abilities of large language models (LLMs). |
| Approach: | They propose to use Chain-of-Thought (CoT) prompting to encourage the LLM to generate intermediate rationales for solving a problem by providing a series of reasoning steps in the demonstrations. |
| Outcome: | The proposed model can generate coherent lines of reasoning even with invalid demonstrations while still generating coherent lines during inference. |
Copied to clipboard
| Challenge: | Argument mining (AM) is a challenging task as it requires recognizing complex argumentation structures involving multiple subtasks. |
| Approach: | They propose a generative framework where expected outputs of AM are framed as a simple target sequence. |
| Outcome: | The proposed framework achieves state-of-the-art on two AM benchmarks. |
Copied to clipboard
| Challenge: | Existing studies on argumentation mining focus on monological argumentation and dialogical argumentation. |
| Approach: | They propose a mutual guidance framework that could guide arguments in one passage . they propose an inter-sentence relation graph to effectively model the inter-relations between two sentences . |
| Outcome: | The proposed method outperforms the current state-of-the-art model. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for large reasoning models are saturated by a lack of reliable and verifiable benchmarks. |
| Approach: | They propose a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions. |
| Outcome: | The proposed benchmark unifies two evaluation paradigms and offers 150 problems formalized in Lean 4 for rigorous process-level evaluation. |
Copied to clipboard
| Challenge: | a failure in one tool can trigger a cascade of errors, leading to complete task failure. |
| Approach: | They propose a framework for tools more broadly which explores a model’s ability to detect “silent” tool errors and reflect on how to plan. |
| Outcome: | The proposed approach shows that the model can detect "silent" tool errors and plan. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) are becoming tool-interactive agents . quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data . et al.: a new approach to scale trajectory diversity improves tool-use generalization . |
| Approach: | They propose a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. |
| Outcome: | Experiments on general tool-use benchmarks and code agent tasks show that TDScaling improves tool-user generalization and inherent coding proficiency. |
Copied to clipboard
| Challenge: | Existing methods for embodied agents focus on directly executing instructions without considering whether objects can be manipulated. |
| Approach: | They propose a benchmark that evaluates embodied agents in dynamic environments . they use plug-and-play module that augments existing planners with explicit affordance reasoning . |
| Outcome: | The proposed benchmark evaluates embodied agents in dynamic environments with unpredictable affordances . ADAPT significantly improves robustness and task success across seen and unseen environments . |
Copied to clipboard
| Challenge: | Low-resource questions pose a significant challenge within the field of Question-Answering (QA) tasks. |
| Approach: | They propose a method that leverages large models' internal knowledge to enhance the quality of augmented data by Prompt Answer, Question Generation, and Question Filter. |
| Outcome: | The proposed method outperforms existing augmentation strategies on high-resource QA tasks like SQUAD1.1 and TriviaQA. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have demonstrated notable inferential capacities via reinforcement learning (RL) however, “zero-RL” approaches relying on fixed prompt templates introduce substantial sampling inefficiencies for weak LLMs. |
| Approach: | They propose a hierarchical metacognitive RL framework that decomposes zero-accuracy problems into subproblems and prompts the policy to refine answers by referencing previous wrong solutions. |
| Outcome: | The proposed framework improves sample utilization and sample efficiency and accelerates convergence compared to baselines. |
Copied to clipboard
| Challenge: | naive prompts can enhance the task performance of large language models, but they are resource-intensive. |
| Approach: | They propose an automatic prompt optimization method that refines naive prompts according to task outputs from in-box testing models. |
| Outcome: | The proposed method is based on a large-scale dataset and performed fairly across multiple models. |
Copied to clipboard
| Challenge: | Recent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus. |
| Approach: | They propose to use crowdsourcing annotations to build a large-scale but quality-unguaranteed corpus for opinion expression identification in Chinese. |
| Outcome: | The proposed model can be trained with a synthetic expert and is highly consistent with the training and testing phase. |
Copied to clipboard
| Challenge: | Argumentation Mining (AM) aims to extract argumentative structures from texts by identifying argumentation components (ACs) and their argumentative relations (ARs). |
| Approach: | They propose a First- Order Logic reasoning framework for AM to capture logical reasoning paths within argumentative texts. |
| Outcome: | The proposed framework outperforms strong baselines while significantly improving explainability. |
Copied to clipboard
| Challenge: | Argumentative Essay Generation (AEG) is a challenging task in computational argumentation, where detailed logical reasoning and effective rhetorical skills are essential. |
| Approach: | They propose an argumentative planning strategy for prompting large language models to generate high-quality essays by sketch planning and dialectical planning. |
| Outcome: | The proposed method generates more dialectical and persuasive essays with higher diversity compared to baselines. |
Copied to clipboard
| Challenge: | Existing studies consider Aspect Sentiment Classification (ASC) as an independent sentence-level classification problem aspect by aspect. |
| Approach: | They propose a Cooperative Graph Attention Networks approach for cooperatively learning aspect-related sentence representation. |
| Outcome: | The proposed approach outperforms the state-of-the-art methods in document-level sentiment classification. |
Copied to clipboard
| Challenge: | Large Multimodal Models (LMMs) are used to capture subtle differences between images but are noisy and coarse summaries. |
| Approach: | They propose a noise-robust approach to image difference capture using large multimodal models . they use LMMs with structured prompts to generate fine-grained change descriptions . |
| Outcome: | The proposed model outperforms streamlined architectures and improves inference efficiency. |
Copied to clipboard
| Challenge: | Prior work to mitigate fairness issues often employs subjective demonstration selection, leading to low controllability and limited stability across different models and tasks. |
| Approach: | They propose to use in-context learning to insert social biases into large language models to create a structured and controllable representation of the relationship between sensitive attributes and predicted labels. |
| Outcome: | Extensive experiments show that Fair-CCD consistently improves fairness metrics without degrading task accuracy. |
Copied to clipboard
| Challenge: | Global scientific publications are growing annually by about 4%-5% (Pinedo et al., 2024). |
| Approach: | They introduce an AI-assisted platform that answers diverse questions from researchers using Retrieval-Augmented Generation (RAG) they develop various tools to understand queries, search from the scientific literature, filter retrieved information, provide accurate and comprehensive answers, and self-refine answers. |
| Outcome: | OpenResearcher is built on Retrieval-Augmented Generation (RAG) to integrate Large Language Models (LLMs) with up-to-date, domain-specific knowledge. |
Copied to clipboard
| Challenge: | Existing research on customer service dialogue generation generates generic responses from sellers . however, such cost prohibits small businesses, and multiturn dialogue generation is becoming more popular. |
| Approach: | They propose a novel and extensible dialogue generation method by leveraging sellers’ historical dialogue information to generate generic seller responses. |
| Outcome: | The proposed model can generate high-quality responses that cater to specific sellers’ characteristics and exhibit consistent superiority over baselines on a real-world multi-turn customer service dialogue dataset. |
Copied to clipboard
| Challenge: | Existing studies on question answer matching focus on formal text . however, there exists many scenarios where the QA text is informal . |
| Approach: | They propose a novel QA matching approach using informal text from a product review site. |
| Outcome: | The proposed approach improves word-level and sentence-level attentions for solving the noisy problem in the informal text. |
Copied to clipboard
| Challenge: | Recent advances in AM models overlook the integration of supplementary discourse structure information, resulting in suboptimal outcomes. |
| Approach: | They propose a framework which generates discourse structure-aware prefixes for each layer of the generation model. |
| Outcome: | The proposed framework achieves state-of-the-art performance on two AM benchmarks. |
Copied to clipboard
| Challenge: | Existing studies on aspect sentiment classification focus on non-interactive reviews . a new task aims to predict sentiment polarities for specific aspects from interactive reviews based on annotated corpus . |
| Approach: | They propose a task to predict aspects from interactive QA style reviews using an annotated corpus. |
| Outcome: | The proposed approach is compared with state-of-the-art methods against a high-quality corpus of data. |
Copied to clipboard
| Challenge: | extractive models can obtain sentence-level attention with high ROUGE scores but less readable. abstractive models generate novel words and phrases not copied from the source text. |
| Approach: | They propose to combine extractive and abstractive models to achieve a unified model that generates readable paragraphs with word-level attention. |
| Outcome: | The proposed model achieves state-of-the-art ROUGE scores while being the most informative and readable summarization on the CNN/Daily Mail dataset in a solid human evaluation. |
Copied to clipboard
| Challenge: | Argumentation relation classification (ARC) is the most challenging subtask of argumentation mining. |
| Approach: | They propose a dual prior graph neural network to explore probing knowledge and syntactical information for comprehensively modeling the relationship between AC pairs. |
| Outcome: | The proposed model outperforms the state-of-the-art models on three public datasets. |
Copied to clipboard
| Challenge: | Existing studies aim to integrate multiple sub-tasks into a unified ABSA model but suffer from major disadvantages . |
| Approach: | They propose a multi-task learning approach to make use of sub-tasks for a unified ABSA. |
| Outcome: | The proposed model can work well when some sub-tasks are absent, and the interactive relations among subtasks not adequate. |
Copied to clipboard
| Challenge: | Long-context modeling capabilities are important for large language models (LLMs) however, training LLMs with long context windows is insufficient since some samples do not exhibit strong semantic dependencies across long contexts. |
| Approach: | They propose a data mining framework ProLong that assigns each training sample with a long dependency score and ranks and filters them according to their results. |
| Outcome: | The proposed framework can rank and filter training samples that exhibit more powerful long-context modeling abilities. |
Copied to clipboard
| Challenge: | Argument pair extraction (APE) aims to extract interactive argument pairs from two passages within a discussion. |
| Approach: | They propose a method to extract interactive argument pairs from two passages . they propose to decompose the probing graph into four sub-graphs based on inter- and intra-passage perspectives . |
| Outcome: | The proposed method improves on strong baselines on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to classify QA text contain rich sentiment information. |
| Approach: | They propose a task/method to address QA sentiment analysis by annotating QA text pair with annotation guidelines. |
| Outcome: | The proposed method can learn the matching vectors of each Q-sentence, A-sentent unit. |
Copied to clipboard
| Challenge: | Existing text-to-SQL parsing methods mainly focus on English, but there is no labeled data available for the language . a larges-scale and pragmatic Chinese dataset is used for cross-domain text- to-Sql task . |
| Approach: | They propose a larges-scale Chinese dataset for a cross-domain text-to-SQL task . they analyze questions from several representative applications and use an effective data construction framework . |
| Outcome: | The proposed dataset contains 200 databases, 813 tables, and 23,797 question/SQL pairs. |
Copied to clipboard
| Challenge: | Existing models for document-level context translation ignore documentlevel context. |
| Approach: | They propose a document-level context encoder to represent document- level context and integrate it into the Transformer model. |
| Outcome: | Experiments on NIST Chinese-English and IWSLT French-English datasets show that the proposed translation model outperforms the Transformer model significantly. |
Copied to clipboard
| Challenge: | a library to facilitate the development, use, and evaluation of large language models (LLMs) is presented. |
| Approach: | They propose a unified library to facilitate the development, use and evaluation of large language models (LLMs). |
| Outcome: | The proposed library is based on extensive experiments in a variety of evaluation settings. |
Copied to clipboard
| Challenge: | Existing mitigation strategies rely on suppressing specific neuron activations or employing computationally expensive contrastive decoding mechanisms, which often result in increased perplexity or significantly elevated inference latency. |
| Approach: | They propose a lightweight inference-time intervention method grounded in the perspective of residual stream signal dynamics to resolve the signal attenuation of external evidence during its propagation through deep networks. |
| Outcome: | The proposed method improves contextual faithfulness across multiple factual consistency and strong knowledge-conflict tasks while maintaining the model’s general language understanding capabilities. |
Copied to clipboard
| Challenge: | Existing Vision-and-Language Navigation benchmarks assume instructions are feasible and the referenced target exists. |
| Approach: | They propose a benchmark with false-premise instructions where the target is absent . they propose supervised room-level navigation with LLM/VLM-driven in-room exploration . |
| Outcome: | The proposed benchmark produces false-premise goals that are plausible but factually incorrect . ROAM achieves the best REV-SPL among compared methods, while baselines often under-explore and terminate prematurely under unreliable instructions. |
Copied to clipboard
| Challenge: | Recent studies have exposed the risk of Large Language Models (LLMs) generating harmful content by jailbreak attacks. |
| Approach: | They propose a framework that exploits AdVersArial meTAphoR to induce LLMs to calibrate harmful metaphors for jailbreaking. |
| Outcome: | The proposed framework can successfully jailbreak Large Language Models (LLMs) by leveraging the AdVersArial meTAphoR (AVATAR) framework achieves state-of-the-art attack success rate across multiple advanced LLMs. |
Copied to clipboard
| Challenge: | Existing methods to detect and safeguard LLMs against knowledge leakage fail to address the long-term challenge of mitigating it. |
| Approach: | They propose a method to reinforce and safeguard existing benchmarks against knowledge leakage by perturbation-based detection and counterfactual rewriting to disrupt memorization while preserving original intent. |
| Outcome: | The proposed method reduces memorization effects in long-context QA benchmarks, providing a more accurate assessment of model reasoning and generalization abilities. |
Copied to clipboard
| Challenge: | Recent neural networks have shown promising results on Document-level Aspect Sentiment Classification (DASC) however, these approaches often offer little transparency w.r.t. their inner working mechanisms and lack interpretability. |
| Approach: | They propose a Hierarchical Reinforcement Learning approach to DASC that incorporates clause selection and word selection strategies to tackle the data noise problem. |
| Outcome: | The proposed approach over the state-of-the-art approaches shows impressive performance over the current baselines. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have incredible instruction-following capabilities while maintaining strong task completion ability. |
| Approach: | They propose a framework to encourage LLMs to Forget Spurious correlations and Learn from In-context information. |
| Outcome: | The proposed framework can mitigate shortcut learning by forging spurious correlations and learning from in-context information. |
Copied to clipboard
| Challenge: | Existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. |
| Approach: | They propose a multi-task instruction-tuned IR benchmark that includes 126 distinct IR tasks across 6 domains. |
| Outcome: | The proposed model performs better on instruction-tuned models than non-instruction-tunned models on MAIR. |
Copied to clipboard
| Challenge: | Large Language Models have been developed to deal with real-world crimes, but it remains unclear whether they internalize authentic knowledge or are forced to simulate toxic language patterns. |
| Approach: | They construct knowledge-intensive Q&A to investigate misuse threats of Large Language Models in terms of dangerous knowledge possession, harmful task planning utility, and harmfulness judgment robustness. |
| Outcome: | The findings raise concerns that jailbreak success is often attributable to a hallucination loop between jailbroken LLM and judger LLM . |