Papers by Dong Jin
Copied to clipboard
| Challenge: | Existing supervised models struggle to make correct predictions on rare word senses due to limited training data. |
| Approach: | They propose a gloss alignment algorithm that can align definition sentences with the same meaning from different sense inventories to collect rich lexical knowledge. |
| Outcome: | The proposed method outperforms previous methods on both frequent and rare word senses. |
Copied to clipboard
| Challenge: | Large Language Models excel at code generation by learning from vast code corpora, but a fundamental semantic gap remains between training on textual patterns and the goal of functional correctness . reinforcement learning with verifiable rewards (RLVR) approaches are inefficient for establishing a well-aligned connection between the textual representation of code and its execution semantics. |
| Approach: | They propose a novel approach that integrates execution semantics alignment into the RLVR training pipeline for code generation. |
| Outcome: | The proposed model outperforms baseline training and RLVR and shows strong applicability across RL and LLMs. |
Copied to clipboard
| Challenge: | Current self-training methods focus on improving model performance on a single task. |
| Approach: | They propose a cross-task self-training framework where models trained to do different tasks are used in iterative training, pseudo-labeling, and retraining processes to help each other for better selection of pseudo-labeled labels. |
| Outcome: | The proposed framework achieves the best performance compared to baselines on two dialogue understanding tasks. |
Copied to clipboard
| Challenge: | IntentCoding captures the influence of user intent by masking out the intent, and integrates seamlessly with existing decoding procedures. |
| Approach: | They propose a decoding strategy that captures the influence of user intent by masking out the intent and applies a multi-strength ensemble mechanism to amplify the effect of user intention during generation. |
| Outcome: | The proposed model significantly improves both constraint satisfaction and functional correctness compared to greedy decoding approaches. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have shown promise in MER, but their internal decision-making mechanisms under modality conflict and missingness remain underexplored. |
| Approach: | They propose a multimodal large language model that can detect and control modality conflicts and missing subsets by a lightweight mechanism that detects and controls modality conflict. |
| Outcome: | The proposed framework improves performance across settings, showing it can handle conflict and missing behaviors. |
Copied to clipboard
| Challenge: | Existing closed-source LLMs have a performance gap in text-to-SQL reasoning tasks. |
| Approach: | They propose a SQL-based approach to synthesize reliable data to enhance text-to-SQL reasoning in LLMs. |
| Outcome: | The proposed model achieves state-of-the-art accuracy on the widely recognized Spider and BIRD benchmarks, significantly narrowing the performance gap with closed-source methods. |
Copied to clipboard
| Challenge: | Long reasoning steps in LLMs improve reasoning abilities, but the correlation between their effectiveness and the length of reasoning steps remains largely unknown. |
| Approach: | They conducted experiments that expand and compress the rationale reasoning steps within CoT demonstrations while keeping all other factors constant. |
| Outcome: | The results show that lengthening the reasoning steps in prompts significantly enhances LLMs’ reasoning abilities across multiple datasets. |
Copied to clipboard
| Challenge: | Existing solutions for math reasoning tasks use semantic parsing or AST decoding, but performance can degrade dramatically even with slight changes to the questions. |
| Approach: | They propose three calibration methods based on self-consistency for math reasoning tasks. |
| Outcome: | The proposed methods bridge model confidence and accuracy better than existing methods based on p(True) or logit. |
Copied to clipboard
| Challenge: | Existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of Large Language Models (LLMs). |
| Approach: | They propose a repository-level benchmark named DevEval to evaluate LLMs' coding abilities in real-world code repositories. |
| Outcome: | The proposed benchmarks show that the LLMs perform better in real-world code repositories than existing benchmarks. |
Copied to clipboard
| Challenge: | Existing datasets for reading comprehension tasks have been used to test the generalization of natural language understanding systems. |
| Approach: | They propose a diagnostic benchmark suite to clarify key issues related to the robustness and systematicity of NLU systems. |
| Outcome: | The proposed benchmark suite clarifies key issues related to the robustness and systematicity of NLU systems. |
Copied to clipboard
| Challenge: | Existing methods to assess and bolster utterance consistency of chat systems have been shown difficult to detect. |
| Approach: | They propose to use annotators to write dialogue responses and recovery utterances to assess and bolster utteration consistency of chat systems. |
| Outcome: | The proposed dataset significantly improves the detection and resolution of inconsistencies in chat conversations. |
Copied to clipboard
| Challenge: | Existing studies on sentence representation learning focus on human annotation, but they neglect the critical property that essential contents should contribute to sentence semantics more than non-essential contents when encoding a sentence. |
| Approach: | They propose a perturbation method for unsupervised semantic analysis that uses a sentence compression metric to adapt sentence compression datasets for automatic evaluation. |
| Outcome: | The proposed method can capture the main semantics of sentences better than several SOTA unsupervised sentence embedding models. |
Copied to clipboard
| Challenge: | Large language models are reshaping internet services, and serving them is costly. |
| Approach: | They propose an efficient distributed LLM serving system that splits prefill and decode requests into smaller chunks . |
| Outcome: | The proposed system reduces TTFT, TPOT, and latency compared to the state-of-the-art system. |
Copied to clipboard
| Challenge: | Document images are characterized by higher resolutions, denser content, and more complex structural layouts. |
| Approach: | They propose a 1.2B-parameter document parsing vision-language model that decouples layout analysis from local content recognition. |
| Outcome: | The proposed model surpasses general-purpose and domain-specific models on multiple benchmarks while maintaining significantly lower computational overhead. |
Copied to clipboard
| Challenge: | Knowledge-based open-domain dialogue generation aims to build chit-chat systems that talk to humans using mined support knowledge. |
| Approach: | They propose a benchmark for evaluating multi-source dialogue knowledge selection and response generation using Wikipedia's wizard of Wikipedia. |
| Outcome: | The proposed benchmark is called multi-source Wizard of Wikipedia (Ms.WoW) it contains clean support knowledge, grounded at the utterance level and partitioned into multiple knowledge sources. |
Copied to clipboard
| Challenge: | Prompt with Actor-Critic Editing (PACE) for LLMs improves performance of different human-written prompts, resulting in significant performance discrepancies. |
| Approach: | They propose to use LLMs as actors and critics to enable automatic prompt editing by taking feedback from both actors performing prompt and criticizing response into account. |
| Outcome: | The proposed model improves the performance of human-written prompts by 98% and compares to high-quality human-writing prompts. |
Copied to clipboard
| Challenge: | Large language models behave consistently with human goals, values and intentions, but are computationally expensive. |
| Approach: | They propose a framework that enables weak-to-strong alignment transfer via concept transplantation. |
| Outcome: | The proposed framework surpasses instruction-tuned models in terms of truthfulness. |
Copied to clipboard
| Challenge: | Experimental results show that Rex can benefit from cross-lingual training and improve the effectiveness of semantic parsers. |
| Approach: | They propose a Representation Mixup Framework for effectively exploiting translations in the cross-lingual Text-to-SQL task. |
| Outcome: | The proposed framework can benefit from cross-lingual training and improve the effectiveness of semantic parsers, achieving state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing approaches to query-relevant content retrieval fail to retrieve contextually relevant data. |
| Approach: | They propose a multi-agent framework for table question answering over long tables . TALON features a planning agent that iteratively invokes a tool agent to access tabular data . |
| Outcome: | The proposed framework achieves average accuracy improvements of 7.5% and 12.0% across all language models. |
Copied to clipboard
| Challenge: | a new study proposes a conversational search system that integrates product attributes and dialog with search . but it faces two real world challenges: imperfect product schema/knowledge and lack of training dialog data . |
| Approach: | They propose an end-to-end conversational search system that integrates search with text . they propose an utterance transfer approach that generates dialogue utterations from other domains . |
| Outcome: | The proposed system outperforms the best tested baseline in a conversational search dataset for online shopping. |
Copied to clipboard
| Challenge: | Extensive experiments on multi-lingual datasets show that our method significantly outperforms multiple baselines and can robustly handle negative transfer. |
| Approach: | They propose to transfer semantic knowledge from rich-resourced languages to low-resource languages by using multilingual transfer learning. |
| Outcome: | The proposed model outperforms baselines and can handle negative transfer. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit impressive natural language capabilities but suffer from hallucination – generating content that does not align with realworld facts. |
| Approach: | They propose to extrapolate critical token probabilities beyond the last layer to improve decoding by manipulating the predicted distributions at inference time. |
| Outcome: | The proposed methods surpass state-of-the-art on multiple datasets by large margins. |
Copied to clipboard
| Challenge: | Real-world RAG applications often encounter long-context input scenarios where redundant information and noise results in higher inference costs and reduced performance. |
| Approach: | They propose an efficient plug-and-play refiner that leverages the structural characteristics of long documents. |
| Outcome: | Experiments on seven QA datasets show that LongRefiner achieves competitive performance in various scenarios while using 10x fewer computational costs and latency compared to baseline. |
Copied to clipboard
| Challenge: | Existing methods focus on designing efficient multimodal fusion frameworks to bridge the semantic gap between images and texts. |
| Approach: | They propose a covariance matrix-driven image channel allocation method that expands the number of original channel maps and assigns importance scores to the expanded channel maps. |
| Outcome: | The proposed method achieves state-of-the-art on three public multimodal fake news detection benchmark datasets. |
Copied to clipboard
| Challenge: | Tool-calling agents are increasingly deployed in real-world customer-facing workflows . but most studies on tool-callers focus on idealized settings with general, fixed, and well-specified tasks. |
| Approach: | They propose a tool-calling agent-based data pipeline that converts trajectories into user-facing tasks with controlled intent adaptations. |
| Outcome: | The proposed pipeline can be used to study tool use under three scenarios. |
Copied to clipboard
| Challenge: | Randomly concatenating data points can lead to cross-contamination due to the significant difference in their subject matter. |
| Approach: | They propose a method that randomly concatenates data of varying lengths until reaching the designed maximum length to optimize context length and reduce padding. |
| Outcome: | The proposed method significantly improves performance on GSM8K and HumanEval, and also improves fairness and accuracy by 15%. |
Copied to clipboard
| Challenge: | Existing speech codecs struggle to balance these objectives at low bitrates . XY-Tokenizer achieves stronger semantic alignment than representative semantic-distillation codec . |
| Approach: | They propose a low-bitrate speech codec that aligns discrete speech representations with text while preserving fine-grained acoustic details for reconstruction. |
| Outcome: | The proposed codec outperforms existing low-bitrate speech codecs in speech understanding and generation tasks. |
Copied to clipboard
| Challenge: | Existing methods for identifying semantic type of entities are incomplete even in large knowledge bases. |
| Approach: | They propose an attributed and predictive entity embedding method which can fully utilize various kinds of information comprehensively. |
| Outcome: | Experiments on two real DBpedia datasets show that the proposed method outperforms 8 state-of-the-art methods with 4.0% improvement in Mi-F1 and 5.2% improvement in Ma-F1. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) have demonstrated impressive long stepwise reasoning capabilities through large-scale reinforcement learning. |
| Approach: | They propose a framework that enhances large reasoning models with an agentic retrieval-augmented generation mechanism and a Reason-in-Documents module for refining retrieved documents. |
| Outcome: | The proposed framework enhances LRMs with an agentic retrieval-augmented generation mechanism and Reason-in-Documents module for refining retrieved documents. |
Copied to clipboard
| Challenge: | Large language models have achieved remarkable success across a wide range of tasks, yet their performance remains heavily biased toward high-resource languages. |
| Approach: | They propose a pipeline for advancing Tibetan language modeling through multilingual continual pre-training with Tibetan, Chinese, and English. |
| Outcome: | The proposed model outperforms open-source and Tibetan-focused models on diverse tasks. |
Copied to clipboard
| Challenge: | Existing OIE systems organize knowledge into subject-relation-object (SRO) triplets, and they use templates to extract such knowledge triplet. |
| Approach: | They propose a framework to handle expressiveness and groundedness in OpenFact . they propose to use templates, extra constraints, and adopt human efforts to ensure that most triplets contain enough details. |
| Outcome: | The proposed framework improves expressiveness and groundedness of OpenFact . it is more accurate and denser than OPIEC-Linked, which is grounded to Wikidata . |
Copied to clipboard
| Challenge: | Existing datasets do not provide enough annotation to explain unsafe behavior . current chatbots generate toxic and offensive responses, which can be dangerous . |
| Approach: | They construct a dataset called SafeConv that provides comprehensive annotations for chatbots . they compare safe alternatives to rewrite unsafe responses . |
| Outcome: | The proposed model can explain unsafe behavior and detoxify chatbots, the authors show . the proposed model is able to detect unsafe utterances, extract unsafe spans, and convert unsafe responses to safe versions. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have demonstrated remarkable capabilities but can still fail frequently on knowledge-intensive tasks. |
| Approach: | They propose a self-endorsement framework that leverages fine-grained fact-level comparisons across multiple sampled responses. |
| Outcome: | The proposed framework can improve factuality of generations with simple prompts across scales of LLMs. |
Copied to clipboard
| Challenge: | Several noise-robust losses have been proposed and evaluated on tasks in computer vision, but they use a single dataset-wise hyperparamter to control the strength of noise resistance. |
| Approach: | They propose to change single dataset-wise hyperparameters of noise resistance to be instance-wise. |
| Outcome: | The proposed frameworks increase noise-robustness on noisy and corrupted NLP datasets. |
Copied to clipboard
| Challenge: | Large Language Models often require domain-specific fine-tuning to address targeted tasks, which risks degrading their general capabilities. |
| Approach: | They propose to use self-generated dis-preferred weakness data to enhance model performance with a targeted training approach that minimizes interference with existing knowledge base. |
| Outcome: | The proposed approach ensures no reduction in generic capacity and achieves superior performance on downstream tasks compared to existing methods. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have demonstrated remarkable performance across a wide range of downstream tasks. |
| Approach: | They propose a framework that leverages a critic-guided agentic workflow to improve RAG capabilities autonomously. |
| Outcome: | The proposed framework improves RAG capabilities autonomously by leveraging a critic-guided agentic workflow. |
Copied to clipboard
| Challenge: | Existing methods for multimodal content detection fail to capture cross-modal semantic inconsistencies and ignore inherent noise in multimodal features. |
| Approach: | They propose a multimodal rumor detection method based on a frequency domain spectral selection method and entropy-guided uncertainty fusion method to capture cross-modal semantic inconsistencies. |
| Outcome: | The proposed method outperforms state-of-the-art methods in multimodal rumor detection . it shows stronger detection capability and robustness on multiple datasets . |
Copied to clipboard
| Challenge: | Current code generation models produce errors concentrated at specific error-prone points, affecting accuracy of code. |
| Approach: | They propose a framework that focuses preference optimization on error-prone areas . focused-DPO improves the accuracy and reliability of code generation by reducing common errors . |
| Outcome: | The proposed framework improves code generation by focusing on error-prone areas. |
Copied to clipboard
| Challenge: | Existing studies show that multimodal large language models extract visual features from the final layers of a pretrained Vision Transformer. |
| Approach: | They propose a feature fusion method that strategically incorporates shallower layers . they propose MLLMs that extract visual features from the final layers of a pretrained Vision Transformer . |
| Outcome: | The proposed method outperforms deep layers on fine-grained visual tasks . it is the first comprehensive study of visual layer selection for MLLMs . |
Copied to clipboard
| Challenge: | Diffusion language models (DLMs) offer advantages in parallel generation and bidirectional context modeling, but they face a critical trade-off between inference speed and output quality for tasks with strict structural constraints such as code generation. |
| Approach: | They propose an efficient sampling algorithm that reduces the number of tokens unmasked per step based on the model’s evolving confidence. |
| Outcome: | The proposed method improves Pass@1 accuracy by 1.9% while achieving 251.4% inference speedup. |
Copied to clipboard
| Challenge: | A critical bottleneck is the lack of ground-truth human data to link personality traits to emotional shifts. |
| Approach: | They propose a large-scale dataset to capture reader-based emotional variations across news, social media, and life narratives. |
| Outcome: | The proposed model captures reader-based emotional variations across news, social media, and life narratives. |
Copied to clipboard
| Challenge: | Recent studies have focused on content repetition, but structural repetition is a more prevalent problem in code generation. |
| Approach: | They propose a decoding approach that eliminates repetition problems in code generation by identifying grammar rules and strategically decaying the likelihood of critical tokens that contribute to repetitions. |
| Outcome: | The proposed approach outperforms baselines and humanEval benchmarks on CodeRepetEval dataset and MBPP benchmarks, effectively reducing repetitions and enhancing the quality of generated code. |
Copied to clipboard
| Challenge: | Existing text-based relational reasoning models lack a symbolic representation of text . performance gap between NLP models and structured models remains . |
| Approach: | They first pre-train a GNN on a reasoning task using structured inputs and then incorporate its knowledge into an NLP model. |
| Outcome: | The proposed model improves on two state-of-the-art NLP models on 13 different inductive reasoning datasets from the CLUTRR benchmark. |
Copied to clipboard
| Challenge: | Existing methods for preference optimization of large language models use pairs of positive and negative samples, but the quality of positive samples may become similar during training, complicating preference learning. |
| Approach: | SeaPO introduces error types commonly occurring in large language models to improve preference learning. |
| Outcome: | SeaPO introduces error types into model Preference Optimization to improve model performance . negative samples are more erroneous than positive samples, and preference-based training mitigates errors . |
Copied to clipboard
| Challenge: | aims to find more accurate syntactic grammars for accompanying text using video data. |
| Approach: | They build a video-aided grammar induction model that can learn video-span correlation without manual features. |
| Outcome: | The proposed model can learn video-span correlation without manual features adopted by previous systems. |
Copied to clipboard
| Challenge: | Existing pruning methods rely on sequential revisions and unreliable critique signals . Existing methods fail to detect the loss of answer-critical data . |
| Approach: | They propose a table pruning framework which transforms table pruning to gold trajectory-supervised parallel search. |
| Outcome: | The proposed framework outperforms the strongest baseline pruning framework by 3.2% on various tabular reasoning tasks. |
Copied to clipboard
| Challenge: | Considering the vast size and wide-ranging sources of LLMs’ training data, it could explicitly or implicitly include test data. |
| Approach: | They propose a Contamination Detection via output Distribution (CDD) which detects data contamination only by identifying the peakedness of LLM's output distribution. |
| Outcome: | The proposed method improves performance by 21.8%-30.2% on humanEval and TED: trustworthy evaluation via output distribution. |
Copied to clipboard
| Challenge: | Existing methods for inferring the fine-grained type of an entity from knowledge base are incomplete and lack type information. |
| Approach: | They propose a novel Deep Learning architecture to infer the fine-grained type of an entity from a knowledge base. |
| Outcome: | The proposed method significantly outperforms four state-of-the-art methods on two large-scale datasets. |
Copied to clipboard
| Challenge: | Abstractive summarization models implicitly learn to capture the salient information from scratch. |
| Approach: | They propose a method that uses salience expectation to guide abstractive summarization by averaging salient content to a fixed threshold. |
| Outcome: | The proposed method can be easily adapted to documents with various abstractiveness and achieves high performance. |
Copied to clipboard
| Challenge: | Experimental results show that dense retrieval models are better at obtaining query-informed representations. |
| Approach: | They propose a dual-encoder approach that computes latent representations of query and document independently, but inference replaces the real query with a generated one. |
| Outcome: | The proposed approach outperforms previous dense retrieval models on in-domain and out-of-domain datasets. |
Copied to clipboard
| Challenge: | Existing domain-specific code benchmarks focus on assessing what knowledge LLMs possess rather than how they acquire and apply new knowledge. |
| Approach: | They propose a benchmark to evaluate domain specialization methods in real-world software development. |
| Outcome: | KOCO-bench is a new benchmark for evaluating domain specialization methods in real-world software development. |
Copied to clipboard
| Challenge: | Lightweight Vision-Language Models (VLMs) are indispensable for resource-constrained applications. |
| Approach: | They propose a framework that retrieves context from a memory bank to enhance alignment . they propose EMI-based approach to align vision and language models . |
| Outcome: | The proposed framework reduces training loss, accelerates convergence, and enhances task performance with negligible computational overhead. |
Copied to clipboard
| Challenge: | Existing studies have studied history-dependent reasoning for question answering . utilizing global conversation history for enhancement is gaining interest . |
| Approach: | They propose to establish long-distance dependency among global utterances in multi-turn conversation. |
| Outcome: | The proposed method improves on QuAC by 1%, yielding the F1 score of 73.7%. |
Copied to clipboard
| Challenge: | Extensive experiments demonstrate that our framework effectively generates both general and domain-specific data. |
| Approach: | They propose a multi-agent simulator that automatically generates diverse text-based scenarios, capturing a wide range of real-world human needs. |
| Outcome: | Experiments show that the proposed model outperforms Meta’s Llama-3-8B-Instruct model on AlpacaEval 2 and Arena-Hard benchmarks with just 20K instruction-response pairs. |
Copied to clipboard
| Challenge: | Recent studies have focused on factual correctness, semantic grounding, visual reasoning, or multimodal large language models. |
| Approach: | They propose a benchmark to assess AICA, which integrates perception, reasoning, and generation into a unified framework. |
| Outcome: | The proposed framework corrects intensity errors and significantly enhances descriptive depth. |
Copied to clipboard
| Challenge: | Existing methods of multi-modal grammar induction focus on grammar inducing from text-image pairs, but videos provide even richer information, such as static objects and actions. |
| Approach: | They propose a video-aided grammar induction model which learns a constituency parser from unlabeled text and its corresponding video. |
| Outcome: | The proposed model outperforms existing systems on three benchmarks. |
Copied to clipboard
| Challenge: | Existing training methods for code generation do not improve code correctness and efficiency. |
| Approach: | They propose a framework that integrates preference learning into code generation to improve code correctness and efficiency. |
| Outcome: | The proposed framework improves code correctness and efficiency by integrating preference learning into code generation. |
Copied to clipboard
| Challenge: | Existing approaches to fine tune a large language model in low-resource settings are limited in their expressiveness or rely on task-independent knowledge. |
| Approach: | They propose a framework where all parameters are finetuned with task-dependent information from the training data only. |
| Outcome: | The proposed framework outperforms baseline models on several classification datasets in low-resource scenarios. |
Copied to clipboard
| Challenge: | Reinforcement Learning with Verifiable Reward (RLVR) has significantly advanced the complex reasoning abilities of Large Language Models (LLMs). |
| Approach: | They propose a hybrid-policy optimization approach that synergizes internal exploitation with external data to achieve stronger reasoning capabilities. |
| Outcome: | The proposed approach achieves state-of-the-art performance on six math reasoning benchmarks and superior performance on out-of distribution reasoning tasks. |
Copied to clipboard
| Challenge: | Extensive research shows that noisy data significantly degrades the performance of table reasoning in real-world applications. |
| Approach: | They propose a dual denoising framework for complex questions and large-scale tables that uses Tree-guided table pruning to remove irrelevant data step by step. |
| Outcome: | The proposed framework achieves outstanding performance on TableQA tasks with complex questions and large-scale tables. |