Papers by Han Guo
Copied to clipboard
| Challenge: | Recent advances on abstractive summarization have allowed substantial improvements in the quality of the model, but there is still scope for improvement. |
| Approach: | They propose novel multi-task architectures with high-level layer-specific sharing across multiple encoder and decoder layers of the three tasks and soft-sharing mechanisms. |
| Outcome: | The proposed model improves on the CNN/DailyMail and Gigaword datasets and on the DUC-2002 transfer setup. |
Copied to clipboard
| Challenge: | Existing toxic language detection models focus on the single utterance level without deeper understanding of context. |
| Approach: | They propose a dataset for in-game toxic language detection enabling joint intent classification and slot filling analysis, which is the core task of Natural Language Understanding (NLU). |
| Outcome: | The proposed framework handles utterance and token-level patterns, and rich contextual chatting history. |
Copied to clipboard
| Challenge: | Prompt Engineering (PE) is renowned for improving IE performance through prompt modifications, but the realm of sample design for downstream fine-tuning remains unexplored. |
| Approach: | They propose a methodical approach to enhancing LLMs’ post-tuning performance by refining input, output, and reasoning designs. |
| Outcome: | The proposed approach outperforms heuristic design strategies on three complex IE tasks with four additional LLMs. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for ORMs are largely text-centric or limited to bimodal tasks . a new study examines the effectiveness of Omni-RewardBench for ORms across modalities . |
| Approach: | They propose a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data. |
| Outcome: | The proposed model is the first benchmark for comprehensive evaluation of ORMs across modalities. |
Copied to clipboard
| Challenge: | Maximum likelihood estimation (MLE) is the predominant method for training text generation models. |
| Approach: | They propose a new RL formulation for text generation from the soft Q-learning perspective using path consistency learning to combine the best of on-/off-policy updates and learn effectively from sparse reward. |
| Outcome: | The proposed approach outperforms MLE and previous RL methods in a wide range of tasks. |
Copied to clipboard
| Challenge: | Existing datasets for sarcasm detection are limited due to the difficulty in acquiring ground-truth annotations. |
| Approach: | They propose a generalized latent optimization strategy that allows different losses to accommodate each other and improves training dynamics. |
| Outcome: | The proposed approach outperforms transfer learning and meta-learning baselines and achieves 10.02% performance gain on the iSarcasm dataset. |
Copied to clipboard
| Challenge: | Contract review is labor-intensive, time-consuming, and costly . a benchmark is proposed to detect potential legal conflicts . |
| Approach: | They propose a benchmark for legal provision recommendation and conflict detection for contract auto-reviewing which aims to recommend the legal provisions related to contract clauses and detect possible legal conflicts. |
| Outcome: | The proposed task recommends legal provisions related to contract clauses and detects legal conflicts. |
Copied to clipboard
| Challenge: | HermEs is a spreadsheet formula prediction language that is difficult for Excel users without programming experience to master. |
| Approach: | They propose a hierarchical approach to formula prediction via HiEraRchical forMulet ExpanSion . they propose generating formulas in a fixed order using hierarchically generated formulas . |
| Outcome: | The proposed approach improves formula prediction accuracy by guaranteeing correct grammar and streamlining token-level decoding with high-level Formulet. |
Copied to clipboard
| Challenge: | Existing autoregressive models for dialogue generation suffer from high latency and stability issues. |
| Approach: | They propose a non-autoregressive (NAR) zero-shot spoken dialogue generation model based on flow-matching. |
| Outcome: | The proposed model outperforms existing models in speech generation due to poor speech intelligibility and turn-taking precision. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas, but their effectiveness in complex mathematical reasoning involving multi-step FOL deductions remains under-explored. |
| Approach: | They propose a self-adaptive solution that enhances the Diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct their proofs. |
| Outcome: | The proposed model improves diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct proofs. |
Copied to clipboard
| Challenge: | Existing methods for aspect-based sentiment analysis of review text use only a few keywords describing each aspect/sentiment without using any labeled examples. |
| Approach: | They propose a weakly-supervised approach for aspect-based sentiment analysis which uses only a few keywords describing each aspect/sentiment without using any labeled examples. |
| Outcome: | The proposed method generates quality joint topics and outperforms baselines significantly on benchmark datasets. |
Copied to clipboard
| Challenge: | Multi-task learning is an inductive transfer mechanism that leverages information from related tasks to improve the primary model's generalization performance. |
| Approach: | They propose a multitask learning pipeline that finds relevant auxiliary tasks and learns their mixing ratio. |
| Outcome: | The proposed model can find relevant auxiliary tasks and learn their mixing ratio . the proposed model achieves significant performance boosts on several primary tasks . |
Copied to clipboard
| Challenge: | Existing role-play fine-tuning techniques improve role adaptability but may degrade safety performance, especially for villainous characters. |
| Approach: | They propose safety-aware Role-Play Fine-Tuning (SaRFT) to balance role-playing capabilities and safety. |
| Outcome: | The proposed method outperforms state-of-the-art baselines under both LoRA and full-parameter fine-tuning settings. |
Copied to clipboard
| Challenge: | Existing evaluation methodologies for MWPs diverge from human judgment and face challenges in recognizing mathematically equivalent answers. |
| Approach: | They propose an evaluation metric rooted in graph edit distance that features benefits such as permutation invariance and more accurate program equivalence identification. |
| Outcome: | The proposed evaluation metric features benefits such as permutation invariance and more accurate program equivalence identification. |
Copied to clipboard
| Challenge: | Recent work has shown that statistical language modeling with transformers can greatly improve the performance in code completion tasks. |
| Approach: | They propose a retrieval-augmented code completion framework that combines a source code retriever and an auto-regressive language model for programming language. |
| Outcome: | The proposed framework achieves state-of-the-art on CodeXGLUE benchmark. |
Copied to clipboard
| Challenge: | Recent sparsity-aware binarization approaches can achieve sub-1-bit compression, but they face performance degradation, mask-management overhead, and limited hardware compatibility. |
| Approach: | They propose a binary quantization framework that leverages binary pattern clustering and weight transformation to overcome performance degradation and mask-management overhead. |
| Outcome: | The proposed framework achieves state-of-the-art compression (1.11–0.7 bits) it maintains high performance with only a 3.1% accuracy drop in zero-shot benchmarks while delivering a 1.6 speedup over FP16. |
Copied to clipboard
| Challenge: | Existing approaches to implicit discourse relation recognition lack connectives as strong linguistic clues. |
| Approach: | They propose a transS-driven joint learning architecture to translate discourse relations in low-dimensional embedding space and exploit the semantic features of arguments to assist discourse understanding. |
| Outcome: | The proposed model outperforms existing systems on the Penn Discourse TreeBank. |
Copied to clipboard
| Challenge: | Existing frameworks for data analysis and insight exploration are lacking in terms of benchmarks . existing frameworks suffer from format inconsistencies, poorly conceived objectives, and redundant insights. |
| Approach: | They propose a data-curation pipeline to construct a new dataset named InsightEval. |
| Outcome: | The proposed benchmarks highlight prevailing challenges in automated insight discovery and raise key findings to guide future research. |
Copied to clipboard
| Challenge: | Large language models suffer from knowledge gaps and hallucinations, resulting in incorrect or poor reasoning. |
| Approach: | They propose Graph retrieval-augmented generation (GraphRAG) which integrates structured knowledge from external graphs to enhance model's reasoning. |
| Outcome: | Experiments on knowledge graph QA tasks show that GraphRAG significantly improves reasoning performance across multiple backbone models. |
Copied to clipboard
| Challenge: | Existing methods for streaming video understanding are query-agnostic and implicitly model video evidence. |
| Approach: | They propose a framework that establishes explicit, structured alignment between the accumulated video evidence and the query’s expected response conditions via scene graphs. |
| Outcome: | The proposed model achieves more interpretable and accurate response timing decisions on both proactive and reactive tasks. |
Copied to clipboard
| Challenge: | Several pre-training models of different modalities are showing a rising trend of homogeneity in their model structures. |
| Approach: | They propose a toolkit that supports pre-training models of different modalities. |
| Outcome: | The proposed toolkit can match the performance of the original implementations on text, vision, and audio benchmarks. |
Copied to clipboard
| Challenge: | Recent advances establish "SFT-then-RL" as the defacto paradigm for enhancing large reasoning mod- els on automatically verifiable tasks. |
| Approach: | They propose an entropy-preserving SFT method to enhance exploration capabilities through intrinsic curiosity. |
| Outcome: | The proposed method outperforms the vanilla method on reasoning tasks by 2.5 points . it also outperformed the vanilla SFT by 2.9 points on out-of-distribution tasks . |
Copied to clipboard
| Challenge: | Existing methods focus on designing efficient multimodal fusion frameworks to bridge the semantic gap between images and texts. |
| Approach: | They propose a covariance matrix-driven image channel allocation method that expands the number of original channel maps and assigns importance scores to the expanded channel maps. |
| Outcome: | The proposed method achieves state-of-the-art on three public multimodal fake news detection benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to training LLMs at ultra-low precisions suffer from convergence instability and substantial training costs. |
| Approach: | They propose a progressive QAT framework with outlier channel splitting to address these issues . they use nested structure of integer quantization grids to enable a "train once, deploy any precision" paradigm . |
| Outcome: | The proposed framework outperforms baselines on both Llama2/3 and W2A16, with an 11 speedup over BF16. |
Copied to clipboard
| Challenge: | Large Speech Language Models (LSLMs) typically operate at high token rates to ensure acoustic fidelity, yet this results in sequence lengths that exceed the underlying semantic content, incurring prohibitive inference costs. |
| Approach: | They propose a token-based token merging mechanism that uses a training-free token pooling mechanism to reduce prefilling FLOPs by 27.48% while maintaining competitive accuracy. |
| Outcome: | The proposed method reduces prefilling FLOPs by 27.48% while maintaining competitive accuracy. |
Copied to clipboard
| Challenge: | generative (tokenby-token) inference is memory-bound and requires a large amount of memory to perform. |
| Approach: | They propose a lookup table engine for weight-quantized large language models that uses offline restructuring of the quantized weight matrix to minimize bit manipulations associated with unpacking. |
| Outcome: | The proposed kernel can be 2-4x faster than existing GEMM kernels while achieving performance gains of 1.5 to 2 times. |
Copied to clipboard
| Challenge: | Prompt tuning is a technique for adapting large-scale pretrained language models for downstream tasks. |
| Approach: | They propose to condition a frozen pretrained language model with soft prompts from data . they propose to use a domain adaptation technique to regularize the decision boundary . |
| Outcome: | The proposed method outperforms full-model tuning in data-scarce settings by a large margin. |
Copied to clipboard
| Challenge: | Existing systems that use long-context modeling incur computational and memory overhead. |
| Approach: | They propose a visual memory framework that pre-rendered text into structured images and stored as visual notes for agentic systems. |
| Outcome: | The proposed system reduces token consumption while preserving effective long-term memory recall. |
Copied to clipboard
| Challenge: | Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation. |
| Approach: | They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. |
| Outcome: | The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models. |
Copied to clipboard
| Challenge: | Existing prompt engineering methods exploit database content and execution feedback to improve text-to-sql performance. |
| Approach: | They propose a framework for large language model-based text-to-sql task that exploits database content and execution feedback to improve execution accuracy. |
| Outcome: | The proposed framework improves execution accuracy and usability by 12.41% and 5.38% on four widely used benchmarks. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for long-form speech are limited to limited domains, creating a significant gap with the diverse downstream applications. |
| Approach: | They propose a benchmark that decomposes "long-form speech quality" into specific, disentangled dimensions. |
| Outcome: | The proposed benchmark decomposes “long-form speech quality” into specific, disentangled dimensions. |
Copied to clipboard
| Challenge: | Current outcome-centric verification paradigms neglect potential errors in the derivation process. |
| Approach: | They propose a process-aware RLVR training paradigm utilizing verifiers selected via **PRIME**. |
| Outcome: | The proposed approach outperforms the baseline verification paradigm on AIME24, AIME25, and Beyond-AIME models. |
Copied to clipboard
| Challenge: | Existing search-augmented approaches rely on indiscriminate whole-image retrieval and lack deep iterative reflection, limiting their effectiveness on complex visual queries. |
| Approach: | They propose a fully autonomous framework that shifts from passive perception to active visual planning and introduces a Selective Gaze mechanism that dynamically chooses whether to glance at global context or gaze into high-value regions. |
| Outcome: | Experiments across six benchmarks demonstrate state-of-the-art performance. |
Copied to clipboard
| Challenge: | Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model. |
| Approach: | They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering. |
| Outcome: | The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies. |
Copied to clipboard
| Challenge: | Existing methods for quantizing large language models focus on breaking down the problem into layer-wise sub-problems and minimizing per-layer error, but this approach lacks theoretical justification and the metrics employed may be sub-optimal. |
| Approach: | They propose a "linearity theorem" establishing a direct relationship between the layer-wise reconstruction error and the model perplexity increase due to quantization. |
| Outcome: | The proposed method outperforms previous data-free methods and improves accuracy-compression trade-offs on Llama-family models. |
Copied to clipboard
| Challenge: | Existing studies have shown that cross-lingual knowledge distillation can improve the performance of pre-trained models for cross-linguistic similarity matching tasks. |
| Approach: | They propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model using contrastive learning, bottleneck, and parameter recurrent strategies. |
| Outcome: | The proposed model can compress the size of XLM-R and MiniLM by more than 50% while the performance is only reduced by about 1%. |
Copied to clipboard
| Challenge: | Existing benchmarks explore aspects of threedimensional spatial reasoning and visual-language reasoning in dynamic environments, but they are unable to perform well on 3D spatial deformation reasoning. |
| Approach: | They propose to use a ladder competition format to assess the model's spatial deformation reasoning abilities to determine its performance. |
| Outcome: | The proposed framework assesses the performance of Vision-Language Models in spatial deformation reasoning tasks. |
Copied to clipboard
| Challenge: | Existing vision-and-language navigation methods do not incorporate environmental feedback into their decision-making processes. |
| Approach: | They propose a framework that incorporates environmental feedback into decision-making and a 3D simulator that renders realistic scenarios using Unreal Engine 5. |
| Outcome: | The proposed framework outperforms existing vision-and-language navigation methods in a zero-shot multi-task setting by 28.1% on average. |
Copied to clipboard
| Challenge: | Existing methods for multimodal content detection fail to capture cross-modal semantic inconsistencies and ignore inherent noise in multimodal features. |
| Approach: | They propose a multimodal rumor detection method based on a frequency domain spectral selection method and entropy-guided uncertainty fusion method to capture cross-modal semantic inconsistencies. |
| Outcome: | The proposed method outperforms state-of-the-art methods in multimodal rumor detection . it shows stronger detection capability and robustness on multiple datasets . |
Copied to clipboard
| Challenge: | Existing evaluations of LLMs in finance are text-only, monolingual, and largely saturated by current models. |
| Approach: | They propose a multilingual and multimodal benchmark for evaluating LLMs in real financial contexts. |
| Outcome: | The first expert-annotated multilingual and multimodal benchmark is released . it evaluates 21 leading LLMs and shows they perform better in multilingual settings . |
Copied to clipboard
| Challenge: | Multimodal Process Reward Models (MPRMs) have emerged as a pivotal framework for enhancing the reasoning capabilities of Multimodal Large Language Models. |
| Approach: | They propose a benchmark specifically designed to evaluate MPRMs’ proficiency in detecting erroneous reasoning steps across diverse error categories. |
| Outcome: | The proposed model achieves up to 4.8% performance improvement through test-time scaling. |
Copied to clipboard
| Challenge: | Existing studies on table reasoning focus on flat tables and hierarchical tables . a new dataset, HiTab, aims to examine numerical reasoning over hierarchic tables based on hierarchically structured tables - a strong challenge for existing baselines and a valuable benchmark for future research. |
| Approach: | They propose a hierarchical question answering and natural language generation dataset to study hierarchic tables. |
| Outcome: | The proposed model shows that it is effective in QA and natural language generation over hierarchical tables. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly integrated into real-world decision-making, but their ability to comprehend and reason about policy-related content remains underexplored. |
| Approach: | They propose a bilingual benchmark evaluating policy comprehension comprising 21K cases across a broad spectrum of policy areas. |
| Outcome: | The proposed model shows stronger performance on application-oriented policy tasks than on memorization or conceptual understanding, and yields the highest accuracy on structured reasoning tasks. |
Copied to clipboard
| Challenge: | Existing supervised fine-tuning (SFT) fails to address these issues, as it trains models on single gold-standard responses without modeling nuanced strategy trade-offs. |
| Approach: | They propose a two-stage framework that optimizes strategy selection preferences at each dialogue turn. |
| Outcome: | The proposed framework improves strategy selection preferences at each dialogue turn. |
Copied to clipboard
| Challenge: | Existing methods for finding the optimal prompt for a task are difficult to optimize. |
| Approach: | They propose an efficient discrete prompt optimization approach with reinforcement learning that generates the optimal discrete stimulus after training with reward. |
| Outcome: | The proposed approach is based on a parameter-efficient policy network that generates the optimal discrete prompt after training with reward. |
Copied to clipboard
| Challenge: | Existing studies treat each transformer encoding layer as a single artificial neuron . layer-level embeddings aggregate multiple types of contextual attention captured by multiple head modules . |
| Approach: | They propose to embed each transformer encoding layer as a single artificial neuron . they propose to couple those ANs with their biological-neuron counterparts in the human brain . |
| Outcome: | The proposed models can be used to link representations to brain activity, the authors say . their results show that the proposed models carry meaningful neurolinguistic information . |
Copied to clipboard
| Challenge: | Entity alignment (EA) methods identify the aligned entities based on cosine similarity, ignoring the semantics underlying the embeddings themselves. |
| Approach: | They propose to model entity alignment as a sequential decision-making task where an agent sequentially decides whether two entities are matched or mismatched based on representation vectors. |
| Outcome: | The proposed framework consistently advances the performance of several state-of-the-art methods, with a maximum improvement of 31.1% on Hits@1. |
Copied to clipboard
| Challenge: | Existing work on long document visual question answering is based on Retrieval-Augmented Generation (RAG) where textual or visual content is encoded into embeddings and relevance is determined by similarity scores with respect to the original query. |
| Approach: | They propose a framework that employs an agentic, vision-aware workflow to address long document visual question answering through iterative information discovery and synthesis. |
| Outcome: | The proposed framework outperforms existing RL systems by 10.4% on the MMLongbench-Doc benchmark and demonstrates superior training performance over GRPO. |
Copied to clipboard
| Challenge: | Activation steering offers training-free defense but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. |
| Approach: | They propose an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. |
| Outcome: | The proposed method outperforms baseline methods across multiple jailbreak attacks with minimal impact on utility. |
Copied to clipboard
| Challenge: | Experimental results demonstrate robust performance of the strategy in Chinese & US market regimes compared to established benchmarks. |
| Approach: | They propose a framework leveraging Large Language Models within a risk-aware multi-agent system for automate strategy finding in quantitative finance. |
| Outcome: | The proposed framework outperforms all benchmarks in Chinese & US market regimes with 53.17% cumulative return on SSE50. |
Copied to clipboard
| Challenge: | Recent work has shown that reinforcement learning with simple rule-based reward functions (RLVR) can induce emergent reasoning behaviors and yield gains in challenging domains such as math problem solving. |
| Approach: | They propose a rollout-alignment-quantization-aware RL which aligns training-side forward with the quantized rollout to minimize mismatch. |
| Outcome: | The proposed approach outperforms quantized-rollout training by +5.5 on Qwen3-30B-A3B MoE for math problems while maintaining low-bit throughput. |
Copied to clipboard
| Challenge: | Existing work on LLM-based planning uses language generation to produce free-style plans . however, these plans are not grounded in an executable set of actions . |
| Approach: | They propose a new task for open grounded planning that asks the model to generate an executable plan based on a variable action set. |
| Outcome: | The proposed task is open grounded planning, which is based on a set of variables. |
Copied to clipboard
| Challenge: | Existing legal judgment prediction methods struggle with logical errors when conducting complex legal reasoning. |
| Approach: | They propose a method which enhances LJP reliability through step-wise verification and correction of the reasoning process. |
| Outcome: | The proposed model significantly improves concordance with court decisions from 72.37 to 80.27 on LLAMA-3.1-70B. |
Copied to clipboard
| Challenge: | Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process. |
| Approach: | They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
| Outcome: | The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
Copied to clipboard
| Challenge: | Using large-scale annotation data, large language models can generate noise, errors and biases, leading to unexpected behaviours. |
| Approach: | They propose a dataset to promote safety alignment in large language models . they separate helpfulness and harmlessness annotations for question-answering pairs . |
| Outcome: | The proposed dataset provides 44.6k prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels, with answers generated by Llama-family models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) face excessive computational and memory requirements due to the commonly used Transformer architecture. |
| Approach: | They propose a method to enhance the flow of hidden information between layers in large language models by selectively integrating shallow-layer hidden states into deeper layers. |
| Outcome: | The proposed method maintains parallelizability and inference efficiency of SSMs while significantly boosting performance on public benchmarks. |
Copied to clipboard
| Challenge: | despite popularity of influence functions, their computational cost does not scale well with model and training data size. |
| Approach: | They propose a fast parallel variant that approximates the “influences” of training data-points for test predictions. |
| Outcome: | The proposed method achieves about 80X speedup while being highly correlated with the original influence values. |
Copied to clipboard
| Challenge: | Existing defenses target single-turn attacks, but real-world usage involves multi-turn dialogues, exposing models to attacks that exploit conversational context to bypass safety measures. |
| Approach: | They propose a framework that tackles multi-turn jailbreaks from both attack and defense angles. |
| Outcome: | Experiments on large language models show that MUSE effectively mitigates multi-turn jailbreaks. |
Copied to clipboard
| Challenge: | Conventional retrieval-augmented generation (RAG) methods encode content in isolated chunks during ingestion, losing structural and cross-page dependencies, and retrieve a fixed number of pages at inference. |
| Approach: | They propose a Layout-Aware Dynamic RAG framework that encodes content in isolated chunks during ingestion and retrieves a fixed number of pages at inference. |
| Outcome: | Experiments on MMLongBench-Doc, LongDocURL, DUDE, and MP-DoxVQA show that LAD-RAG improves retrieval, achieving over 90% perfect recall on average without any top-k tuning, and outperforming baseline retrievers by up to 20% in recall at comparable noise levels. |
Copied to clipboard
| Challenge: | Recent studies indicate that large language models (LLMs) may exhibit risks, including threats to the protection of private data and the generation of hallucinations. |
| Approach: | They propose to evaluate LLMs from a legal perspective using the SafeLawBench benchmark. |
| Outcome: | The proposed framework categorizes safety risks into three levels based on legal standards and includes 24,860 multi-choice questions and 1,106 open-domain question-answering tasks. |
Copied to clipboard
| Challenge: | a study of slow reasoning models for multimodal reasoning finds that they are more prone to fabricating plausible yet false details when confronted with incomplete or misleading visual inputs. |
| Approach: | They conduct the first systematic study of the inverse scaling law in slow-thinking paradigms for multimodal reasoning. |
| Outcome: | The findings suggest that slower reasoning models are more prone to fabricating false details . the study analyzed 5,000-sample hierarchical prompt dataset by 50 participants . |
Copied to clipboard
| Challenge: | Existing approaches to address hallucinations in large vision-language models require substantial computational cost and time. |
| Approach: | They propose to leverage sparse autoencoders to identify semantic directions closely associated with faithfulness or hallucination, extracting more precise and disentangled hallucinian-related representations. |
| Outcome: | The proposed method outperforms existing decoding approaches while maintaining transferability across different model architectures with negligible additional time overhead. |
Copied to clipboard
| Challenge: | Large language models lack reliability in scientific domains that require strict adherence to physical constraints. |
| Approach: | They propose a large-scale dataset constructed via a task-adaptive strategy and a hybrid verification protocol that combines deterministic solvers with semantic auditing to guarantee scientific rigor. |
| Outcome: | The proposed model outperforms baselines and general-purpose preference models and is competitive with proprietary models. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on casual conversation or task-oriented dialogue, failing to capture “long-term project-oriented” interactions where agents must track evolving goals. |
| Approach: | They propose a benchmark that simulates the dynamic evolution of memory in real-world projects. |
| Outcome: | The proposed benchmarks simulate the dynamic evolution of memory in real-world projects. |
Copied to clipboard
| Challenge: | a library to facilitate the development, use, and evaluation of large language models (LLMs) is presented. |
| Approach: | They propose a unified library to facilitate the development, use and evaluation of large language models (LLMs). |
| Outcome: | The proposed library is based on extensive experiments in a variety of evaluation settings. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) demonstrate remarkable capabilities across diverse applications, but concerns regarding their security persist. |
| Approach: | They propose an adversarial game that leverages agent learning to extend knowledge to defend against jailbreaks. |
| Outcome: | The proposed game shows that LLMs safeguarded by ICAG exhibit significantly reduced jailbreak success rates across various attack scenarios. |
Copied to clipboard
| Challenge: | Existing studies focus on prompting and developing workflows with frozen LLMs. |
| Approach: | They propose a multi-agentic framework for collaborative LLMs with reinforcement learning that leverages multi-gendered frameworks to enhance collaboration. |
| Outcome: | The proposed model improves collaboration performance across multiple datasets with generalization to unseen domains compared to existing models. |
Copied to clipboard
| Challenge: | Recent advances in end-to-end neural networks-based approaches have shown wide success in sequence generation tasks. |
| Approach: | They propose to optimize multiple metric rewards simultaneously using a multi-armed bandit approach . they empirically show the effectiveness of their approaches via various automatic metrics and human evaluation . |
| Outcome: | The proposed approach improves on question generation and data-to-text generation using a bandit approach. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) excel in tasks ranging from image captioning to complex reasoning. |
| Approach: | They propose a contrastive decoding framework that dynamically calibrates each token generation by mining the model’s internal perceptual discrepancies. |
| Outcome: | The proposed framework mitigates hallucination while enhancing general reasoning capabilities. |
Copied to clipboard
| Challenge: | Existing studies on large language models focus on ethical reviews, failing to capture the diversity of national values. |
| Approach: | They propose a national value extraction pipeline to efficiently construct value assessment datasets and a model-based model with instruction tagging to process raw data sources. |
| Outcome: | The proposed benchmark evaluates the alignment of LLMs with the values of five major nations: China, the United States, the UK, France, and Germany. |
Copied to clipboard
| Challenge: | Sentence simplification is the task of improving readability and understandability of an input text. |
| Approach: | They propose a strong pointer-copy mechanism based sequence-to-sequence sentence simplification model and a novel ‘multi-level’ soft sharing approach where each auxiliary task shares different (higher versus lower) level layers of the model. |
| Outcome: | The proposed model outperforms competing simplification systems in SARI and FKGL automatic metrics, and human evaluation. |
Copied to clipboard
| Challenge: | Despite various proposed data construction methods, their practical utility in real-world pipelines remains underexplored. |
| Approach: | They conduct a comprehensive analysis of open-source datasets and data synthesis techniques for mathematical reasoning under a unified pipeline designed to mirror training and deployment scenarios. |
| Outcome: | The proposed pipelines mirror training and deployment scenarios and are suitable for industrial applications. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have experienced rapid development in recent years, but there is a notable lack of effective and specialized multimodal evaluation datasets in the financial domain. |
| Approach: | They introduce FinMME, a multimodal large language model with 11,000 financial research samples and 20 annotators. |
| Outcome: | The proposed model performs better than state-of-the-art models, highlighting its challenging nature. |
Copied to clipboard
| Challenge: | Existing routers that use hardcoded tools are limited by scalability and generality bottlenecks. |
| Approach: | They propose a pipeline for training history-aware routers to empower precise navigation in large-scale ecosystems. |
| Outcome: | The proposed pipeline can train routers with dynamic context understanding to create the plug-and-play Light Routing Agent. |