Papers by Han Han
Copied to clipboard
| Challenge: | Existing topic models that analyze documents from multiple platforms are not able to capture the authentic topics due to platform-induced biases. |
| Approach: | They propose to use a platform-invariant contrastive learning algorithm to reduce platform influence in topic models by removing platform-specific jargon word sets. |
| Outcome: | The proposed model reduces platform influence in topic models by developing a platform-invariant contrastive learning algorithm and removing platform-specific jargon word sets. |
Copied to clipboard
| Challenge: | Existing models do not detect PII in user prompts, despite their convenience . current models show significant limitations in determining PI I query relevance . |
| Approach: | They propose a query-unrelated PII masking strategy and propose PIi-Bench . they propose 'quick-and-easy' PI I masking with a user query and context description . |
| Outcome: | The proposed model performs well in basic PII detection, but shows significant limitations in query relevance. |
Copied to clipboard
| Challenge: | Recent advances on abstractive summarization have allowed substantial improvements in the quality of the model, but there is still scope for improvement. |
| Approach: | They propose novel multi-task architectures with high-level layer-specific sharing across multiple encoder and decoder layers of the three tasks and soft-sharing mechanisms. |
| Outcome: | The proposed model improves on the CNN/DailyMail and Gigaword datasets and on the DUC-2002 transfer setup. |
Copied to clipboard
| Challenge: | Recent studies show that large language models generate harmful content, but the potential for generating harmful content is an escalating concern. |
| Approach: | They propose to fine-tune LLMs with preference learning to emphasize the preference for timely course-correction by using an automated pipeline. |
| Outcome: | The proposed model improves course-correction skills without affecting general performance and resists jailbreak attacks. |
Copied to clipboard
| Challenge: | Existing methods for Named entity recognition (NER) rely on labeled data, which is labor-intensive. |
| Approach: | They propose a method to de-biase DS-NER models by a structural Causal Model . they propose to use a causal invariance regularizer to make them more robust . |
| Outcome: | The proposed method significantly improves DS-NER models on four datasets and three DS NER models. |
Copied to clipboard
| Challenge: | On-policy distillation (OPD) requires expensive on-the-fly sampling of the student policy during training, which substantially increases training cost. |
| Approach: | They propose to use on-policy distillation to sample trajectories from student model . they propose to terminate the sampling early during distillation . |
| Outcome: | The proposed method matches the performance of full OPD in long reasoning outputs while reducing training FLOP by 2x–40x. |
Copied to clipboard
| Challenge: | Continual pre-training is the paradigm where pre-trained language models acquire fresh knowledge and gradually get upgraded. |
| Approach: | They propose to use adapted weights to recycle old PLMs for continual pre-training . they propose to combine initialization and distillation methods to achieve better performance . |
| Outcome: | The proposed method improves the convergence and performance of the upgraded PLM. |
Copied to clipboard
| Challenge: | Vision Language Models (VLMs) have demonstrated promise in generating visually grounded responses, but their application in the medical domain is hindered by unique challenges. |
| Approach: | They propose a vision language model with versatile visual grounding for medicine that generates semantic segmentation masks and instance-level bounding boxes. |
| Outcome: | The proposed model can generate semantic segmentation masks and instance-level bounding boxes, and accommodates various imaging modalities, including both 2D and 3D data. |
Copied to clipboard
| Challenge: | Existing models that use textual features and sentiments to make stock predictions are poor explainability and low signal-to-noise ratio. |
| Approach: | They propose a bi-level event detection model that detects corporate events from news articles and an elaborately-annotated dataset EDT for corporate event detection and news-based stock prediction benchmark. |
| Outcome: | The proposed strategy outperforms baselines in winning rate, excess returns over the market, and the average return on each transaction. |
Copied to clipboard
| Challenge: | Existing tabular data synthesis methods fail to account for cross-modal heterogeneity of real-world tables, where structured continuous and discrete attributes coexist with unstructured long-text columns. |
| Approach: | They propose a framework that synergistically trains an LLM-based text generator and a deep-learning-based non-textual generator to quantify cross-modal semantic alignment. |
| Outcome: | The proposed framework outperforms state-of-the-art frameworks in fidelity, diversity, and task utility. |
Copied to clipboard
| Challenge: | Experimental results show that Large Language Models can generate rule-based data in long contexts without following all specified rules. |
| Approach: | They propose a novel prompting strategy Multi-Lingual Prompt which automatically translates the error-prone rule that an LLM struggles to follow into another language, thus drawing greater attention to it. |
| Outcome: | The proposed framework outperforms state-of-the-art prompting methods on public datasets across various tasks, with a specific case study in text-to-MIP instances. |
Copied to clipboard
| Challenge: | Existing safety evaluations rely on coarse success rates and domain-specific setups, making it difficult to diagnose why and where these models fail. |
| Approach: | They propose a framework for systematically evaluating the physical safety of LLMs in embodied decision making. |
| Outcome: | The proposed framework assesses the physical safety of LLMs in embodied decision making. |
Copied to clipboard
| Challenge: | Existing toxic language detection models focus on the single utterance level without deeper understanding of context. |
| Approach: | They propose a dataset for in-game toxic language detection enabling joint intent classification and slot filling analysis, which is the core task of Natural Language Understanding (NLU). |
| Outcome: | The proposed framework handles utterance and token-level patterns, and rich contextual chatting history. |
Copied to clipboard
| Challenge: | Recent advances in activation quantization methods cause outliers in tokens, causing extra overhead and speedup . a method to quantize per-tensor activation is currently challenging due to the outlier activation outlier. |
| Approach: | They propose a method to find a set of key-value cache which mitigates outliers in subsequent tokens when inserted as a prefix. |
| Outcome: | The proposed method surpasses the established baseline of per-tensor activation quantization and can be seamlessly integrated with the recent activation quantitative method. |
Copied to clipboard
| Challenge: | Existing methods to optimize large language models suffer from high computational costs and produce uninterpretable, high-perplexity inputs. |
| Approach: | They propose a sparse index-based intervention that bypasses guardrails via sparser logit editing. |
| Outcome: | The proposed method bypasses guardrails by modifying pre-softmax logits without gradients or auxiliary models. |
Copied to clipboard
| Challenge: | Prompt Engineering (PE) is renowned for improving IE performance through prompt modifications, but the realm of sample design for downstream fine-tuning remains unexplored. |
| Approach: | They propose a methodical approach to enhancing LLMs’ post-tuning performance by refining input, output, and reasoning designs. |
| Outcome: | The proposed approach outperforms heuristic design strategies on three complex IE tasks with four additional LLMs. |
Copied to clipboard
| Challenge: | Reward models capture values and preferences of humans and are used in Reinforcement Learning with Human Feedback (RLHF) Traditionally, training large language models relies on extensive human-annotated preference data, which poses significant challenges in terms of scalability and cost. |
| Approach: | They propose a method that enhances RM training using unlabeled data. |
| Outcome: | The proposed approach improves reward models without incurring additional labeling costs on unlabeled datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive language capabilities, but most of them have very unbalanced performance across different languages. |
| Approach: | They propose to use question translation data to enhance LLMs' multilingual capabilities by using mechanistic interpretability methods. |
| Outcome: | The proposed method improves multilingual alignment even with unannotated answers in English and a wide range of languages even with instruction-tuned LLMs. |
Copied to clipboard
| Challenge: | Existing research to improve CoT efficiency falls into three categories, each with distinct limitations. |
| Approach: | They propose a training-free framework that addresses both dimensions of CoT reasoning by applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. |
| Outcome: | Empirical results show that the proposed framework achieves 11.3 efficiency gain without compromising accuracy. |
Copied to clipboard
| Challenge: | Existing safety mechanisms for large language models (LLMs) are inadequate to fully leverage their internal cognitive processes. |
| Approach: | They propose a framework that regulates unsafe outputs by utilizing the prober-based internal state monitor that actively detects harmful intentions. |
| Outcome: | The proposed framework reduces harmful outputs by approximately 80% while maintaining strong utility. |
Copied to clipboard
| Challenge: | Spreadsheets are characterized by their extensive two-dimensional grids, flexible layouts, and varied formatting options, which pose significant challenges for large language models (LLMs). |
| Approach: | They propose a structural-anchor-based compression, inverse index translation, and data-format-aware aggregation module to compress spreadsheets effectively. |
| Outcome: | The proposed method outperforms the existing model in GPT4 and achieves a state-of-the-art 78.9% F1 score. |
Copied to clipboard
| Challenge: | Existing fine-tuning algorithms for vision-language models are restricted by patient privacy concerns and can contain imperceptible noise. |
| Approach: | They propose a framework to mitigate adversarial noise and mitigate upstream noise during fine-tuning. |
| Outcome: | The proposed framework improves model robustness and transferability while decreasing noise levels negatively impact downstream performance. |
Copied to clipboard
| Challenge: | Existing approaches to assess and improve model fairness have been inconsistent and inconsistent. |
| Approach: | They propose an open-source python library for assessing and improving model fairness. |
| Outcome: | The proposed framework can be used for natural language, images, and audio. |
Copied to clipboard
| Challenge: | Autoregressive (AR) and diffusion language models (DLMs) suffer from insufficient reasoning capabilities. |
| Approach: | They propose a fully connected Diffusion Language Model that uses a concept-level causal graph to guide attention to learn causal relationships between concepts. |
| Outcome: | The proposed model achieves a 12% improvement and 3.2 training speedup on the COT-OrderPerturb task, along with an average gain of 1.31% across six downstream reasoning tasks. |
Copied to clipboard
| Challenge: | Existing fashion recommendation systems struggle with the unique challenges of the fashion domain. |
| Approach: | They propose a sequential fashion recommendation framework that leverages a pre-trained large language model enhanced with recommendation-specific prompts. |
| Outcome: | The proposed framework significantly improves fashion recommendation performance on Amazon fashion. |
Copied to clipboard
| Challenge: | Existing code sandboxes fail to provide accurate verification and efficiency under high-concurrency workloads. |
| Approach: | They propose a high-fidelity code verification system that provides sandbox feedback for RL training and evaluation. |
| Outcome: | The proposed system outperforms heuristic-matching baselines on LiveCodeBench and training stability on high-concurrency workloads. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely used in commercial applications . low latency is crucial due to system latency, query concurrency, and computational resources constraints. |
| Approach: | They propose a system that can be resource-efficiently served by addressing bottlenecks beyond LLM inference . they propose 4.3 speed up over vLLM and 1.5 higher throughput . |
| Outcome: | The proposed system outperforms state-of-the-arts with 1.5 higher throughput . it achieves 4.3 speed up with 64 concurrent requests on Mixtral 8x7B . |
Copied to clipboard
| Challenge: | Abusive text is a serious problem in social media and causes many issues among users . a model that detects text abusiveness in context without explicit abusive words is challenging . |
| Approach: | They propose to use an abusive lexicon to determine the existence of an abusive word in text . they combine local and global features to evaluate the model using benchmark data . |
| Outcome: | The proposed model outperforms all previous models for detecting abusiveness in text without abusive words. |
Copied to clipboard
| Challenge: | Existing methods of open-domain dialogue evaluation are labor-intensive and inefficient. |
| Approach: | They propose to use open-domain dialogues to evaluate different aspects of dialogues using holistic evaluation metrics. |
| Outcome: | The proposed metrics show strong correlations with human judgments. |
Copied to clipboard
| Challenge: | Existing methods to integrate extracted knowledge from the Web to knowledge graphs (KGs) however, the predictions are made independently, which can be mutually inconsistent. |
| Approach: | They propose a relation integration model that aligns free-text relations to relations in a target KG . they propose combining two stages to make independent predictions and a collective model that accesses all candidate predictions. |
| Outcome: | The proposed model outperforms baseline models on two datasets and improves AUC from .677 to .748 and from 1.716 to 1.780. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for ORMs are largely text-centric or limited to bimodal tasks . a new study examines the effectiveness of Omni-RewardBench for ORms across modalities . |
| Approach: | They propose a hybrid automatic-annotation and human-verification pipeline to construct high-quality evaluation data. |
| Outcome: | The proposed model is the first benchmark for comprehensive evaluation of ORMs across modalities. |
Copied to clipboard
| Challenge: | Existing data synthesis methods focus on general-purpose tasks and fail to capture domain-specific terminology and reasoning patterns. |
| Approach: | They propose a framework that generates domain-specific instruction datasets without human supervision by pairing task-informed keywords with different cognitive levels from Bloom’s Taxonomy. |
| Outcome: | The proposed framework generates domain-specific instruction datasets without human supervision and achieves significant improvements over existing methods. |
Copied to clipboard
| Challenge: | Existing Paper2Video systems are monolingual and often rely on single-pass pipelines. |
| Approach: | They propose a multilingual agentic Paper2Video system that decomposes the task into planning, audience-oriented critique, layout-aware slide generation, and multilingual figure interpretation. |
| Outcome: | The proposed system improves question-answering accuracy relative to previous systems while maintaining affordable cost and latency. |
Copied to clipboard
| Challenge: | Existing methods for debiasing protected attributes have been limited to binary attributes in isolation, however many corpora involve multiple such attributes, possibly with higher cardinality. |
| Approach: | They propose to evaluate a bias-constrained model which is new to NLP and an extension of the iterative nullspace projection technique which can handle multiple identities. |
| Outcome: | The proposed model is based on a new iterative nullspace projection technique which can handle multiple identities. |
Copied to clipboard
| Challenge: | Neural Machine Translation (NMT) generates translations in isolation, resulting in translation inconsistency and ambiguity. |
| Approach: | They propose to incorporate referring process into translation decoding of NMT by using local coordinates coding to obtain global context vectors containing monolingual and bilingual contextual information. |
| Outcome: | The proposed model improves translation quality with lightweight computation cost on Chinese-English and English-German translation tasks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have impressive capabilities but their application in open-ended, knowledge-intensive, complex reasoning scenarios is limited. |
| Approach: | They propose a framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval-augmented generation within a Monte Carlo tree search paradigm. |
| Outcome: | The proposed framework outperforms the state-of-the-art KAR methods by up to 23.10% and the latest RAG-equipped large reasoning models by upto 25.37%. |
Copied to clipboard
| Challenge: | Recent advances in vision-language models have accelerated research into models capable of advanced reasoning based on images. |
| Approach: | They propose a method that leverages vision-language models to convert charts into table format . they use Large Language Model (LLM) for reasoning to extract only the essential information . |
| Outcome: | The proposed method extracts only the elements necessary for chart reasoning without the need for additional annotations or datasets. |
Copied to clipboard
| Challenge: | Existing IE tools lack multi-task support and automatic updates for KG and EKG construction. |
| Approach: | They propose a human-machine-cooperative IE toolkit for KG and EKG construction that unifies different IE subtasks and integrates LLMs as the assistant machine. |
| Outcome: | The proposed tool improves annotation quality, efficiency, and stability simultaneously. |
Copied to clipboard
| Challenge: | Traditional approaches to truncate inputs, sparse self-attention, and chunking often lead to information loss and hinder the model’s ability to capture long-range dependencies. |
| Approach: | They propose a novel chunk representation method that uses unsupervised keyphrase extraction to group input tokens to retain core document content while reducing input length. |
| Outcome: | The proposed method minimizes information loss and improves the efficiency of Transformer-based models. |
Copied to clipboard
| Challenge: | Existing methods rely on model uncertainty but lack interpretability and data imbalance. |
| Approach: | They propose a lightweight model that predicts relevant database schemas to detect unanswerable questions, enhancing interpretability and addressing the data imbalance in binary classification tasks. |
| Outcome: | The proposed model improves interpretability and improves accuracy in binary classification tasks. |
Copied to clipboard
| Challenge: | MLLMs are deployed on limited image-text pairs, which makes them more vulnerable to catastrophic forgetting of their original abilities during safety fine-tuning. |
| Approach: | They propose a plug-and-play strategy that detects harmful visual inputs and transforms harmful ones into harmless ones. |
| Outcome: | The proposed approach mitigates the risks posed by malicious visual inputs without compromising the original performance of MLLMs. |
Copied to clipboard
| Challenge: | MELLE is a novel language modeling approach for text-to-speech synthesis that generates continuous tokens from text . authors demonstrate that it reduces the need for vector quantization and improves model robustness . |
| Approach: | They propose to autoregressively generate continuous mel-spectrogram frames directly from text condition, bypassing vector quantization. |
| Outcome: | The proposed model achieves superior performance across multiple metrics and is more streamlined. |
Copied to clipboard
| Challenge: | SHARP is a new attack method for structured prediction models that solves several challenges. |
| Approach: | They propose a black-box adversarial attack method that uses a search-based optimization problem to attack adversarials. |
| Outcome: | The proposed method performs more potent attack than pioneer arts on two structured prediction tasks. |
Copied to clipboard
| Challenge: | Existing mRAG systems suffer from a language bias during reranking, systematically favoring English and the query’s native language. |
| Approach: | They propose a language-agnostic utility-driven reranker alignment technique to mitigate language bias during re-ranking. |
| Outcome: | The proposed approach mitigates language bias and consistently improves mRAG performance across languages. |
Copied to clipboard
| Challenge: | EmpathyEar is an open-source, avatar-based multimodal empathetic chatbot . currently, ERG systems rely on text, sound, and vision . |
| Approach: | They propose an open-source, avatar-based multimodal empathetic chatbot to fill the gap in traditional text-only ERG systems. |
| Outcome: | The proposed system enables users to generate emotional responses to user queries . it can also generate avatars with talking faces and synchronized speeches . |
Copied to clipboard
| Challenge: | Existing studies on inflectional morphology disagree on whether or not it makes languages harder to model. |
| Approach: | They propose to use a corpus of 145 Bible translations in 92 languages to investigate whether inflectional morphology makes languages harder to model. |
| Outcome: | The proposed model trains with linguistically motivated subword segmentation strategies and reduces the impact of morphology on language modeling. |
Copied to clipboard
| Challenge: | Existing models for text-rich networks do not take inter-document structure into account. |
| Approach: | They propose a pretraining framework for a text-rich network using a masked language model and a masking node prediction framework. |
| Outcome: | The proposed model outperforms baselines on four tasks in academic and e-commerce domains. |
Copied to clipboard
| Challenge: | ACE 2005 2 is the first large-scale event extraction dataset with 205K event mentions and 3,465 different types. |
| Approach: | They propose to use the DWD Overlay to map PropBank rolesets to a large distantlysupervised training dataset with partial labels to make event extraction more accessible. |
| Outcome: | The proposed model performs better than baselines including InstructGPT and ACE 2005 2 despite being 18 years old . key limitations of ACE include its small event ontology of 33 types, small dataset size of around 600 documents and restricted domain (with a significant portion concentrated on military conflicts). |
Copied to clipboard
| Challenge: | Existing Relation extraction models require extensive annotated training data, which is costly and labor-intensive to collect. |
| Approach: | They propose a new zero-shot RE task where only relation definitions are provided instead of seen-unseen relation instances. |
| Outcome: | The proposed task significantly improves cost-effective zero-shot performance by large margins. |
Copied to clipboard
| Challenge: | Maximum likelihood estimation (MLE) is the predominant method for training text generation models. |
| Approach: | They propose a new RL formulation for text generation from the soft Q-learning perspective using path consistency learning to combine the best of on-/off-policy updates and learn effectively from sparse reward. |
| Outcome: | The proposed approach outperforms MLE and previous RL methods in a wide range of tasks. |
Copied to clipboard
| Challenge: | Existing approaches struggle with mapping questions to precise logical forms . Existing frameworks struggle with complex mapping of questions to logical form . |
| Approach: | They propose a framework that leverages a hierarchical multi-task learning paradigm to enhance the performance of logical form generation. |
| Outcome: | The proposed framework outperforms supervised fine-tuning methods and training-free ones on large language models. |
Copied to clipboard
| Challenge: | Existing methods to recognize entities in text are limited by the diversity of entity types and the lack of high-quality annotations. |
| Approach: | They propose an in-context learning-based NER approach that can inject in-const NER ability into PLMs and recognize entities of novel types on-the-fly using only a few demonstrative instances. |
| Outcome: | The proposed method outperforms the PLMs+fine-tuning counterparts on 4 few-shot NER datasets and significantly outperformed the Plms+initialized extractors. |
Copied to clipboard
| Challenge: | Existing datasets for sarcasm detection are limited due to the difficulty in acquiring ground-truth annotations. |
| Approach: | They propose a generalized latent optimization strategy that allows different losses to accommodate each other and improves training dynamics. |
| Outcome: | The proposed approach outperforms transfer learning and meta-learning baselines and achieves 10.02% performance gain on the iSarcasm dataset. |
Copied to clipboard
| Challenge: | Existing LLMs model overly capable learners who over-apply feedback, resulting in pedagogically implausible behavior. |
| Approach: | They propose a framework that decouples cognitive ability from writing proficiency and models their interaction during writing and revision. |
| Outcome: | The proposed model produces distinguishable proficiency levels and is consistent with instructional theories. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of tasks. |
| Approach: | They explore how LLMs can be extended to interact with and reason about the physical world through IoT sensors and actuators, a concept that they call "Penetrative AI". |
| Outcome: | The proposed approach extends LLMs' capabilities to interact with and reason about the physical world through IoT sensors and actuators. |
Copied to clipboard
| Challenge: | Experimental results show that CLORE is superior to baselines on zero-shot classification tasks. |
| Approach: | They propose a framework for classification by logically parsing and reasoning on natural language explanations. |
| Outcome: | The proposed framework outperforms baselines on zero-shot classification tasks. |
Copied to clipboard
| Challenge: | Existing models that understand spatial concepts and compositional language are inadequate for executing natural language instructions in a physically grounded domain. |
| Approach: | They propose to use knowledge-free auxiliary signals to help the model understand compositional instructions and provide supervision for the instruction's components. |
| Outcome: | The proposed model correctly identifies the source block while the existing model fails on this example. |
Copied to clipboard
| Challenge: | Explicit /think> tags are used to expose intermediate reasoning and enable hybrid thinking behaviors. |
| Approach: | They propose a training-free prompting format that combines these triggers to achieve intermediate-budget reasoning, outperforming fixed-token and prompt-based baselines in terms of the accuracy–length trade-off. |
| Outcome: | The proposed method outperforms fixed-token and prompt-based prompts in accuracy–length trade-offs while improving Qwen3-8B on AIME from 69.8% to 72.4% and on GPQA from 58.5% to 61.1%. |
Copied to clipboard
| Challenge: | Multi-turn response selection models have shown comparable performance to humans in several benchmark datasets, but in the real environment, they often have weaknesses, such as giving the highest score to the wrong response candidate containing several keywords related to the context. |
| Approach: | They propose to build a robust multi-turn response selection model in an adversarial environment and to use it to evaluate weaknesses. |
| Outcome: | The proposed model makes incorrect predictions based heavily on superficial patterns without a comprehensive understanding of the context. |
Copied to clipboard
| Challenge: | Visual storytelling aims to automatically generate a coherent story based on a given image sequence. |
| Approach: | They propose a framework that represents the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge. |
| Outcome: | The proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations. |
Copied to clipboard
| Challenge: | Existing methods for early rumor detection on social media platforms are limited, incomplete and noisy. |
| Approach: | They propose a novel hybrid neural network architecture which combines a task-specific character-based bidirectional language model and stacked Long Short-Term Memory (LSTM) networks to represent textual contents and social-temporal contexts of input source tweets. |
| Outcome: | The proposed model achieves state-of-the-art for detecting unseen rumors on large augmented data which covers more than 12 events and 2,967 rumors. |
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) show strong System-2-style reasoning, but at the cost of significant computational overhead. |
| Approach: | They propose a two-stage curriculum distillation framework which builds a robust internal problem-solving student model and then teaches the student model to externalize this knowledge as explicit reasoning. |
| Outcome: | The proposed model outperforms single-stage baselines on mathematical benchmarks and significantly outperformed LRMs on complex tasks. |
Copied to clipboard
| Challenge: | Admin (Adaptive model initialization) is more stable, converges faster, and leads to better performance. |
| Approach: | They propose a model initialization algorithm to stabilize early training and unleash its full potential in the late stage. |
| Outcome: | The proposed model initialization method stabilizes early training and unleashes full potential in late stage. |
Copied to clipboard
| Challenge: | Automated essay scoring (AES) is a useful tool in English as a foreign language (EFL) writing education. |
| Approach: | They propose a large-scale, standard dataset for rubric-based automated essay scoring with 48.9K samples in total. |
| Outcome: | The proposed system improves the baseline scores by 45.44%. |
Copied to clipboard
| Challenge: | Existing models for implicit hate speech detection do not have significant advantage over cross-entropy loss-based learning. |
| Approach: | They propose a label-aware hard negative sampling strategy that encourages the model to learn detailed features from hard negative samples instead of random batch. |
| Outcome: | The proposed models outperform existing models for implicit hate speech detection both in- and cross-datasets. |
Copied to clipboard
| Challenge: | End-to-end speech translation models have limited training data and are often inefficient due to the inconsistency of length and representation between speech and text. |
| Approach: | They find that the "modality gap" between speech and text data is not a major problem in E2E ST . they decouple the encoder to speech encoder and text encoder, and they find that there is a 'capacity gap' |
| Outcome: | The proposed model achieves 29.0 for en-de and 40.3 for fr on the MuST-C dataset. |
Copied to clipboard
| Challenge: | Existing methods for analyzing social media data lack a systematic integration of medical knowledge, causing a critical treatment gap. |
| Approach: | They propose a framework that leverages Large Language Models to integrate medical knowledge into social media data. |
| Outcome: | The proposed framework can be used to distinguish depression from transient mood changes. |
Copied to clipboard
| Challenge: | Existing benchmarks for logical reasoning in large language models lack language naturalness or limited complexity. |
| Approach: | They propose to use first-order logic annotations to evaluate logical reasoning capabilities of large language models. |
| Outcome: | The proposed dataset evaluates the FOL reasoning ability of supervised fine-tuning on medium-sized language models. |
Copied to clipboard
| Challenge: | Existing methods for camouflaged object segmentation are limited to vision-only mask prediction under fixed task assumptions. |
| Approach: | They propose a language-guided reasoning camouflaged object segmentation task that generates an intent-consistent segmentation mask from an image and an implicit query text instruction. |
| Outcome: | The proposed task can generate an intent-consistent segmentation mask from a camouflaged image and an implicit query text instruction. |
Copied to clipboard
| Challenge: | Modern deep learning models for NLP are notoriously opaque, and this has motivated efforts to design example-specific approaches to interpret such models. |
| Approach: | They propose to use influence functions to explain models by highlighting important words in input text to provide models with an explanation. |
| Outcome: | The proposed approach is particularly useful for natural language inference, a task in which ‘saliency maps’ may not have clear interpretation. |
Copied to clipboard
| Challenge: | Recent automated taxonomies over-rely on a specific corpus, sacrificing generalizability, or depend heavily on the general knowledge of large language models (LLMs) . |
| Approach: | They propose a framework that dynamically adapts an LLM-generated taxonomy to a given corpus across multiple dimensions. |
| Outcome: | The proposed framework performs iterative hierarchical classification, expanding both the taxonomy width and depth based on corpus’ topical distribution. |
Copied to clipboard
| Challenge: | Large Language Model-based Multi-Agent Systems (LLM-MAS) have revolutionized complex problem-solving capability by enabling agent collaboration through message-based communications. |
| Approach: | They propose an attack that exploits communication mechanisms in Large Language Model-based Multi-Agent Systems (LLM-MAS) by intercepting and manipulating inter-agent messages. |
| Outcome: | The proposed attack exploits communication mechanisms in large language model-based multi-agent systems by intercepting and manipulating inter-agencies. |
Copied to clipboard
| Challenge: | In-context learning (ICL) is a form of learning that provides a handful of examples at inference time, but it is not well understood why it emerges as the model has never been specifically trained on such demonstrations. |
| Approach: | They adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL and compare it with random subsets of pretrain data. |
| Outcome: | The proposed method improves the model's ICL ability by 18% if it is continued on a small subset of pretraining data. |
Copied to clipboard
| Challenge: | a recent study examined the effects of media framing on public perception and understanding of news articles. |
| Approach: | They propose to extract framing devices employed by media to assess their role in framating the narrative. |
| Outcome: | The proposed method surpasses baseline models and offers a more detailed and explainable analysis of media framing effects. |
Copied to clipboard
| Challenge: | Numerical Question Answering is the task of answering questions that require numerical capabilities. |
| Approach: | They propose to conduct numerical capability diagnosis on a series of Numerical Question Answering systems and datasets. |
| Outcome: | The proposed approach relieves existing systems’ lack of robust numerical capabilities. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for RAG systems are lacking due to high costs of data construction and lack of factual accuracy. |
| Approach: | They propose a framework to evaluate RAG systems in specialized scenarios . they propose three new metrics to evaluate LLM-generated responses . |
| Outcome: | The proposed framework outperforms zero-shot and one-shot methods in terms of clarity, safety, conformity, and richness of generated samples. |
Copied to clipboard
| Challenge: | Existing methods for integrating knowledge graphs rely on entity and relation embeddings . Fig. 1 shows how to decode knowledge graph in under 6 seconds . |
| Approach: | They propose a framework that only utilizes entity embeddings to decode knowledge graphs. |
| Outcome: | The proposed framework reconstructs KG representation by maximizing smoothness of entity embeddings. |
Copied to clipboard
| Challenge: | Current neural event detection approaches focus on trigger-centric representations, which work well on distilling discrimination knowledge, but poorly on learning generalization knowledge. |
| Approach: | They propose a Delta-learning approach to distill discrimination and generalization knowledge by incrementally learning and adaptively fusing event representation. |
| Outcome: | The proposed method significantly outperforms previous approaches on unseen/sparse trigger words and achieves state-of-the-art performance on ACE2005 and KBP2017 datasets. |
Copied to clipboard
| Challenge: | Instruction fine-tuning (IFT) is a crucial phase in building large language models (LLMs). |
| Approach: | They propose a knowledge intervention framework to decouple the potential underlying factors of IFT and enable individual analysis of different factors. |
| Outcome: | The proposed framework decouples the potential underlying factors of IFT, enabling individual analysis of different factors. |
Copied to clipboard
| Challenge: | Contract review is labor-intensive, time-consuming, and costly . a benchmark is proposed to detect potential legal conflicts . |
| Approach: | They propose a benchmark for legal provision recommendation and conflict detection for contract auto-reviewing which aims to recommend the legal provisions related to contract clauses and detect possible legal conflicts. |
| Outcome: | The proposed task recommends legal provisions related to contract clauses and detects legal conflicts. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for natural language generation are dominated by similarity-based metrics. |
| Approach: | They propose a multi-dimensional evaluator for natural language generation that integrates multiple dimensions into one evaluer. |
| Outcome: | The proposed evaluator improves on three typical NLG tasks and improves with external knowledge. |
Copied to clipboard
| Challenge: | Recent studies have focused on the use of large language models (LLMs) for table-based reasoning, but most approaches struggle with scalability when applied to large tables. |
| Approach: | They propose a framework to harness latent augmentation potential in tabular data . they use only a small subset of relevant data from the table to supplement it with schema . |
| Outcome: | The proposed framework outperforms all other approaches and exhibits robustness and efficiency against perturbations in large-table scenarios. |
Copied to clipboard
| Challenge: | entailment : absence of questions classified based on their rewriting hardness or difficulty . enactment of QR system to rewrite context-dependent questions in CQA requires context knowledge . |
| Approach: | They propose a heuristic method to automatically classify questions into subsets of varying hardness . they then conduct a human evaluation to annotate the rewriting hardness of questions . |
| Outcome: | The proposed learning framework improves the overall performance compared to baselines. |
Copied to clipboard
| Challenge: | Real-world machine learning systems are achieving excellent performance in terms of coarse-grained metrics like overall accuracy and F-1 score. |
| Approach: | They extend slice-based learning (SBL) with a mixture of attentions to learn slice-aware dual attentive representations. |
| Outcome: | The proposed approach outperforms the baseline method and the original SBL approach on monitored slices with two natural language understanding tasks. |
Copied to clipboard
| Challenge: | Existing work describes paragraph-level counter-argument generation task as paragraph-based . however, sentence-level generation can be quite different due to its unique constraints and brevity-focused challenges. |
| Approach: | They propose a benchmark framework for sentence-level counter-argument generation . they use an annotated debate forum dataset to generate high-quality counter-argments . |
| Outcome: | The proposed framework and evaluator are competitive in counter-argument generation tasks. |
Copied to clipboard
| Challenge: | Conventional text embedding methods suffer from information loss if directly adapted to hyper-documents. |
| Approach: | They propose an embedding approach for hyper-documents that incorporates four criteria to preserve necessary information for embeddable models. |
| Outcome: | The proposed model outperforms several existing models on two tasks in the academic domain. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable performance across a diverse set of domain-specific tasks. |
| Approach: | They propose a non-monolithic LLM querying system that seamlessly integrates various LLM experts into a single query interface and dynamically routes incoming queries to the most high-performant expert based on query’s requirements. |
| Outcome: | The proposed model improves query efficiency by 40% and costs by 30% while maintaining or enhancing model performance by 10%. |
Copied to clipboard
| Challenge: | Multi-agent systems based on large language models are limited by high computational overhead, information loss, and robustness. |
| Approach: | They propose a Residual Mixture-of-Agents (RMoA) that integrates residual connections to optimize efficiency and reliability. |
| Outcome: | The proposed model achieves state-of-the-art performance on benchmarks of alignment, mathematical reasoning, code generation, and multitasking understanding, while significantly reducing computational overhead. |
Copied to clipboard
| Challenge: | Existing methods for annotating instruction data are expensive and difficult to scale. |
| Approach: | They propose a method to automatically build instruction data from an unlabeled corpus without heavy reliance on proprietary LLMs and human annotation. |
| Outcome: | The proposed method outperforms existing methods on AlpacaEval leaderboard and other open-source methods. |
Copied to clipboard
| Challenge: | Video-guided Machine Translation (VMT) uses short video clips to enhance translation quality, but many samples are text-sufficient. |
| Approach: | They propose a framework that integrates multimodal large language models’ multimodal reasoning into video-guided machine translation by using a pipeline for constructing training data based on multimodal relevance to translation. |
| Outcome: | The proposed framework improves multimodal information utilization in video-guided machine translation, yielding gains in translation quality and computational efficiency. |
Copied to clipboard
| Challenge: | HermEs is a spreadsheet formula prediction language that is difficult for Excel users without programming experience to master. |
| Approach: | They propose a hierarchical approach to formula prediction via HiEraRchical forMulet ExpanSion . they propose generating formulas in a fixed order using hierarchically generated formulas . |
| Outcome: | The proposed approach improves formula prediction accuracy by guaranteeing correct grammar and streamlining token-level decoding with high-level Formulet. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown increasing power on NLP tasks. however, tuning these models for downstream tasks usually requires exorbitant costs. |
| Approach: | They propose a black-box tuning technique that optimizes task-specific prompts without accessing gradients and hidden representations. |
| Outcome: | The proposed method improves performance under few-shot learning scenarios. |
Copied to clipboard
| Challenge: | Existing quality filtering methods rely on a high-quality dataset as reference . Existing methods introduce potential biases and compromise diversity . |
| Approach: | They propose a method that evaluates text quality based on the perplexity difference between two language models trained on the same data. |
| Outcome: | The proposed approach improves performance of pre-trained models without increasing training costs. |
Copied to clipboard
| Challenge: | Existing methods for multi-hop reasoning assume that every relation has enough triples for training . however, performance drops significantly on few-shot relations . |
| Approach: | They propose a meta-based multi-hop reasoning method that learns meta parameters from high-frequency relations that could quickly adapt to few-shot scenarios. |
| Outcome: | The proposed method outperforms state-of-the-art methods in few-shot scenarios on two public datasets from Freebase and NELL. |
Copied to clipboard
| Challenge: | Existing methods for extracting medical decision trees rely on manual annotation . PI-LoRA is a low-rank adaptation method for extract medical decision tree from clinical guidelines and textbooks . |
| Approach: | They propose a low-rank adaptation method for automatically extracting medical decision trees from clinical guidelines and textbooks. |
| Outcome: | The proposed method outperforms existing methods for the Text2MDT task while maintaining a lightweight architecture. |
Copied to clipboard
| Challenge: | Existing work on grounding events into a precise timeline has been limited due to the inherent ambiguity of language and the requirement for information propagation over inter-related events. |
| Approach: | They propose a 4-tuple temporal representation for entity slot filling to ground events into a timeline using a graph attention network approach. |
| Outcome: | The proposed approach yields 7.0% match rate over contextualized embedding approaches and 16.3% higher match rate compared to sentence-level manual event time argument annotation. |
Copied to clipboard
| Challenge: | Recent LLM-based search agents often concatenate the full interaction history into the context, producing long and noisy inputs and increasing compute cost and memory overhead. |
| Approach: | They propose an agent framework that maintains a compact memory during multi-turn interactions. |
| Outcome: | The proposed framework outperforms strong history-concatenation (ReAct-style) baselines on a range of public datasets while maintaining nearly constant token counts across multi-turn interactions. |
Copied to clipboard
| Challenge: | Existing name entity recognition methods combine pre-trained language models with supervised models such as BiLSTM/LSTM-CRF to perform poorly in a spoken dialogue context. |
| Approach: | They propose a logic-guided fine-grained address recognition method that softly applies the logic rule to improve the accuracy of FGAER. |
| Outcome: | The proposed method improves fine-grained address entity recognition from multi-turn spoken dialogues. |
Copied to clipboard
| Challenge: | Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling where reliability depends on preserving consistent roles, personas, and goals across long horizons. |
| Approach: | They propose a framework that decomposes LLM–LLM conversations into a modular, stability-first framework that allows for a stable persona-driven agent simulation for multi-turn dialogue generation. |
| Outcome: | The proposed framework decomposes the LLM-based model into four main components: persona creation, plausibility validation, and natural-language persona crafting. |
Copied to clipboard
| Challenge: | Existing autoregressive models for dialogue generation suffer from high latency and stability issues. |
| Approach: | They propose a non-autoregressive (NAR) zero-shot spoken dialogue generation model based on flow-matching. |
| Outcome: | The proposed model outperforms existing models in speech generation due to poor speech intelligibility and turn-taking precision. |
Copied to clipboard
| Challenge: | Using word-based models, we compare word-oriented models with char-based ones . word-driven models are more vulnerable to data sparsity and the presence of out-of-vocabulary words . |
| Approach: | They benchmark word-based models with char-based model which does not involve word segmentation in four NLP benchmark tasks. |
| Outcome: | The proposed model outperforms char-based models in four NLP benchmark tasks. |
Copied to clipboard
| Challenge: | supervised dependency parsers can reach a very high accuracy, but they require treebanks for training. |
| Approach: | They propose a second-order extension of unsupervised neural dependency models that incorporate grandparent-child or sibling information. |
| Outcome: | The proposed model achieves 10% improvement over the previous state-of-the-art model on the full WSJ dataset. |
Copied to clipboard
| Challenge: | Despite the critical role of software requirements, these criteria have not been studied actively in previous code generation works. |
| Approach: | They propose a framework that leverages in-context learning to organize and extrapolate unexpressed requirements from textual descriptions. |
| Outcome: | The proposed framework generates functional requirements from textual descriptions and extrapolates unexpressed requirements from them. |
Copied to clipboard
| Challenge: | Triton is a high-level Python-like programming language for building efficient GPU kernels. |
| Approach: | They propose a TritonBench benchmark that provides a comprehensive evaluation of Tritonic operators on widely deployed GPUs. |
| Outcome: | The proposed benchmarks show that current LLMs struggle to generate efficient Triton operators on widely deployed GPUs aligned with industry applications. |
Copied to clipboard
| Challenge: | Existing knowledge distillation techniques for neural machine translation lack special treatment on the top-1 information, which is limiting the potential of KD. |
| Approach: | They propose a method to distill knowledge from top-1 predictions of teachers and a technique to infuse more additional knowledge by distilling on the data without ground-truth targets. |
| Outcome: | The proposed method outperforms the vanilla word-level KD and outperfies the existing methods on three different students with different capacity gaps. |
Copied to clipboard
| Challenge: | Existing methods for visual semantic representation learning struggle to address semantic uncertainty, especially in the medical domain. |
| Approach: | They propose a hyperbolic density embedding based image-text representation learning approach tailored for specific medical domain data. |
| Outcome: | The proposed method performs better than baseline methods on zero-shot tasks and fine-tuning tasks on different datasets. |
Copied to clipboard
| Challenge: | PromptSource is a system for creating, sharing, and using natural language prompts . prompts are used to train and query language models in zero-shot learning settings . |
| Approach: | PromptSource is a system for creating, sharing, and using natural language prompts . et al.: using prompts to train and query language models is emerging area in NLP . they propose a templating language for defining data-linked prompts, a user interface that iterates on prompt development . |
| Outcome: | PromptSource is a system for creating, sharing, and using natural language prompts . it has a templating language for defining data-linked prompts and a community-driven set of guidelines . |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) achieve surprising performance on the Choice of Plausible Alternatives (COPA) task. |
| Approach: | They propose to add a regularization loss to the existing COPA models to mitigate the problem of semantic similarity bias by adding a normalization loss. |
| Outcome: | The proposed model improves generalization ability and performs better on a challenging dataset, BCOPA-CE, which has unbiased token distribution and is more difficult for models to distinguish cause and effect. |
Copied to clipboard
| Challenge: | Video-guided machine translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips. |
| Approach: | They propose a plug-and-play framework for video-guided machine translation with multimodal large language models. |
| Outcome: | The proposed framework improves performance of MLLMs while reducing computational cost. |
Copied to clipboard
| Challenge: | Existing frameworks for Augmented Language Models lack flexibility, democratization, and holistic evaluation. |
| Approach: | They propose a lightweight and extensible framework for Augmented Language Models called Gentopia. |
| Outcome: | The proposed framework integrates language models, task formats, prompting modules, and plugins into a unified paradigm. |
Copied to clipboard
| Challenge: | Existing methods to improve NLU are laborintensive and expensive. |
| Approach: | They propose a scalable and automatic approach to improving NLU in a large-scale conversational AI system by leveraging implicit user feedback. |
| Outcome: | The proposed framework improves NLU in a large-scale conversational AI system across 10 domains. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas, but their effectiveness in complex mathematical reasoning involving multi-step FOL deductions remains under-explored. |
| Approach: | They propose a self-adaptive solution that enhances the Diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct their proofs. |
| Outcome: | The proposed model improves diversity and REAsonability of LLMs’ generation strategies by introducing an Axiom-Driven Strategy Diversification mechanism and a Sub-Proposition Error Feedback to help LLM reflect on and correct proofs. |
Copied to clipboard
| Challenge: | Existing red-teaming methods for large language models often discover safety risks without addressing them. |
| Approach: | They propose a multi-round automatic red-teaming method that incorporates both adversarial prompt writing and safe response generation. |
| Outcome: | The proposed method significantly increases red-teaming scalability and the safety of the target LLM. |
Copied to clipboard
| Challenge: | Existing models struggle to detect elaborately disguised malicious URLs, despite their ability to process malicious URL's. |
| Approach: | They propose a benchmark to evaluate LLMs’ vulnerabilities to malicious URLs and a lightweight defense module to mitigate the vulnerability. |
| Outcome: | The proposed framework analyzes 61,845 attack instances spanning 10 real-world scenarios and 7 categories of real malicious websites. |
Copied to clipboard
| Challenge: | Existing stance detection research on news content is limited to short texts and high-resource languages. |
| Approach: | They propose a dataset for article-level stance detection that integrates viewpoints into recommendation algorithms and a framework that employs a language model agent to predict the stances of key structural segments. |
| Outcome: | The proposed framework outperforms existing methods in identifying article stances and uncovering patterns of media bias. |
Copied to clipboard
| Challenge: | Existing topic models adopt a fully unsupervised setting and their discovered topics may not reflect user preferences well due to their unsupervised nature. |
| Approach: | They propose a framework that allows out-of-vocabulary seeds to be used to find latent topics from text corpora. |
| Outcome: | The proposed framework can find topics that are never seen in the corpus and can benefit from the general knowledge of pre-trained language models. |
Copied to clipboard
| Challenge: | Existing methods for aspect-based sentiment analysis of review text use only a few keywords describing each aspect/sentiment without using any labeled examples. |
| Approach: | They propose a weakly-supervised approach for aspect-based sentiment analysis which uses only a few keywords describing each aspect/sentiment without using any labeled examples. |
| Outcome: | The proposed method generates quality joint topics and outperforms baselines significantly on benchmark datasets. |
Copied to clipboard
| Challenge: | lexical overlap is a common evaluation metric for extractive summarization, but recent studies reveal its limitations. |
| Approach: | They propose a facet-aware evaluation setup for better assessment of information coverage in extractive summaries. |
| Outcome: | The proposed evaluation setup improves human correlation with extractive summarization datasets and improves comparative analysis. |
Copied to clipboard
| Challenge: | Existing frameworks for Large Language Models (LLMs) for Click-Through Rate prediction require a careful balance between computational efficiency and predictive accuracy. |
| Approach: | They propose a framework that integrates Retrieval-Augmented Generation with a novel multi-head early exit architecture to address both challenges. |
| Outcome: | The proposed framework reduces retrieval time while maintaining high model performance. |
Copied to clipboard
| Challenge: | Currently, large language models (LLMs) train on short text segments due to the computational overhead quadratic in the input lengths of their Transformer architectures. |
| Approach: | They propose a method that allows LLMs pre-trained with 2K or 4K-long segments to generalize to up to 200M length inputs while retaining perplexity. |
| Outcome: | The proposed method achieves 2.7 decoding speed up and 7.5 memory saving over the original model. |
Copied to clipboard
| Challenge: | Recent advances in the field of computer vision have enabled more effective and sophisticated interactions between humans and machines. |
| Approach: | They propose a reasoning-based object detection paradigm that leverages state-of-the-art multi-modal models and open-vocabulary object detectors to perform reasoning within the context of the user’s instructions and the visual scene. |
| Outcome: | The proposed method enables users to interact with the system using natural language instructions, allowing for a higher level of interactivity. |
Copied to clipboard
| Challenge: | Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found. |
| Approach: | They investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path. |
| Outcome: | The proposed model can be used to find low-loss paths between two minima, and to understand how their mode connectivity affects their task knowledge. |
Copied to clipboard
| Challenge: | Existing surveys on scientific LLMs focus on one or two fields or a single modality. |
| Approach: | They survey 260 scientific LLMs and examine their architectures and pre-training techniques . they also discuss commonalities and differences between LLM architectures . |
| Outcome: | The proposed model architectures and evaluation techniques are used to improve scientific discovery. |
Copied to clipboard
| Challenge: | Existing Large Language models with text inputs lack the capability to evolve with non-expert interactions with environments. |
| Approach: | They propose a novel learning paradigm that generates robots’ executable actions in the form of text, derived solely from visual observations. |
| Outcome: | The proposed learning paradigm surpasses baselines and can adapt to the target tasks effectively. |
Copied to clipboard
| Challenge: | Recent advances in multimodal reasoning may pose new safety risks . evaluators neglect reasoningbased safety, where harm emerges only through MLLMs . |
| Approach: | They introduce a benchmark for multi-image reasoning safety that includes 2,676 instances . they find that models with more advanced multi- image reasoning are more vulnerable . |
| Outcome: | The proposed benchmark consists of 2,676 instances covering 9 multi-image relations . the results show that models with more advanced multi- image reasoning are more vulnerable . |
Copied to clipboard
| Challenge: | a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks. |
| Approach: | They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples . |
| Outcome: | The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset. |
Copied to clipboard
| Challenge: | Existing large-scale pre-trained language models are mainly trained from scratch individually, ignoring that many well-taught PLMs are available. |
| Approach: | They propose a pre-training framework called knowledge inheritance and propose auxiliary supervision to efficiently learn larger PLMs. |
| Outcome: | The proposed framework can be used to train large-scale language models with huge parameters and a large dataset can be adapted to domain adaptation and knowledge transfer. |
Copied to clipboard
| Challenge: | Existing prompt optimization methods rely on extensive manual effort or meta-cognitive abilities, making them less effective for LwLLMs. |
| Approach: | They propose a direct behavior optimization parameter that transforms the optimization of complex prompts into discrete, quantifiable execution sequences using a gradient-free Monte Carlo Tree Search. |
| Outcome: | The proposed method outperforms current prompt optimization methods on seven challenging tasks where state-of-the-art LLMs excel but LwLLMs generally underperform. |
Copied to clipboard
| Challenge: | Existing datasets exhibit data scarcity and limited coverage of general-domain events. |
| Approach: | They present a MAssive eVENt detection dataset which contains 4,480 Wikipedia documents and 168 event types. |
| Outcome: | The proposed dataset shows that existing methods cannot achieve promising results on the small datasets. |
Copied to clipboard
| Challenge: | Prompt-based learning has been an effective paradigm for large pretrained language models (LLMs), enabling few-shot or even zero-shot learning. |
| Approach: | They propose a black-box prompt search method that clusters and prunes the search space to focus exclusively on influential prompt tokens. |
| Outcome: | The proposed method achieves state-of-the-art performance across tasks and LLMs while significantly reducing search costs. |
Copied to clipboard
| Challenge: | Existing black-box-like deep learning methods for depression detection focus on improving classification performance, but it is impossible to explain and interpret those models that rely on state-of-the-art (SOTA) deep learning techniques. |
| Approach: | They propose to use hierarchical attention mechanisms and feed-forward neural networks to encode a model for depression detection on Twitter that leverages metaphorical concept mappings as input. |
| Outcome: | The proposed model leverages metaphorical concept mappings as input to detect depressed individuals and identify features of such users’ tweets. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have revolutionized the field of natural language processing and artificial intelligence, creating new SOTAs and reaching human-level language understanding performance on a series of tasks and benchmarks. |
| Approach: | They propose to use an algorithm test set sourced from Introduction to Algorithm to assess LLMs' code execution abilities. |
| Outcome: | The proposed model can execute programs described in natural language as long as no heavy numeric computation is involved. |
Copied to clipboard
| Challenge: | First-order logic (FOL) is often used to represent logical entailment, but determining natural language (NL) enanglement using FOL remains a challenge. |
| Approach: | They propose an Entailment-Preserving FOL representations task and a method which trains an NL-to-FOL translator by using the natural language entailment labels as verifiable rewards. |
| Outcome: | The proposed method achieves 1.8–2.7% improvement in EPR and 17.4–20.6% increase in E PR@16 compared to baselines in three datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) generate coherent, human-like text at scale, but raises concerns about authenticity and trust. |
| Approach: | They propose a threat of watermark spoofing that allows a malicious model to generate text containing the authentic-looking watermark of a trusted, victim model. |
| Outcome: | The proposed attack repurposes watermark radioactivity from a discoverable trait into an attack vector and replicates it. |
Copied to clipboard
| Challenge: | Deep learning models lacking interpretability and interactivity, authors say . lack of interactive mechanisms prevents clinicians from incorporating their own knowledge into decision-making process. |
| Approach: | a new deep learning model is proposed to improve interpretability and interactivity . authors propose a knowledge-enhanced agent-driven causal discovery framework . |
| Outcome: | a new model improves interpretability and interactivity on EHR data . the proposed model improve interpretability through explicit reasoning and causal analysis . |
Copied to clipboard
| Challenge: | Existing datasets for event understanding have limited coverage due to complexity of tasks. |
| Approach: | They propose a dataset that augments MAVEN datasets with event argument annotations . they propose 98,591 events and 290,613 arguments obtained with laborious human annotation . |
| Outcome: | The proposed dataset is the first all-in-one dataset supporting event detection, event argument extraction, and event relation extraction. |
Copied to clipboard
| Challenge: | Event schemas encode knowledge of stereotypical structures of events and their connections . previous work on event schema induction focuses on atomic events or linear temporal sequences . |
| Approach: | They propose a Temporal Complex Event Schema: a graph-based schema representation that encompasses events, arguments, temporal connections and argument relations. |
| Outcome: | The proposed model outperforms existing models on HITS@1 by 17.8%. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) achieve high accuracy on established Classical Chinese Poetry benchmarks, but it remains challenging to distinguish transferable Linguistic-Aesthetic Reasoning from reliance on familiar pre-training patterns. |
| Approach: | They propose a benchmark that combines a constructionist Out-of-Sample dataset with reverse understanding probes to evaluate large-scale large-format models. |
| Outcome: | The proposed model performs well on classical Chinese poetry benchmarks, but a performance gap persists . the model can complete famous couplets and can be used to understand a variety of texts. |
Copied to clipboard
| Challenge: | Existing tasks to assess LMs’ efficacy as KBs do not adequately consider multiple large-scale updates. |
| Approach: | They propose a task where multiple large-scale updates are made to language models and plug-in modules are used to handle the updates. |
| Outcome: | The proposed method outperforms existing methods on zsRE QA and NQ datasets and is 4x more effective in terms of updates/forgets ratio compared to a fine-tuning baseline. |
Copied to clipboard
| Challenge: | Existing methods for learning continual tasks do not cache history data, which makes the problem more challenging. |
| Approach: | They propose a method that allocates a small portion of private parameters and learns them with a shared pre-trained model. |
| Outcome: | The proposed method is comparable to existing methods and comparable to those using historical data. |
Copied to clipboard
| Challenge: | Recent advances in natural language processing have demonstrated societal bias in existing NLP models. |
| Approach: | They propose to use contrastive learning to learn fair representations for text classification . they conduct experiments on two text datasets to demonstrate their methods are stable . |
| Outcome: | The proposed methods balancing task performance and bias mitigation are stable in different hyperparameter settings. |
Copied to clipboard
| Challenge: | Existing methods for textual and structural retrieval ignore mutual reinforcement and only use structural retrievals for text-rich Graph Knowledge Bases (TG-KBs). |
| Approach: | They propose a Mixture of Structural-and-Textual Retrieval to retrieve textual and structural knowledge via a Planning-Reasoning-Organizing framework. |
| Outcome: | Experiments show that the proposed framework performs better than existing methods in analyzing TG-KBs and integrating structural trajectories for candidate reranking. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can generate the same sequences contained in the pre-train corpus, known as memorization. |
| Approach: | They analyze the relationship between memorization and outputs from Large Language Models (LLMs) they show a sudden drop and increase in the frequency of input tokens when generating memorized/unmemorized sequences . |
| Outcome: | The proposed model can generate the same sequences contained in the pre-train corpus, and it can predict unmemorized tokens. |
Copied to clipboard
| Challenge: | Existing approaches to multimodal affective computing learn spurious correlations from training data rather than genuine causal relationships, harming generalization under distribution shifts or noisy modalities. |
| Approach: | They propose a causal modality-invariant representation framework that separates each modality into ‘causal invariant’ and ‘environment-specific spurious representation’ from a modal inference perspective. |
| Outcome: | Experiments on multiple multimodal benchmarks show that the proposed framework achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Large Language Models (LMMs) struggle with simple tasks such as geometry, e.g., arithmetic, and reasoning. |
| Approach: | They propose to leverage code as supervision for cross-modal alignment . they propose to use FigCodifier and ImgCode-8.6M to synthesize novel mathematical figures . |
| Outcome: | The proposed model surpasses GPT-4o and Claude 3.5 Sonnet in the geometry problem-solving subset of MathVista, achieving improvements of 8.9% and 9.2%. |
Copied to clipboard
| Challenge: | Large Language Models exhibit strong capabilities in single-turn instruction following but suffer from Lost-in-Conversation (LiC) when instructions are revealed progressively in multi-turn settings, models get "Lost in Conversation" |
| Approach: | They propose a framework that encourages models to generate correct answers and judge solvability in multi-turn conversations. |
| Outcome: | The proposed framework improves models' ability to balance problem-solving with abstention . it reduces premature answering behaviors that cause lost-in-conversation (LiC) |
Copied to clipboard
| Challenge: | Existing studies have explored how LLMs handle positional relevance, but how they handle it remains unexplored. |
| Approach: | They propose to enforce certain computational mechanisms to allow for the tolerance in position perturbations in large language models (LLMs) they also find a pattern in intermediate features that allows this effect to be observed . |
| Outcome: | The proposed models can understand text with position perturbations and generalize to longer sequences than those seen during training with the latest techniques. |
Copied to clipboard
| Challenge: | Despite advances in artificial intelligence, building social intelligence remains a challenge. |
| Approach: | They propose a task to explain why people laugh in a video and a dataset to do this. |
| Outcome: | The proposed dataset generates plausible explanations for laughter in video and in-the-wild videos. |
Copied to clipboard
| Challenge: | Existing studies have raised concerns about data contamination from psychometric inventories . however, there is no systematic attempt to quantify the extent of data contamination . |
| Approach: | They propose a framework to measure data contamination in psychometric evaluations of Large Language Models by item memorization, evaluation memorisation and target score matching. |
| Outcome: | The proposed framework evaluates item memorization, evaluation memorisation, and target score matching in 21 models from major families and four widely used psychometric inventories. |
Copied to clipboard
| Challenge: | Existing methods for temporal knowledge graphs de-emphasize temporal correlations between facts sequences and ignore inferring clues from missing facts. |
| Approach: | They propose a Temporal PAth-based reasoning model that is robust to ambiguous temporal data. |
| Outcome: | The proposed model outperforms SOTA methods on the link prediction task. |
Copied to clipboard
| Challenge: | Information extraction suffers from its varying targets, heterogeneous structures, and demand-specific schemas. |
| Approach: | They propose a unified text-to-structure generation framework, namely UIE, which can universally model different IE tasks, adaptively generate targeted structures, and collaboratively learn general IE abilities from different knowledge sources. |
| Outcome: | The proposed framework can model different IE tasks, generate targeted structures, and learn general IE abilities from different knowledge sources. |
Copied to clipboard
| Challenge: | Multi-task learning is an inductive transfer mechanism that leverages information from related tasks to improve the primary model's generalization performance. |
| Approach: | They propose a multitask learning pipeline that finds relevant auxiliary tasks and learns their mixing ratio. |
| Outcome: | The proposed model can find relevant auxiliary tasks and learn their mixing ratio . the proposed model achieves significant performance boosts on several primary tasks . |
Copied to clipboard
| Challenge: | Existing presentation agents rely on predefined workflows and fixed templates to generate presentations. |
| Approach: | They propose an agentic framework that adapts to diverse user intents and iterative refinement based on observation. |
| Outcome: | The proposed framework can be used to generate presentations with environmental observations. |
Copied to clipboard
| Challenge: | High-quality, complex question-answer pairs are pivotal for training and evaluating capable deep search agents. |
| Approach: | They propose a pipeline that generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level. |
| Outcome: | The proposed pipeline generates high-quality, difficulty-controlled deep search question-answer pairs for a given corpus and a target difficulty level. |
Copied to clipboard
| Challenge: | Existing models for fake news detection are often insufficient or lacking in features . a novel structure-aware multi-head attention network can detect fake news in 4 hours . |
| Approach: | They propose a structure-aware multi-head attention network to detect fake news in mass news . they use credibility of publishers and users as prior weakly supervised information . |
| Outcome: | The proposed model can detect fake news in 4 hours with an accuracy of over 91% . the proposed model is faster than the state-of-the-art models . |
Copied to clipboard
| Challenge: | Existing efforts to generate Wikipedia articles for new events fall short of real-world application. |
| Approach: | They propose a benchmark to generate Wikipedia articles for new events under real-world scenarios . they use systematic metrics and LLM-based metrics to assess verifiability, organization, and other aspects aligned with real-life scenarios. |
| Outcome: | The proposed benchmarks show that hierarchical-based methods generate more comprehensive content while fine-tuned methods achieve better verifiability. |
Copied to clipboard
| Challenge: | Existing systems for knowledge extraction from natural language sentences are lacking for all languages. |
| Approach: | They propose a Korean knowledge extraction system and web interface for enriching a KBox knowledge base based on the Korean DBpedia. |
| Outcome: | The proposed system can extract factual knowledge from natural language sentences . the endpoint can be used to add knowledge to a KBox knowledge base anytime and anywhere . |
Copied to clipboard
| Challenge: | Existing studies have found that the test loss of LLMs scales as power-laws with model size, computational budget, and dataset size. |
| Approach: | They propose a concept of Temporal Scaling Law to study test loss of LLMs . they break down test loss into fine-grained token positions and develop a dynamic hyperbolic-law . |
| Outcome: | The proposed model predicts the test loss of LLMs as the training steps scale up. |
Copied to clipboard
| Challenge: | Existing studies focus on developing models that exploit the unification of multiple modalities. |
| Approach: | They propose to maintain modality independence by using a multi-modal transformer model that fuses all modalities. |
| Outcome: | The proposed model outperforms state-of-the-art models in multi-modal emotion recognition. |
Copied to clipboard
| Challenge: | Current PP methods face severe bottlenecks, including pipeline bubbles and memory footprint. |
| Approach: | They propose a sequence-level one-forward-one-backward (1F1B) PP method for training LLMs on long sequences with high throughput and memory efficiency. |
| Outcome: | The proposed method achieves 1.14X training throughput with half memory footprint compared to baseline methods . it trains an LLM with 30B parameters on sequences up to 64k tokens using 64X NVIDIA A100 GPUs . |
Copied to clipboard
| Challenge: | Existing methods to reduce bias have been shown to be effective over real-world datasets. |
| Approach: | They propose two new training objectives which directly optimise for the widely-used criterion of equal opportunity. |
| Outcome: | The proposed training objectives directly optimise for the widely-used criterion of equal opportunity while maintaining high performance over two classification tasks. |
Copied to clipboard
| Challenge: | Text-Attributed Graphs (TAGs) are widely used in the real world. |
| Approach: | They propose to use Large Language Models to generate OOD-nodes with high quality . they also use LLMs to integrate existing nodes with LLM-generated edges . |
| Outcome: | The proposed method performs well on samples outside the In-Distribution (ID) data, but it is difficult to obtain high-quality OOD samples in the real world. |
Copied to clipboard
| Challenge: | Existing role-play fine-tuning techniques improve role adaptability but may degrade safety performance, especially for villainous characters. |
| Approach: | They propose safety-aware Role-Play Fine-Tuning (SaRFT) to balance role-playing capabilities and safety. |
| Outcome: | The proposed method outperforms state-of-the-art baselines under both LoRA and full-parameter fine-tuning settings. |
Copied to clipboard
| Challenge: | Modern large-scale Pre-trained Language Models focus on text reconstruction, but have not sought to learn latent-level interpretable representations of sentences. |
| Approach: | They propose a new pre-training objective that enables the model to learn latent types . the objective allows the model a self-supervised way to extract sentence-level keywords . |
| Outcome: | The proposed model learns interpretable latent type categories without external knowledge and improves downstream tasks. |
Copied to clipboard
| Challenge: | Late-interaction based multi-vector retrieval systems rely on a naive summation of token-level similarity scores . this leads to inaccurate relevance estimation due to tokenization of semantic units and the influence of low-content words. |
| Approach: | They propose a late-interaction-based multi-vector retrieval system that uses token relations and token importance in relevance scoring. |
| Outcome: | Extensive tests show that TRIAL achieves state-of-the-art accuracy compared to existing methods. |
Copied to clipboard
| Challenge: | LoRA-Flow uses lightweight modules to customize large language models for downstream tasks . previous work on LoRA combination relied on task-level weights for each involved LoRA . |
| Approach: | They propose a LoRA-Flow approach that uses dynamic weights to adjust the impact of different LoRAs. |
| Outcome: | The proposed method outperforms baselines with task-level weights on six generative tasks. |
Copied to clipboard
| Challenge: | Existing models for enhancing knowledge updating are prone to performance degradation due to incomplete knowledge preservation mechanisms. |
| Approach: | They propose a model for locate-then-edit that decomposes long-term constrained programming into tractable stepwise subproblems for efficient solving. |
| Outcome: | The proposed framework achieves asymptotic optimal editing performance while meeting the constraints of long-term knowledge preservation. |
Copied to clipboard
| Challenge: | Existing methods for extracting relational facts from text have been successful . but with explosion of Web text, human knowledge is increasing drastically . |
| Approach: | They propose to improve relation extraction methods to extract relational facts from text . they analyze existing methods and show promising directions towards more powerful RE . |
| Outcome: | The proposed methods can extract relational facts from text, but they are still lacking in the current field. |
Copied to clipboard
| Challenge: | Existing benchmarks for large language models (LLMs) fail to capture these dynamics, focusing on static, open-ended evaluations. |
| Approach: | They propose a benchmark to assess lifelong learning in large language models . they use two episodic datasets rich in narrative structure and character interactions . |
| Outcome: | Experiments on LLMs show that non-parametric methods outperform parametric ones in managing stateful learning. |
Copied to clipboard
| Challenge: | Existing methods for learning natural language understanding are limited in low-resource settings. |
| Approach: | They propose to use rules of grammar to construct and expand rules of grammatical structure of data without human involvement. |
| Outcome: | The proposed approach outperforms state-of-the-art methods in three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to feature engineering relied on domain expertise to build features. |
| Approach: | They propose a framework that leverages large language models to automatically construct features in a string format and generate semantic explanations based on dataset descriptions. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on real-world datasets. |
Copied to clipboard
| Challenge: | Existing large language models (LLMs) have a tendency to hallucinate and provide creative and fluent responses that are not factually accurate. |
| Approach: | They propose a tool that automatically extracts factual claims from text, gathers evidence from external knowledge sources, evaluates the factuality of each claim, and suggests revisions for identified errors. |
| Outcome: | The proposed tool detects errors in text and evaluates their factuality and suggests revisions based on the collected evidence. |
Copied to clipboard
| Challenge: | Multi-hop question answering (MQA) is one of the challenging tasks to evaluate machine’s comprehension and reasoning abilities, where large language models (LLMs) have widely achieved the human-comparable performance. |
| Approach: | They propose a framework to edit multi-hop question models to update model with up-to-date facts while avoiding expensive re-training or fine-tuning. |
| Outcome: | The proposed framework outperforms all competitors in multi-hop question answering tasks and consistently produces reliable reasoning process. |
Copied to clipboard
| Challenge: | Existing methods for detecting hate speech ignore misalignment and uncertainty between modalities . social media platforms have become conduits for the rapid dissemination of hate speech . |
| Approach: | They propose an uncertainty-aware cross-modal alignment framework for hate speech detection that minimizes the misalignment of image and text in memes. |
| Outcome: | The proposed framework produces a competitive performance compared with existing methods. |
Copied to clipboard
| Challenge: | Existing multi-modal large language models (MLLMs) are able to process visual inputs by converting them into visual tokens that share the same latent space as language tokens in LLMs. |
| Approach: | They propose a benchmark that assesses the visual illusion level given spurious images and a pipeline that converts visual inputs into visual tokens. |
| Outcome: | The proposed benchmark shows that MLLMs suffer from an instinctive bias to varying degrees when presented with spurious images. |
Copied to clipboard
| Challenge: | Existing evaluation methodologies for MWPs diverge from human judgment and face challenges in recognizing mathematically equivalent answers. |
| Approach: | They propose an evaluation metric rooted in graph edit distance that features benefits such as permutation invariance and more accurate program equivalence identification. |
| Outcome: | The proposed evaluation metric features benefits such as permutation invariance and more accurate program equivalence identification. |
Copied to clipboard
| Challenge: | SylloBase is a benchmark for syllogistic reasoning, a critical capability widely required in natural language understanding tasks, such as text entailment and question answering. |
| Approach: | They propose to use a benchmark to learn syllogistic reasoning on a set of templates and to use them to generate and understand slogisms. |
| Outcome: | The proposed benchmark covers a complete taxonomy of syllogism reasoning patterns, and contains both automatically and manually constructed samples. |
Copied to clipboard
| Challenge: | Language models (LMs) have demonstrated impressive reasoning capabilities across domains . but their ability to handle PSPACE-complete problems remains underexplored . a new benchmark for regex minimization is proposed to evaluate LMs' reasoning capabilities . |
| Approach: | They propose a benchmark for regex minimization to evaluate LMs' reasoning power . they use a million regexes paired with their minimal equivalents to evaluate their performance . |
| Outcome: | The proposed model can solve NP-complete problems, but their ability to handle PSPACE-complete ones remains underexplored. |
Copied to clipboard
| Challenge: | Recent work has shown that statistical language modeling with transformers can greatly improve the performance in code completion tasks. |
| Approach: | They propose a retrieval-augmented code completion framework that combines a source code retriever and an auto-regressive language model for programming language. |
| Outcome: | The proposed framework achieves state-of-the-art on CodeXGLUE benchmark. |
Copied to clipboard
| Challenge: | Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words. |
| Approach: | They propose to use semantic role labeling to provide additional guidance for multi-turn dialogue rewriting models. |
| Outcome: | The proposed model outperforms existing models on multi-turn dialogue rewriting tasks. |
Copied to clipboard
| Challenge: | Existing methods to detect mental disorders focus on the presence of symptoms, but the context of symptoms is often ignored, leading to errors in symptom identification. |
| Approach: | They propose to use large language models to extract contextual information while introducing an uncertainty-aware decision fusion network that combines predictions of multiple models based on quantified uncertainty values. |
| Outcome: | The proposed model detects mental disorders even in situations where symptom information is incomplete. |
Copied to clipboard
| Challenge: | Existing approaches to meeting summarization are limited due to noise, lengthy transcripts, and scattered salient information. |
| Approach: | They propose a two-step framework for meeting summarization that leverages a self-supervised paradigm to reconstruct transcripts and a relative positional bucketing algorithm to equip models to generate the summary. |
| Outcome: | The proposed method significantly reduces memory consumption and processing time on two meeting summarization datasets. |
Copied to clipboard
| Challenge: | Current mitigation strategies fail to preserve contextual reasoning capabilities in risky scenarios, leading to systemic risks for legal compliance. |
| Approach: | They propose to use reinforcement learning with a rule-based reward to incentivize contextual reasoning capabilities while enhancing compliance with safety and privacy norms. |
| Outcome: | The proposed model outperforms Qwen2.5-7B-Instruct model in safety and privacy benchmarks and achieves +8.58% accuracy improvement. |
Copied to clipboard
| Challenge: | Existing studies focus on auto-generated syntactic knowledge to enhance semantic role labeling . experimental results show that map memories can enhance SRL . |
| Approach: | They propose to map memories to enhance semantic role labeling by encoding auto-generated syntactic knowledge from off-the-shelf toolkits. |
| Outcome: | The proposed model outperforms baselines and achieves state-of-the-art results on two English benchmark datasets. |
Copied to clipboard
| Challenge: | Recent sparsity-aware binarization approaches can achieve sub-1-bit compression, but they face performance degradation, mask-management overhead, and limited hardware compatibility. |
| Approach: | They propose a binary quantization framework that leverages binary pattern clustering and weight transformation to overcome performance degradation and mask-management overhead. |
| Outcome: | The proposed framework achieves state-of-the-art compression (1.11–0.7 bits) it maintains high performance with only a 3.1% accuracy drop in zero-shot benchmarks while delivering a 1.6 speedup over FP16. |
Copied to clipboard
| Challenge: | Latent variable models for text capture global semantic and syntactic features when trained correctly. |
| Approach: | They propose a short run dynamics for inference that initializes from the prior distribution of the latent variable and runs a small number of Langevin dynamics steps guided by its posterior distribution. |
| Outcome: | The proposed model is able to generate coherent sentences with smooth transition and shows no sign of posterior collapse. |
Copied to clipboard
| Challenge: | Existing safety alignment methods leave Large Language Models vulnerable to sophisticated jailbreak attacks. |
| Approach: | They propose a safety reasoning internalization framework that internalizes safety reasoning into an implicit computational pathway using Low-Rank Adaptation (LoRA). |
| Outcome: | The proposed framework achieves a 43% lower Attack Success Rate (ASR) against distinct jailbreak attacks compared to strong baselines. |
Copied to clipboard
| Challenge: | Vision-language models often generate excessive visual tokens, leading to poor performance . a novel training-free visual token pruning method is proposed to improve performance despite the computational cost associated with VLMs. |
| Approach: | They propose a training-free visual token pruning method that reduces biased token pruning . they plan to open-source the code upon publication . |
| Outcome: | The proposed method reduces biased token pruning and enhances model robustness with limited visual token budget. |
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning with verifiable rewards (RLVR) rely on static objective functions and rigid clipping strategies that misalign with the model’s evolving reasoning capabilities. |
| Approach: | They propose to incorporate Power-Mean Policy Optimization (PMPO) and Feedback-Adaptive Clipping (FAC) to overcome limitations of static mechanisms. |
| Outcome: | Extensive experiments on nine reasoning tasks show the proposed paradigm outperforms state-of-the-art methods. |
Copied to clipboard
| Challenge: | Current AI-powered code assistance tools struggle with ambiguous problem statements . failures on such ambiguously requests are highly correlated with longer trajectories . |
| Approach: | They propose a contextual query refinement approach that transforms ambiguous user requests into comprehensive, actionable problem statements through lightweight pre-exploration of the target codebase. |
| Outcome: | Empirical results show that CodeScout improves resolution rates with 27 additional issues resolved compared to baseline method. |
Copied to clipboard
| Challenge: | Few-shot named entity recognition (NER) aims to identify entities of target types with limited number of illustrative instances. |
| Approach: | They propose a superposition concept discriminator which solves the intrinsic generalization problem by an active learning paradigm. |
| Outcome: | The proposed model significantly improves few-shot named entity recognition (FS-NER) with minimal additional efforts. |
Copied to clipboard
| Challenge: | Existing bilingual or multi-lingual MWE corpora are limited for multilingual use . only 871 pairs of English-German MWEs are available for research . |
| Approach: | They present a collection of bilingual and multi-lingual MWEs extracted from parallel corpora. |
| Outcome: | The available bilingual or multi-lingual MWE corpus is very limited . the collection is a small collection of 871 pairs of English-German MWEs . |
Copied to clipboard
| Challenge: | Existing efforts to improve reasoning efficiency of large language models focus on modifying the reinforcement learning reward, such as adding length penalties. |
| Approach: | They propose a training framework that elicits efficient reasoning through reasoning vectors and a framework that allows the model to generate high-quality responses during reinforcement learning. |
| Outcome: | The proposed framework reduces reasoning length by 30% while maintaining stability, while retaining high accuracy. |
Copied to clipboard
| Challenge: | Existing methods to generate medical records using Causal Language Modelling are limited due to privacy concerns. |
| Approach: | They propose a method for generating medical records using Masked Language Modelling using Causal language models. |
| Outcome: | The proposed method produces high-quality synthetic data with a re-identification risk of only 3.5% and a patient recall of 96%. |
Copied to clipboard
| Challenge: | Existing unsupervised reinforcement learning methods lack the capacity to adapt to the model’s evolving reasoning capabilities during training. |
| Approach: | They propose an unsupervised reinforcement learning algorithm that adapts rewards to balance consensus and exploration based on the Free Energy Principle. |
| Outcome: | Empirical evaluations on nine datasets show that FREIA outperforms baseline methods on reasoning tasks. |
Copied to clipboard
| Challenge: | Existing offline preference optimization methods rely on preference labels to optimize large language models. |
| Approach: | They propose an offline method for enhancing large language models in reasoning tasks that utilizes value signals at individual reasoning steps. |
| Outcome: | The proposed framework outperforms offline preference optimization techniques by 4% to 6% on math reasoning, commonsense reasoning, and coding tasks. |
Copied to clipboard
| Challenge: | Existing approaches to implicit discourse relation recognition lack connectives as strong linguistic clues. |
| Approach: | They propose a transS-driven joint learning architecture to translate discourse relations in low-dimensional embedding space and exploit the semantic features of arguments to assist discourse understanding. |
| Outcome: | The proposed model outperforms existing systems on the Penn Discourse TreeBank. |
Copied to clipboard
| Challenge: | Existing comparative summarization methods focus on surface-level semantic differences, which may not capture the most relevant distinctions. |
| Approach: | They propose a framework which transforms scientific papers into LLM personas that debate their respective novelties. |
| Outcome: | The proposed framework generates informative arguments and effectively contrasts papers, and supports researchers in their literature review. |
Copied to clipboard
| Challenge: | Detecting and identifying events is an important subtask of event extraction. |
| Approach: | They build a large event-related candidate set with good coverage and apply an adversarial training mechanism to iteratively identify informative instances from the candidate set and filter out those noisy ones. |
| Outcome: | The proposed method significantly outperforms the state-of-the-art methods on two real-world datasets. |
Copied to clipboard
| Challenge: | Reaction Miner is a system designed to extract chemical reactions from raw scientific PDFs. |
| Approach: | They propose a system that extracts chemical reactions directly from raw scientific PDFs. |
| Outcome: | The proposed system can extract chemical reactions from raw scientific PDFs. |
Copied to clipboard
| Challenge: | Existing studies have failed to assess RAG leakage risks for large language models . constructing and maintaining highquality RAG knowledge databases has become increasingly costly . |
| Approach: | They propose a framework for controlled evaluation of RAG leakage using query generation and adversarial instructions. |
| Outcome: | The proposed framework compares six existing attacks across fourteen LLMs, four datasets, and diverse RAG systems. |
Copied to clipboard
| Challenge: | Figure 1 illustrates the challenges of monolingual word alignment. |
| Approach: | They propose to use the family of optimal transport (OT) to achieve unbalanced word alignment that values alignment and null alignment on unsupervised datasets. |
| Outcome: | The proposed methods are competitive against the state-of-the-art methods on challenging datasets with high null alignment frequencies. |
Copied to clipboard
| Challenge: | Existing methods for text style transfer only focus on the transformation between styles, yet they do not take into account that this transformation can be achieved via different hidden transfer patterns. |
| Approach: | They propose a novel approach which automatically mines hidden transfer patterns to improve TST . they use a clustering module to automatically discover hidden transfer pattern from the data . |
| Outcome: | The proposed method can be applied in a plug-and-play manner to enhance other methods to further improve their performance. |
Copied to clipboard
| Challenge: | Existing VG models make unrealistic assumptions about how to ground video segments . a recent study has shown that video grounding can be useful for downstream applications . |
| Approach: | They propose a new task: Weakly-Supervised temporal Article Grounding (WSAG) given an article and a relevant video, WSAG aims to localize all "groundable" sentences to the video. |
| Outcome: | The proposed method is simple but effective, and it can be used in real-world applications. |
Copied to clipboard
| Challenge: | Medical Vision-Language Models are predominantly trained on professional literature, limiting their ability to communicate findings in the lay register required for patient-centered care. |
| Approach: | They propose a multimodal benchmark dedicated to expert-lay semantic alignment that enforces strict semantic equivalence by integrating unified medical language system (UMS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. |
| Outcome: | The proposed benchmark enforces strict semantic equivalence by integrating unified medical language system (UMLS) Concept Unique Identifiers (CUIs) with micro-level entity constraints. |
Copied to clipboard
| Challenge: | Existing classification models only consider the temporal variations of existing data . current models focus on English corpora, leaving time as domains unexplored . |
| Approach: | They propose a framework to generalize classifiers over time on four languages, English, Danish, French, and German. |
| Outcome: | The proposed framework can generalize classifiers over time on four languages, English, Danish, French, and German. |
Copied to clipboard
| Challenge: | Existing approaches to search for images using single-modality are limited by representation space fragmentation. |
| Approach: | They propose a unified representation framework that achieves efficient query-target alignment . they introduce a multi-level Chain-of-Thought prompting strategy that guides MLMs to generate discriminative, semantically compatible captions for target images . |
| Outcome: | The proposed framework achieves efficient query-target alignment through synergistic components. |
Copied to clipboard
| Challenge: | Recent advances in prompt engineering have created impediments for end users to adopt . however, prompt engineering remains an impedance due to rapid advances in models, tasks, and associated best practices. |
| Approach: | They propose to define APO as a 5-part unifying framework and categorize all relevant works based on their salient features. |
| Outcome: | The proposed framework aims to improve the performance of large language models on various tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are evolving towards autonomous agents . retrieval capabilities are well-benchmarked, but post-retrieval synthesis is under-evaluated due to open-ended writing. |
| Approach: | They propose a benchmark to evaluate information consolidation capabilities using survey papers as gold standards. |
| Outcome: | The proposed benchmark analyzes the post-retrieval synthesis stage of large language models . it leverages high-quality survey papers as gold standards and reverse-engineers research requests . the proposed benchmark outperforms single-turn generation and reduces hallucinations . |
Copied to clipboard
| Challenge: | Existing models for fine-grained speaking styles are limited in terms of accuracy, coverage, and naturalness. |
| Approach: | They propose a model that pre-trains with coarse captions and annotates with a pipeline that grounds captions in audio. |
| Outcome: | The proposed model outperforms existing models with fine-grained style annotations . it integrates global and fine-granular supervision, enabling unified representations based on the proposed model . |
Copied to clipboard
| Challenge: | Parallel Coordinated Reasoning (PaCoRe) overcomes a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. |
| Approach: | They propose a training-and-inference framework to overcome a central limitation of language models: their inability to scale test-time compute (TTC) under a fixed context window. |
| Outcome: | The proposed model scales to multi-million-token effective TTC without exceeding context limits. |
Copied to clipboard
| Challenge: | Existing paper search systems lack detailed information to support finer-grained queries. |
| Approach: | They propose a paper-based index that transforms abstract-based corpus index into hierarchical index tree and offline can support paper search queries. |
| Outcome: | The proposed system achieves the SOTA performance and excels in fine-grained scenarios. |
Copied to clipboard
| Challenge: | Document-level relation extraction (DocRE) aims to extract semantic relations among entity pairs in a document. |
| Approach: | They propose an evidence-enhanced framework that empowers document-level relation extraction (DocRE) Eider efficiently extracts evidence and effectively fuses extracted evidence in inference. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on three benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods on understanding the capabilities of LLMs in logical reasoning rely on binary entailment classification or synthetically derived rationales. |
| Approach: | They propose to annotate a human-annotated dataset consisting of diverse and complex reasoning chains for a set of realistic logical reasoning stories also written by humans. |
| Outcome: | The proposed model outperforms existing methods on understanding the capabilities of LLMs in logical reasoning by 10% or more. |
Copied to clipboard
| Challenge: | a new pipeline for personality-based synthetic dialogues is being developed in Korea . a dataset curated by large language models is needed to generate human-like dialogues . |
| Approach: | They propose a personality-based synthetic dialogue data pipeline to elicit responses from large language models via prompting. |
| Outcome: | The proposed pipeline generates human-like dialogues considering real-world scenarios when users engage with chatbots. |
Copied to clipboard
| Challenge: | Existing datasets in operations research domain lack detailed annotations of the modeling process, focusing only on objective values. |
| Approach: | They propose an annotation-based tree-of-thought tree-based reasoning algorithm that integrates reinforcement learning into a tree- of-though. |
| Outcome: | The proposed algorithm outperforms state-of-the-art methods on StructuredOR, NL4OPT, and MAMO-ComplexLP datasets. |
Copied to clipboard
| Challenge: | Existing methods for calibration of large reasoning models (LRMs) focus on clean inputs, leaving noise unexplored. |
| Approach: | They propose a confidence calibration framework for character-level noisy inputs that extracts uncertainty signals from both the empirical answer distribution and the model’s predictive distribution and integrates them via a learned calibrator. |
| Outcome: | Experiments on multiple mathematical reasoning benchmarks show that DisCal outperforms existing calibration methods under noisy inputs, reducing expected calibration error (ECE) by up to 39.21% and improving Area Under the Receiver Operating Characteristic Curve (AUROC) by 31.44%. |
Copied to clipboard
| Challenge: | Experimental results show that fine-tuning of large language models for specific tasks can be challenging . distribution shift during fine-timing can lead to performance degradation in general task capabilities . |
| Approach: | They propose a new approach that bridges the distribution gap between task datasets and LLMs by guiding fine-tuning with a distilled dataset generated by the model itself. |
| Outcome: | The proposed approach achieves comparable or superior performance on downstream tasks compared to the vanilla approach. |
Copied to clipboard
| Challenge: | Experimental results show that Synchronous Semantic Decoding (SSD) can achieve state-of-the-art unsupervised semantic parsing performance on multiple datasets. |
| Approach: | They propose an unsupervised method which solves the semantic gap and the structure gap by leveraging paraphrasing and grammar-constrained decoding. |
| Outcome: | The proposed method can solve the semantic gap and structure gap on multiple datasets. |
Copied to clipboard
| Challenge: | Text embedding models show strong performance on generic benchmarks, but their effectiveness diminishes when applied to private datasets. |
| Approach: | They propose a method for adapting general-purpose text embedding models to private datasets . they construct supervisory signals from the ranking of keyword-based retrieval results . |
| Outcome: | The proposed method improves retrieval performance across domains, datasets, and models. |
Copied to clipboard
| Challenge: | Existing methods for detecting hallucination in long-form tasks focus on limited domains or rely heavily on external fact-checking tools, which may not always be available. |
| Approach: | They propose a new paradigm that augments fine-tuning with an auxiliary task for the model to jointly learn with the main task of hallucination detection. |
| Outcome: | The proposed method outperforms existing methods for detecting hallucination in open-domain long-form generation and is more accurate than random guessing. |
Copied to clipboard
| Challenge: | Current paradigms rely on holistic scoring and static leaderboards to disentangle fine-grained competencies. |
| Approach: | They propose a framework to shift the focus from ranking to fine-grained diagnosis. |
| Outcome: | The proposed framework surpasses the strongest baseline by 7.92%. |
Copied to clipboard
| Challenge: | Detection problems involving positive instances are often deficient in information extraction tasks . a number of researches have employed neural network models to solve detection problems . |
| Approach: | They propose an algorithm which can handle positive sparsity problem and directly optimize over F-measure . they borrow the idea of marginal utility from economics and propose a theoretical framework for instance importance measuring . |
| Outcome: | The proposed algorithm improves on positive sparsity problem and over F-measure . it leads to more effective and stable training of neural network based detection models. |
Copied to clipboard
| Challenge: | Recent studies have focused on instruction tuning to show cross-lingual generalization . a novel non-English meta-dataset is used to study instruction tuning . |
| Approach: | They perform instruction tuning individually for two distinct language meta-datasets and assess the performance on unseen tasks in a non-English language. |
| Outcome: | The proposed model outperforms baseline training in English and Korean by 20.7% and 13.6%. |
Copied to clipboard
| Challenge: | Existing methods focus on optimizing document features, overlooking the potential of high-quality label features to enhance classification performance. |
| Approach: | They propose a multi-label document classification paradigm that utilizes large language models to expand the label content and generate pseudo-samples for the tail categories. |
| Outcome: | The proposed method significantly outperforms state-of-the-art models. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are vulnerable to jailbreak attacks that exploit weaknesses in traditional safety alignment. |
| Approach: | They propose a framework that trains models to engage in explicit safe reasoning before response . they propose RATIONAL, which allows models to reject harmful prompts while providing meaningful and context-aware responses. |
| Outcome: | The proposed framework fine-tunes models to reason about query intent, ethics, and potential harm. |
Copied to clipboard
| Challenge: | Existing studies on semantic parsing use Maximum Likelihood Estimation (MLE) to train discriminative semantic parses. |
| Approach: | They propose a semantic-aware contrastive learning algorithm which can learn to distinguish fine-grained meaning representations and take the overall sequence-level semantic into consideration. |
| Outcome: | The proposed algorithm improves on two standard datasets and gets state-of-the-art performance over existing methods. |
Copied to clipboard
| Challenge: | Existing meta-learning models rely on implicit instance statistics and are unreliability and weak interpretability. |
| Approach: | They propose a meta-information guided meta-learning framework that uses semantics to guide meta- learning . experimental results demonstrate the effectiveness of the proposed framework . |
| Outcome: | The proposed framework can establish connections between instance-based information and semantic-based data, enabling faster initialization and adaptation. |
Copied to clipboard
| Challenge: | Existing methods for continual learning in language models suffer catastrophic forgetting when learning sequential tasks. |
| Approach: | They propose an orthogonal low-rank adaptation approach for continual learning in language models that uses orthogons to learn sequentially. |
| Outcome: | The proposed approach outperforms state-of-the-art methods on continual learning benchmarks and preserves generalization ability of LLMs on unseen tasks. |
Copied to clipboard
| Challenge: | Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and fail to generalize across unseen attacks. |
| Approach: | They propose a new approach that disrupts representations underlying harmful behaviors in Large Language Models by using loss-based fine-tuning. |
| Outcome: | The proposed approach outperforms existing methods such as Circuit Breaker, RMU, and NPO with 95% reduction in attack success rates across diverse jailbreak benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for semantic parsing are difficult to design and learn, especially in wideopen domains. |
| Approach: | They propose a neural semantic parsing approach which models semantic par- sing as an end-to-end semantic graph generation process. |
| Outcome: | The proposed model achieves state-of-the-art performance on Overnight dataset and gets competitive performance on Geo and Atis datasets. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax. |
| Approach: | They propose to use a model averaging technique to find the most powerful alignment-forging Pareto front among RLHF algorithms. |
| Outcome: | The proposed method achieves the strongest alignment-forging Pareto front among competing methods. |
Copied to clipboard
| Challenge: | Recent techniques employ pretrained language models to improve topic quality. |
| Approach: | They propose a topic-based model that uses contrastive learning and term weighting to learn from a pretrained language model and discover influential terms from semantically coherent clusters. |
| Outcome: | The proposed model outperforms baselines across multiple topic coherence measures and can be used as an add-on to existing topic models and improves their performance. |
Copied to clipboard
| Challenge: | Existing methods for fingerprinting model ownership traces are vulnerable to illegal plagiarism and are not reliable. |
| Approach: | They propose a rule-driven fingerprinting framework that encodes contextual correlations across multiple dialogue turns. |
| Outcome: | The proposed framework achieves stronger stealth and robustness than previous work. |
Copied to clipboard
| Challenge: | Neural machine translation (NMT) is pivotal for crosslingual conversation and trade . traditional solutions that penalize text redundancy or token reoccurrence have shown limited efficacy . |
| Approach: | They propose an algorithm that modulates suppression of tokens dynamically, informed by attention weights and inter-token distances. |
| Outcome: | The proposed algorithm outperforms existing methods in precision and generalizability. |
Copied to clipboard
| Challenge: | Existing approaches for few-shot text classification rely on exploitation of lexical features and distributional signatures on training data, while neglecting to strengthen the model's ability to adapt to new tasks. |
| Approach: | They propose a meta-learning framework integrated with an adversarial domain adaptation network to improve the model's adaptive ability and generate high-quality text embedding for new classes. |
| Outcome: | The proposed framework outperforms the state-of-the-art models on four datasets and shows clear superiority over existing models. |
Copied to clipboard
| Challenge: | Existing concept reasoning related datasets suffer from modeledge leakage and context leakage. |
| Approach: | They propose a concept reasoning for large language models with modeledge leakage prevention and context leakage preventive methods to improve the models' conceptual reasoning abilities. |
| Outcome: | The proposed method significantly improves the existing models and reasoning methods, achieving a 7% increase in accuracy compared to CoT and showing better granularity. |
Copied to clipboard
| Challenge: | Existing methods for product attribute value identification face critical challenges . seller-provided attribute values are often incomplete or inaccurate . |
| Approach: | They propose a retrieval-based method that uses taxonomy-aware contrastive learning . they use product profiles and candidate values to encode and retrieve attributes based on similarity . |
| Outcome: | The proposed method is based on a taxonomy-aware, hard negative sampling and adaptive inference with dynamic thresholds. |
Copied to clipboard
| Challenge: | Existing studies focus on matching candidate responses with every context utterance, but it also brings noise signals and unnecessary information. |
| Approach: | They propose a multi-hop selector network to match context with candidate responses . they propose to use a selector to filter the relevant utterances as context . |
| Outcome: | The proposed model outperforms state-of-the-art methods on three public multi-turn dialogue datasets. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit prompt leakage vulnerabilities, raising intellectual property and confidentiality concerns. |
| Approach: | They use probing techniques to capture LLMs’ intent-related internal representations and show that they internalize prompt leakage intents in their hidden states before generating tokens. |
| Outcome: | The proposed probes achieve 90%+ AUROC across all tested models, even when applied to new system prompts and attacks. |
Copied to clipboard
| Challenge: | Existing embedding-based retrieval systems rely on heuristic and suboptimal cutoffs for item retrieval. |
| Approach: | They propose a probabilistic Embedding-Based Retrieval framework that learns a shared semantic representation space for both queries and items. |
| Outcome: | The proposed framework improves retrieval precision and recall, and ablation studies show it captures the differences between head-to-tail queries. |
Copied to clipboard
| Challenge: | Existing approaches to review scientific papers are limited by their content or quality . SEA is a framework for automated scientific review, but its contents are generic or partial. |
| Approach: | They propose a framework for automated scientific review using large language models . they propose to use a standardized review dataset to fine-tune an LLM to generate high-quality reviews. |
| Outcome: | The proposed framework can generate high-quality reviews from standardized datasets and improves on the existing feedback mechanisms. |
Copied to clipboard
| Challenge: | Existing methods for chain-of-thought prompting rely on manual demonstrations . experimental results show that GCR outperforms baseline methods without performance degradation . |
| Approach: | They propose a method that uses random samples to generate demonstrations in zero-shot settings. |
| Outcome: | The proposed method outperforms baseline methods on ten datasets without demonstration bias. |
Copied to clipboard
| Challenge: | Existing methods to build named entity recognition systems with limited labeled data are lacking. |
| Approach: | They propose three orthogonal schemes to build named entity recognition systems when labeled data is limited. |
| Outcome: | The proposed NER systems outperform existing methods on few-shot and training-free settings. |
Copied to clipboard
| Challenge: | Large language models generate long and verbose reasoning traces at inference time . short context post-training alone induces substantial reasoning compression . |
| Approach: | They propose a step-level advantage selection approach that reduces reasoning length by over 30% . they propose to use GRPO without any length-aware objective to train models in a shorter context window . |
| Outcome: | The proposed approach reduces average reasoning length by over 30% while improving Pass@1 accuracy by 3.79 points over the strongest length-aware baseline. |
Copied to clipboard
| Challenge: | evaluating commonsense in dialogue systems remains an open challenge . despite the success of open-domain dialogue systems, systems struggle to produce commonsensical responses as humans do. |
| Approach: | They propose an event commonsense evaluation metric empowered by commonsensence knowledge bases. |
| Outcome: | The proposed metric achieves higher correlations with human judgments than baselines. |
Copied to clipboard
| Challenge: | Existing research has focused on post-training knowledge editing (KE) for language models to ensure that knowledge remains accurate and up-to-date. |
| Approach: | They propose to use a GradSim indicator to detect when and why updated knowledge ripples in language models. |
| Outcome: | The proposed indicator GradSim shows that LMs that fail to handle ripple effects have low GradSIM. |
Copied to clipboard
| Challenge: | Existing methods for dynamic web navigation rely on greedy strategies or value estimation, struggle to achieve effective backtracking and are heavily dependent on proprietary models. |
| Approach: | They propose a cognitive multi-agent collaboration framework that enhances cyberspace exploration capability through In-Context Exploration. |
| Outcome: | The proposed framework surpasses the proprietary model Claude-3.5 Sonnet on the WebArena benchmark. |
Copied to clipboard
| Challenge: | Current text classification methods require a large number of labeled documents as training data. |
| Approach: | They propose a model that uses only the label name of each class to train classification models on unlabeled data without using any labeled examples. |
| Outcome: | The proposed model achieves 90% accuracy on four benchmark datasets using label names as the only supervision . |
Copied to clipboard
| Challenge: | Existing approaches to optimize retrieval using search-only metrics ignore downstream utility and fine-tune entire LLM to jointly reason and retrieve limit retrieval utility and compatibility with frozen or proprietary models. |
| Approach: | They propose a lightweight, model-agnostic framework that decouples the searcher from the generator and trains the search user using a Gain Beyond RAG reward. |
| Outcome: | The proposed framework outperforms baselines trained on over 70 more data with 2.4k training samples. |
Copied to clipboard
| Challenge: | Recurrent exchange of model updates in FL can result in prohibitively high communication costs, hindering the distributed learning process. |
| Approach: | They propose a federated fine-tuning framework that uses a round-robin segment sharing scheme to reduce network bandwidth and adaptive sparsification methods tailored to LoRA’s training dynamics. |
| Outcome: | The proposed framework reduces communication overhead without compromising performance on question-answering and value-alignment tasks. |
Copied to clipboard
| Challenge: | Existing prompting methods that require white-box access to the model or substantial training fail to simultaneously lessen toxicity and bias. |
| Approach: | They propose a strategy that encourages LLMs to integrate diverse human perspectives and self-regulate their responses by incorporating diverse human viewpoints. |
| Outcome: | The proposed approach can significantly diminish toxicity (up to 89%) and bias (up 73%) in LLMs’ responses. |
Copied to clipboard
| Challenge: | Long video content understanding poses a challenging set of research questions as it involves long-distance, cross-media reasoning and knowledge awareness. |
| Approach: | They propose a framework which extracts events, entities, and relations from the rich multimedia content in long videos to pre-construct movie knowledge graphs. |
| Outcome: | The proposed framework performs competitively for both the new DeepMovieQA and the pre-existing MovieQA dataset. |
Copied to clipboard
| Challenge: | Recent approaches to fine-tuning of large language models suffer from task interference and catastrophic forgetting. |
| Approach: | They propose a fine-tuning framework that adapts isolation decisions based on online estimates of parameter importance. |
| Outcome: | The proposed framework reduces interference and forgetting while releasing outdated parameters to recover plasticity. |
Copied to clipboard
| Challenge: | Existing automated layout models are ill-suited for spreadsheets, authors say . existing layout models treat components as rectangles with continuous coordinates . authors: spreadsheets are powerful tools for organizing and analyzing data . |
| Approach: | They formalize a spreadsheet layout generation task and introduce a framework for spreadsheet layouts . they use multimodal large language models to combine rule and vision reflection . |
| Outcome: | The proposed framework outperforms baselines in a spreadsheet layout generation task by 22.6%. |
Copied to clipboard
| Challenge: | a lightweight module for tuning large multimodal models is introduced . CaMML integrates contextual samples into large models, enabling them to make inferences . |
| Approach: | They introduce a lightweight module for tuning large multimodal models . they have developed two models that have shown exceptional performance . |
| Outcome: | The proposed model outperforms LLaVA-1.5 on ten widely recognized datasets with a noticeable margin. |
Copied to clipboard
| Challenge: | Attention mechanism has been used in Vision-and-Language (VL) tasks to bridge the semantic gap between visual and textual clues. |
| Approach: | They conduct a comprehensive analysis on understanding the role of attention alignment by looking into attention score calculation methods and checking how it represents the visual region’s and textual token’s significance for the global assessment. |
| Outcome: | The attention score calculation methods represent visual region’s and textual token’s significance for the global assessment. |
Copied to clipboard
| Challenge: | Existing methods for debiasing use uniform bias corrections across all input queries . weak debiases retains bias in sensitive queries, while weak dealiases in biased ones . |
| Approach: | They propose a framework that selectively applies debiasing based on input sensitivity . RG-TTA adaptively triggers fairness regularization based upon bias sensitivity of each input . |
| Outcome: | Experiments show that debiasing improves zero-shot performance while maintaining fairness . weak debiased queries distort semantically meaningful information while weak ones fail to mitigate stereotypes . |
Copied to clipboard
| Challenge: | Existing knowledge injection frameworks focus on knowledge memorization and retrieval, but static nature of large language models leads to outdated information as the real world evolves or when adapting to domain-specific knowledge. |
| Approach: | They propose a four-tier knowledge injection framework that defines the levels of knowledge injection: memorization, retrieval, reasoning, and association. |
| Outcome: | The proposed framework defines the levels of knowledge injection: memorization, retrieval, reasoning, and association. |
Copied to clipboard
| Challenge: | Large language models can produce unreliable or misleading outputs, posing challenges for real-world applications. |
| Approach: | They employ an auxiliary LLM to analyze the patterns of disagreement among LLMs . they validate their framework on AmbigQA, OpenBookQA, and MMLU-Pro . |
| Outcome: | The proposed model can be used to diagnose uncertainty sources in a model with an auxiliary model. |
Copied to clipboard
| Challenge: | Existing methods for inductive reasoning over knowledge graphs lack the ability to model the logical structures of complex queries. |
| Approach: | They propose a structure-modeled textual encoding framework for inductive logical reasoning over KGs that encodes linearized query structures and entities using pre-trained language models to find answers. |
| Outcome: | The proposed framework encodes query structures and entities using pre-trained language models to find answers. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental task of information extraction. |
| Approach: | They propose to perform randomization tests on standard NER benchmarks to examine name regularity, mention coverage and context diversity. |
| Outcome: | The proposed model performs better on standard NER benchmarks than other models on open datasets. |
Copied to clipboard
| Challenge: | Large language model agents have enabled GUI-based automation, but their deployment is limited by noisy data, poor generalization, and lack of support for non-English GUIs. |
| Approach: | They propose an 8B-parameter GUI agent built for robust and efficient on-device GUI interaction. |
| Outcome: | The proposed GUI agent achieves promising performance on five public benchmarks and proposed Chinese benchmark CAGUI. |
Copied to clipboard
| Challenge: | Existing approaches to few-shot named entity recognition (NER) focus on coarse-grained entities with few examples, while most unseen entities are fine-grounded. |
| Approach: | They present a human-annotated few-shot named entity recognition dataset . they construct benchmark tasks to assess the generalization capability of models . |
| Outcome: | The proposed model is the first few-shot NER dataset and the largest human-crafted NER data set. |
Copied to clipboard
| Challenge: | Existing video fake news detection benchmarks focus on the detection accuracy, while failing to provide fine-grained assessments for the entire detection process. |
| Approach: | They propose a process-oriented video fake news detection benchmark that evaluates MLLMs' perception, understanding, and reasoning capabilities in VFND. |
| Outcome: | The proposed model achieves sota performance on video fake news detection tasks. |
Copied to clipboard
| Challenge: | Existing RAG frameworks face critical limitations due to text chunking and semantic similarity. |
| Approach: | They propose a framework that incorporates causal graphs into the retrieval process. |
| Outcome: | The proposed framework preserves contextual continuity and improves retrieval precision, leading to more accurate and interpretable responses. |
Copied to clipboard
| Challenge: | Existing LLMs lack immersion and adaptability, resulting in limited character orchestration and on-the-fly character introduction. |
| Approach: | They propose an LLM-based framework that allows actors to interact with users in an ongoing narrative. |
| Outcome: | The proposed framework outperforms commercial LLMs in character consistency, environment grounding, and narrative coherence. |
Copied to clipboard
| Challenge: | Existing incomplete multimodal learning frameworks are inadequate for integrating multimodal data. |
| Approach: | They propose a framework for incomplete multimodal learning that is deficiency-resistant and provides two modules to address fine-grained deficiencies. |
| Outcome: | The proposed framework outperforms the SOTA models on two well-known multimodal benchmarks. |
Copied to clipboard
| Challenge: | Existing methods to train language models on diverse text corpora have brought up performance improvements on several natural language understanding (NLU) tasks. |
| Approach: | They propose a method to automatically generate domain- and task-adaptive maskings of a given text for self-supervised pre-training. |
| Outcome: | The proposed framework outperforms rule-based masking strategies on question answering and text classification datasets on which it outperformed rule-driven masking techniques. |
Copied to clipboard
| Challenge: | Existing methods for retrieval-augmented generation (RAG) are limited and fine-tuning incurs prohibitive costs of external signals. |
| Approach: | They propose a self-supervised framework that enhances RAG systems through efficient model adaptation. |
| Outcome: | The proposed framework achieves 90% of the performance gain obtained through GPT-4-supervised adaptation while relying entirely on self-annotation of much smaller models. |
Copied to clipboard
| Challenge: | Existing Dynamic topic models are either fully supervised, requiring expensive human annotations, or fully unsupervised, producing topic evolutions that often do not cater to a user’s needs. |
| Approach: | They propose to use a framework that ensembles semantic similarity, category indicative, and time indicative scores to produce informative topic evolutions. |
| Outcome: | The proposed framework can be used to discover topic evolutions from temporal corpora that align with user-provided category names and uniquely capture topics at each time step. |
Copied to clipboard
| Challenge: | Existing video moment retrieval methods rely on sparse frame sampling, risking information loss. |
| Approach: | a new video-based framework enhances memory efficiency while maintaining high information resolution . SMORE uses query-guided captions to encode semantics aligned with user intent . |
| Outcome: | a new framework improves memory efficiency while maintaining high information resolution . it achieves state-of-the-art performance on QVHighlights, Charades-STA, and ActivityNet-Captions benchmarks . |
Copied to clipboard
| Challenge: | Existing QA methods lack scalability and performance is difficult to solve with document-level contexts. |
| Approach: | They propose an end-to-end deep network model that sequentially reads the input contexts into an external memory while replacing memories that are less important for answering unseen questions. |
| Outcome: | The proposed model improves on a synthetic dataset and real-world large-scale textual and video QA datasets. |
Copied to clipboard
| Challenge: | Visual arguments rely on images to persuade viewers to do or believe something . |
| Approach: | They propose three tasks for evaluating visual argument understanding . they use visual premises, commonsense premises and reasoning trees to analyze visual arguments . |
| Outcome: | The proposed tasks evaluate visual argument understanding using a dataset of 1,611 images annotated with 5,112 visual premises (with regions), 5,574 commonsense premises, and reasoning trees connecting them into structured arguments. |
Copied to clipboard
| Challenge: | Existing few-shot named entity recognition (NER) models capture information from limited instances while transferring useful knowledge from external resources. |
| Approach: | They propose a self-describing mechanism for few-shot NER which can universally describe mentions using concepts and automatically map novel entity types to concepts. |
| Outcome: | The proposed model can universally describe mentions using concepts and automatically map novel entity types to concepts and adaptively recognize entities on-demand. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have produced models that exhibit remarkable performance across a variety of NLP tasks. |
| Approach: | They analyze a large-scale collection of user-GPT conversations to identify a significant gap between academic research in NLP and the needs of real-world NLP applications. |
| Outcome: | The proposed model outperforms existing models in a large-scale collection of user-GPT conversations and identifies a significant gap between the tasks that users frequently request from LLMs and the tasks commonly studied in academic research. |
Copied to clipboard
| Challenge: | Incomplete learning is widespread and heterogeneous in large language models . authors identify five recurrent sources of incomplete learning: missing prerequisite knowledge, conflicts between SFT supervision and pre-training knowledge, internal inconsistencies within SFT data, left-side forgetting during sequential fine-tuning, and insufficient optimization for rare or complex patterns. |
| Approach: | They propose a diagnostic-first framework that maps incomplete learning to causes . they identify five recurrent sources of incomplete learning: missing prerequisite knowledge, conflicts between supervision and pre-training knowledge, internal inconsistencies, left-side forgetting during sequential fine-tuning, and insufficient optimization for rare or complex patterns. |
| Outcome: | The proposed framework maps incomplete learning to causes using observable training and inference signals. |
Copied to clipboard
| Challenge: | Existing methods to reduce memory usage for large language models neglect inter-layer dependency between layers and huge memory consumption in pre-computation. |
| Approach: | They propose a method that compresses the KV cache by layer-wise retaining crucial context. |
| Outcome: | The proposed method reduces memory usage by layer-wise retaining crucial context . it can improve 2.2x throughput compared to Accelerate with over 54% memory reduction . |
Copied to clipboard
| Challenge: | Existing instruction data synthesis methods focus on single-turn instructions and neglect cross-turn coherence, resulting in context drift and reduced task completion rates. |
| Approach: | They propose a framework that constrains multi-turn instruction synthesis by explicitly modeling human conversational intent. |
| Outcome: | The proposed framework outperforms existing models trained on single-turn and multi-turn instruction datasets. |
Copied to clipboard
| Challenge: | Existing approaches to handwritten mathematical expression recognition are limited by CFGs and pre-generated triplet data. |
| Approach: | They propose an architecture that integrates recognition and language features to output corrected sequences while optimizing with a string decoder recognition model. |
| Outcome: | The proposed architecture outperforms state-of-the-art methods on CROHME datasets. |
Copied to clipboard
| Challenge: | Existing studies regard auto-generated knowledge instances as gold references, which limits their effectiveness since they are not always accurate and inferior instances can lead to incorrect predictions. |
| Approach: | They propose to use regularized decoding and adversarial training to appropriately learn from noisy knowledge instances for Arabic diacritization. |
| Outcome: | The proposed model outperforms existing models on two benchmark datasets even with flawed auto-generated knowledge. |
Copied to clipboard
| Challenge: | Existing efforts to train pre-trained language models have brought significant improvements to various NLP applications. |
| Approach: | They propose to compress bulky LMs while preserving useful information for a specific task. |
| Outcome: | The proposed method can detach any layer without affecting others, and stretch shallow and wide LMs to be deep and narrow. |
Copied to clipboard
| Challenge: | Existing work on conversational recommendation systems lacks high-quality data . existing datasets lack large-scale and high-level data based on human annotators . |
| Approach: | They propose an automatic dataset synthesis approach that generates large-scale recommendation dialogues using structured graphs based on user-item information from the real world. |
| Outcome: | The proposed approach can generate large-scale and high-quality recommendation dialogues . it exploits user preferences, knowledge graphs, and conversation ability from existing datasets based on real-world data . |
Copied to clipboard
| Challenge: | Graph Attention Networks (GATs) are a promising model that takes advantage of localized attention mechanism to perform knowledge representation learning (KRL) on graph-structure data. |
| Approach: | They propose to incorporate global information into the GAT family of models by using an attention-based global random walk algorithm. |
| Outcome: | Experimental results on KG entity prediction against the state-of-the-arts demonstrate the effectiveness of the proposed model. |
Copied to clipboard
| Challenge: | Existing frameworks for data analysis and insight exploration are lacking in terms of benchmarks . existing frameworks suffer from format inconsistencies, poorly conceived objectives, and redundant insights. |
| Approach: | They propose a data-curation pipeline to construct a new dataset named InsightEval. |
| Outcome: | The proposed benchmarks highlight prevailing challenges in automated insight discovery and raise key findings to guide future research. |
Copied to clipboard
| Challenge: | Existing methods for paraphrase generation lack reliable supervision signals. |
| Approach: | They propose an unsupervised paradigm for paraphrase generation based on contextual language models, candidate filtering and paraphrase model training based upon the selected candidates. |
| Outcome: | The proposed paradigm outperforms existing paraphrase generation methods in supervised and unsupervised setups. |
Copied to clipboard
| Challenge: | Recent advances in artificial intelligence have led to the creation of highly capable large language models (LLMs) that can perform tasks in a human-like manner, but lack infant-level cognitive abilities in certain areas. |
| Approach: | They designed a text-based multi-choice QA scenario similar to the A-Not-B error to test their inhibitory control abilities. |
| Outcome: | The proposed model shows that state-of-the-art LLMs perform well with in-context learning but make errors and show a drop of as many as 83.3% in reasoning tasks when the context changes trivially. |
Copied to clipboard
| Challenge: | Large language models suffer from knowledge gaps and hallucinations, resulting in incorrect or poor reasoning. |
| Approach: | They propose Graph retrieval-augmented generation (GraphRAG) which integrates structured knowledge from external graphs to enhance model's reasoning. |
| Outcome: | Experiments on knowledge graph QA tasks show that GraphRAG significantly improves reasoning performance across multiple backbone models. |
Copied to clipboard
| Challenge: | Long-context inference is crucial for advancing large language models, but its prefill speed remains a bottleneck. |
| Approach: | They propose an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed. |
| Outcome: | The proposed framework achieves speedups of 9.2, 4.2, and 1.6 without any degradation in performance. |
Copied to clipboard
| Challenge: | Efficient resume parsing is critical for global hiring, yet the lack of dedicated benchmarks for evaluating large language models (LLMs) on multilingual, structure-rich resumes hinders progress. |
| Approach: | They propose to use a human-in-the-loop pipeline to generate 2,500 synthetic resumes spanning 50 templates, 30 career fields, and 5 languages to evaluate large language models. |
| Outcome: | The proposed benchmarks show that the models perform poorly on multilingual resumes and lack of standardized templates. |
Copied to clipboard
| Challenge: | Existing methods to improve inference efficiency target to reduce per-layer latency, but ignore cumulative latency due to number of layers. |
| Approach: | They propose to identify quasi-independent layers that can be concurrently computed to significantly decrease inference latency. |
| Outcome: | Empirical results show that the proposed method reduces latency by 48.3% on LLaMA-33B while maintaining close level of performance. |
Copied to clipboard
| Challenge: | Structured product information is a major bottleneck for the efficiency of e-commerce platforms. |
| Approach: | They propose a data-driven approach to generate product structured representations using product metadata. |
| Outcome: | Extensive experiments show that GSID can generate better product representations on real-world e-commerce platforms. |
Copied to clipboard
| Challenge: | Existing approaches to improve the reasoning performance of large language models rely on intuitive instance-level feedback, which limits the reasoning capabilities. |
| Approach: | They propose a framework that pushes LLMs toward System-2-like critic capability by using a step-wise CoT reasoning paradigm and automatic construction of weak-supervision data without human annotation. |
| Outcome: | The proposed model significantly improves task-solving performance by filtering out invalid solutions or iterative refinement. |
Copied to clipboard
| Challenge: | Existing systems that translate optimization formulas manually are cumbersome and time-consuming. |
| Approach: | They propose a system that converts optimization formulas from TeX document to solver language. |
| Outcome: | The proposed system helps operations research practitioners convert optimization formulations into solver modeling languages. |
Copied to clipboard
| Challenge: | Existing approaches for optimizing domain-level sampling strategies struggle with maintaining intra-domain consistency and accurately measuring domain impact. |
| Approach: | They propose to use a Fisher-Information Matrix-guided metric to measure domain impact to ensure intra-domain consistency and accuracy. |
| Outcome: | The proposed model achieves 3.4% higher average performance while maintaining comparable training efficiency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have significantly impacted various domains, especially through organized LLM-driven autonomous agents. |
| Approach: | They propose a framework that enables orchestrated teams to jointly propose various task-oriented solutions and interact with their insights in a self-independence while cross-team collaboration environment for superior solutions generation. |
| Outcome: | Experiments show that the framework can generate better software quality compared to state-of-the-art frameworks. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have demonstrated better safety performance in high-resource languages than in low-resourced languages. |
| Approach: | They propose language-agnostic semantic alignment (LASA) which anchors safety alignment directly in semantic bottlenecks. |
| Outcome: | The proposed approach significantly improves safety across all languages: average attack success rate drops from 24.7% to 2.8% on LLaMA-3.1-8B-Instruct and remains within 3–4% across Qwen2.5 and Qwend3 Instruct models (7B–32B). |
Copied to clipboard
| Challenge: | XLM-R and mBART have advanced multilingualism in NLP, but low-resource languages such as Tibetan, Uyghur, Kazakh, and Mongolian are underserved. |
| Approach: | They propose a framework for adapting multilingual encoders to text generation in extremely low-resource languages by reusing the weights between the encoder and the decoder. |
| Outcome: | The proposed framework performs better on various downstream tasks even when compared with much larger models. |
Copied to clipboard
| Challenge: | Existing metrics for evaluating functional correctness of SQL queries are prone to false positives due to inadequately prepared test databases. |
| Approach: | They propose a graph-based metric that uses a relational operator tree to extract rich semantic information from the logical execution plan of SQL queries and embed it into a diagram. |
| Outcome: | The proposed method eliminates the need for extensive test database preparation and performs graph matching on unseen SQL queries. |
Copied to clipboard
| Challenge: | Existing studies on medication recommendation mainly rely on EHRs, but some details of interactions between doctors and patients may be ignored or omitted in EHR. |
| Approach: | They propose to use medical dialogues to recommend medications with medical dialogue data . they propose to model dialogue structure and disease knowledge aware network . |
| Outcome: | The proposed method is a promising solution to recommend medications with medical dialogues. |
Copied to clipboard
| Challenge: | Existing GUI Agents face challenges in multi-step reasoning and reliance on textual annotations, limiting their effectiveness. |
| Approach: | They propose an MLLM-based GUI Agent with a two-stage supervised fine-tuning pipeline that enhances GUI understanding and grounding. |
| Outcome: | InfiGUIAgent achieves competitive performance on several GUI benchmarks, highlighting the impact of native reasoning skills in enhancing GUI interaction for automation tasks. |
Copied to clipboard
| Challenge: | Existing methods to augment knowledge graph completion require factual triples or manual prompts to extract knowledge from a pre-trained language model. |
| Approach: | They propose a tool that generates quality query prompts and retrieves support information from large text corpora to probe knowledge from a pre-trained language model. |
| Outcome: | The proposed method outperforms embedding-based, graph-based and PLM-based methods on two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods to extract information from evidence are unable to grasp relational and logical information among the evidence. |
| Approach: | They propose a graph-based evidence aggregating and reasoning framework to integrate evidence from multiple pieces of evidence. |
| Outcome: | The proposed framework achieves significant performance improvements on a large-scale benchmark dataset. |
Copied to clipboard
| Challenge: | Detecting LLM-generated text is crucial for academic integrity, preventing plagiarism, protecting copyrights, ethical research practices. |
| Approach: | They propose a method specifically designed for Korean language to detect LLM-generated text . they examine spacing patterns, part-of-speech diversity, and comma usage . |
| Outcome: | The proposed method achieves an average of 19.78% higher AUC-ROC compared to the best-performing detection method. |
Copied to clipboard
| Challenge: | Existing methods for embedding knowledge graphs implicitly memorize relation rules to infer missing links, but they are difficult to memorize due to the inherent deficiencies of such implicit memorization strategy. |
| Approach: | They propose a vertical learning paradigm that allows to explicitly copy target information from related factual triples for more accurate prediction. |
| Outcome: | The proposed model improves generalization ability and makes distant link prediction significantly easier. |
Copied to clipboard
| Challenge: | Existing strategies to circumvent safety constraints face significant trade-offs between effectiveness and efficiency. |
| Approach: | They propose a framework that allows to infer model refusal behaviors without expensive parameter updates or training. |
| Outcome: | The proposed framework outperforms baselines in multiple safety-aligned open-source LLMs. |
Copied to clipboard
| Challenge: | Reverse-Enhanced Thinking (RevThink) is a framework for large language models to perform reverse thinking. |
| Approach: | They propose a framework for enhancing forward-backward reasoning by collecting data from a teacher model and employing three objectives to train a student model in a multi-task learning fashion. |
| Outcome: | The proposed framework outperforms a fine-tuning method trained on 10x more forward reasoning on 12 datasets covering commonsense, math, and logical reasoning. |
Copied to clipboard
| Challenge: | generative AI is expanding in education, yet empirical analyses of large-scale and real-world interactions between students and AI systems remain limited. |
| Approach: | They present a dataset based on a semester-long experiment with 212 college students in English as Foreign Language (EFL) writing courses. |
| Outcome: | The proposed dataset includes conversation logs, students’ intent, students' self-rated satisfaction, and students’ essay edit histories. |
Copied to clipboard
| Challenge: | Existing methods for reinforcement learning (RL) on self-generated data are limited in many domains. |
| Approach: | a new framework combines plan-based search with Step-level Advantage Preference Optimization to optimize plan learning. |
| Outcome: | The proposed framework improves in-domain performance and out-of-domain benchmarks. |
Copied to clipboard
| Challenge: | Existing models do not distinguish genuine users from social bots, and their failure in identifying rumors timely. |
| Approach: | They propose to account for social bots’ behavior and construct a Social Bot-Aware Graph Neural Network to model early propagation of posts and then use it to detect rumors. |
| Outcome: | The proposed method achieves significant improvements over baselines and identifies rumors within 3 hours while maintaining more than 90% accuracy. |
Copied to clipboard
| Challenge: | Existing approaches to long-term dialogue memory management fail to capture the natural semantic structure of conversations, leading to fragmented and incomplete representations. |
| Approach: | They propose a mechanism that integrates forward- and backward-looking reflections into a personalized memory bank for effective future retrieval. |
| Outcome: | The proposed mechanism outperforms state-of-the-art benchmarks on a long-term dialogue memory model. |
Copied to clipboard
| Challenge: | acquiring and representing commonsense in machines has posed a long-standing challenge (Li et al., 2021; Zhang e t al, 2022; Zhou e al. 2023) . |
| Approach: | They use a commonsense-based LLM to evaluate ChatGPT's commonsensing abilities by analyzing 11 datasets and generating knowledge descriptions. |
| Outcome: | The proposed model can achieve good QA accuracies while still struggling with certain domains of datasets. |
Copied to clipboard
| Challenge: | Existing studies have shown that word embedding improves accuracy on NLP tasks. |
| Approach: | They propose a causal diagram based on the evaluation results of word embeddings using partial least squares path modeling. |
| Outcome: | The proposed model proves that word embedding contributes to solving downstream tasks. |
Copied to clipboard
| Challenge: | a new method to detect clickbait posts on the Web is needed to detect such posts. |
| Approach: | They propose a method to detect clickbait posts on the Web using latent factors . they use features in multiple modalities to characterize the posts and causal inference to eliminate noise . |
| Outcome: | The proposed method can detect clickbait posts on popular social media platforms with good generalization ability. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL)-based post-training often improves the reasoning performance of large language models beyond the training domain, while supervised fine-tuning (SFT) frequently leads to general capabilities forgetting. |
| Approach: | They propose a feature-level mechanistic analysis methodology to probe RL generalization using a controlled experimental setup. |
| Outcome: | The proposed method identifies a compact, task-agnostic set of features that directly mediate generalization across diverse tasks. |
Copied to clipboard
| Challenge: | Existing methods for streaming video understanding are query-agnostic and implicitly model video evidence. |
| Approach: | They propose a framework that establishes explicit, structured alignment between the accumulated video evidence and the query’s expected response conditions via scene graphs. |
| Outcome: | The proposed model achieves more interpretable and accurate response timing decisions on both proactive and reactive tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on character-centric approach and fail to reflect real-world applications. |
| Approach: | RMTBench is a user-centric bilingual role-playing benchmark featuring 80 diverse characters and over 8,000 dialogue rounds. |
| Outcome: | RMTBench features 80 diverse characters and over 8,000 dialogue rounds. |
Copied to clipboard
| Challenge: | Recent studies show that pre-trained masked language models can be factual knowledge bases. |
| Approach: | They conduct a rigorous study to explore the underlying predicting mechanisms of MLMs . they find that previous decent performance mainly owes to the biased prompts which overfit dataset artifacts a . |
| Outcome: | The proposed model improves on illustrative cases and external contexts . the results question the previous findings that MLMs can be reliable factual knowledge bases . |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) have shown their superiority by pre-training on unstructured text corpus and then fine-tuning on downstream tasks. |
| Approach: | They propose a Knowledge-Enhanced Pre-trained LanguagE model with Topic entity awareness that incorporates the interactions between tokens and mentioned entities in pre-training. |
| Outcome: | The proposed model incorporates the interactions between tokens and mentioned entities in pre-training and is more effective on entity-centric tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks assess only the final answer with a wide numerical tolerance, overlooking systematic reasoning failures and potentially causing serious clinical misjudgments. |
| Approach: | They propose a new step-by-step evaluation pipeline that assesses formula selection, entity extraction, and arithmetic computation. |
| Outcome: | The proposed method improves the accuracy of large language models on medical benchmarks from 16.35% to 53.19%. |
Copied to clipboard
| Challenge: | Existing methods for opinion expression detection are based on token-level sequence labeling . |
| Approach: | They propose to use BERT and conditional random field embedders to detect opinion expressions. |
| Outcome: | The proposed model outperforms ELMo embedders in opinion expression detection. |
Copied to clipboard
| Challenge: | In few-shot text classification, self-training relies on pseudo-labels to expand data, which has shown success, but can accumulate errors due to noisy pseudo-labeled data. |
| Approach: | They propose a method to mitigate noise in noisy pseudo-labeled data by applying superficial learning to noisy data and fine-tuning to less noisy data. |
| Outcome: | The proposed framework improves the classifier accuracy for few-shot text classification by 18.5% at most and 8% in average, compared with the state-of-the-art SSL baselines. |
Copied to clipboard
| Challenge: | Scientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers . previous attempts failed to scale up due to heavy human annotation and domain expertise . |
| Approach: | They propose a method to automatically extract TLDR summaries from scientific papers . they propose 'citeSum' with no human annotation, which is 30 times larger than SciTLDR . |
| Outcome: | The proposed approach outperforms most fully-supervised methods on SciTLDR without fine-tuning and achieves state-of-the-art results with only 128 examples. |
Copied to clipboard
| Challenge: | Increasing number of parameters can be challenging under resource-constrained environments. |
| Approach: | They propose a parameter-efficient fine-tuning method with fewer parameters and finer granularity that can adaptively select important parameters for each task. |
| Outcome: | The proposed method can fine-tune important parameters for each task, while maintaining the same weights. |
Copied to clipboard
| Challenge: | Existing distillation approaches target Small Language Models (SLMs) or Conventional Recommendation Models, but face a critical trade-off between computational cost and semantic reasoning capacity. |
| Approach: | They propose a framework that establishes a text encoder as the optimal student architecture for scalable recommendation. |
| Outcome: | Experiments on four datasets show that the proposed framework outperforms state-of-the-art models and achieves significantly reduced latency. |
Copied to clipboard
| Challenge: | Traditional machine translation evaluation relies on reference written by humans . reference-free evaluation gets rid of labor-intensive annotations, which can pivot easily to new domains . |
| Approach: | They propose a reference-free evaluation approach that characterizes evaluation as two aspects: fluency and faithfulness. |
| Outcome: | The proposed approach outperforms SOTA reference-fee metrics on machine translation datasets. |
Copied to clipboard
| Challenge: | Chain-of-Thought prompting has improved the reasoning capabilities of Large Language Models (LLMs) but it is ineffective or detrimental to the performance on reasoning tasks in Smaller Language Model (SLMs) with less than 10 billion parameters. |
| Approach: | They propose a Dialogue-guided Chain-of-Thought method to improve the reasoning capabilities of Large Language Models (LLMs) by generating intermediate reasoning steps in a dialogue format to guide the model to the final answer. |
| Outcome: | The proposed method can achieve significant performance gains over state-of-the-art competitors on four arithmetic reasoning datasets. |
Copied to clipboard
| Challenge: | Existing Large Vision-Language Models (LVLMs) lack integrated commonsense knowledge . lack of integrated common knowledge limits their robustness and accuracy in VQA . |
| Approach: | They propose a framework to enhance multimodal inference by integrating commonsense reasoning. |
| Outcome: | MAGIC-VQA improves comprehensive benchmark datasets, surpassing existing models in tasks requiring advanced commonsense reasoning. |
Copied to clipboard
| Challenge: | Existing methods for event argument extraction cannot adequately model the correlation between event arguments and their roles. |
| Approach: | They propose a Bayesian model to jointly extract event arguments using Gibbs sampling . they train two neural networks to model prior distribution and conditional distribution over event arguments . |
| Outcome: | The proposed model can achieve comparable results to existing methods on two widely-used datasets. |
Copied to clipboard
| Challenge: | Existing contrastive methods that ignore the context of a large language model (LLM) fail to handle instances that vary in their amount of conflict, with static methods over-adjusting when conflict is absent. |
| Approach: | They propose a fine-grained, instance-level approach called AdaCAD which dynamically adjusts the degree of conflict based on the degree. |
| Outcome: | The proposed approach outperforms baselines and improves factuality of summaries by 6.19. |
Copied to clipboard
| Challenge: | Pre-trained language models (PTLMs) have achieved noticeable success on many NLP tasks, but struggle for tasks that require event temporal reasoning. |
| Approach: | They propose a continual pre-training approach that equips PTLMs with targeted knowledge about event temporal relations by focusing on masked-out event and temporal indicators and discriminating sentences from their corrupted counterparts. |
| Outcome: | The proposed framework improves the PTLMs’ fine-tuning performances across five relation extraction and question answering tasks and achieves new or on-par state-of-the-art in most of our downstream tasks. |
Copied to clipboard
| Challenge: | Open-domain question answering aims at locating answers to user-generated questions in massive collections of documents. |
| Approach: | They propose an algorithm with a novel reader-retriever design that differs from both families of algorithms. |
| Outcome: | The proposed algorithm outperforms retrieval-based methods with two large-scale datasets and is state-of-the-art. |
Copied to clipboard
| Challenge: | Recent studies have shown that large language models (LLMs) have impressive capabilities in dealing with new tasks with the help of in-context learning (ICL). |
| Approach: | They propose to concate the image and text embeddings to enhance the retrieval performance of a visual-language task and to calculate a list-wise ranking loss for training the embeddable model. |
| Outcome: | The proposed framework fine-tunes the CLIP embedding model to better meet the needs of the large vision-language models. |
Copied to clipboard
| Challenge: | evaluators of simple factoid question answering using different datasets are not able to solve SimpleQuestions. |
| Approach: | They evaluate the progress of the field toward solving simple factoid questions over a knowledge base. |
| Outcome: | The proposed model is nearly solved on the most popular dataset, but not on the robustness of existing systems. |
Copied to clipboard
| Challenge: | Recent advances in text-to-speech (TTS) models have led to improvements in speaker prosody and voices modeling. |
| Approach: | They propose an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies. |
| Outcome: | The proposed model surpasses state-of-the-art models in both naturalness and similarity while reducing inference speed by 90%. |
Copied to clipboard
| Challenge: | Existing studies have used general approaches to alleviate the overfitting of supervised models based on video data with sentiment annotations. |
| Approach: | They propose to capture common sentimental patterns in unlabeled videos using sentiment knowledge and non-verbal behavior to embed sentiment information into pre-trained multimodal representations. |
| Outcome: | The proposed model outperforms the baseline and achieves new State-Of-The-Art (SOTA) results. |
Copied to clipboard
| Challenge: | Several pre-training models of different modalities are showing a rising trend of homogeneity in their model structures. |
| Approach: | They propose a toolkit that supports pre-training models of different modalities. |
| Outcome: | The proposed toolkit can match the performance of the original implementations on text, vision, and audio benchmarks. |
Copied to clipboard
| Challenge: | Existing definition generation techniques have faced various problems such as the out-of-vocabulary problem and over/under-specificity problems. |
| Approach: | They propose to leverage a pre-trained encoder-decoder model and introduce a re-ranking mechanism to model specificity in definitions. |
| Outcome: | The proposed method significantly outperforms the state-of-the-art method on standard evaluation datasets and shows that it addresses the over/under-specificity problems. |
Copied to clipboard
| Challenge: | Existing offline approaches to improve an LLM-based customer support system rely on batch annotations. |
| Approach: | They propose an agent-in-the-loop framework that integrates four key types of annotations directly into live customer operations: (1) pairwise response preferences, (2) agent adoption and rationales, (3) knowledge relevance checks, and (4) identification of missing knowledge. |
| Outcome: | The proposed framework reduces retraining cycles from months to weeks by integrating four key types of annotations directly into live customer operations. |
Copied to clipboard
| Challenge: | Grapheme-to-phoneme conversion (g2p) is a task of predicting the pronunciation of words from their orthographic representation. |
| Approach: | They propose to leverage audio data as an auxiliary modality in a multi-task training process to learn a more optimal grapheme representation. |
| Outcome: | The proposed model reduces phoneme error rate to 2.46% on in-domain test set compared to unimodal spelling- pronunciation model. |
Copied to clipboard
| Challenge: | Existing models for task progress estimation lack long-horizon and dynamic reasoning . estimating how much of a task has been completed requires long-term reasoning based on partial information. |
| Approach: | They propose a benchmark for evaluating progress reasoning from a single observation . they instantiate a two-stage paradigm that combines episodic retrieval with mental simulation . |
| Outcome: | The proposed benchmark improves on 14 VLMs on a small scale and shows common failure patterns. |
Copied to clipboard
| Challenge: | Large Reasoning Models suffer from the "over-thinking" problem, causing performance degradation. |
| Approach: | They propose a unified model that balances reasoning performance and efficiency across multiple formats through a reinforcement learning framework augmented with length-aware optimization. |
| Outcome: | The proposed model reduces token costs while preserving performance compared to traditional models. |
Copied to clipboard
| Challenge: | Existing methods for mixing-of-agents (MoA) lack model selection criteria and struggle with large model pools. |
| Approach: | They propose a mixture-of-agents framework with dynamic routing that uses a lightweight scorer to perform initial screening and refines the model scores through self- and cross-assessment. |
| Outcome: | The proposed framework outperforms existing methods for large model pools and tasks . it reduces cost by 89.8% and latency by 63.6% in the large-scale model pool. |
Copied to clipboard
| Challenge: | Existing studies rely on semantic similarity to retrieve knowledge but ignore fine-grained information within documents. |
| Approach: | They propose a fine-grained knowledge enhancement method to fill knowledge gaps with retrieved external information by a Chain-of-Thought prompting procedure and a decoding enhancement strategy to constrain the document-based decoding process. |
| Outcome: | The proposed method can be applied in a plug-and-play manner to enhance its performance with no additional modules or training process. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced tasks like text summarization, but their size and computational demands limit their use in resource-constrained and privacy-centric settings. |
| Approach: | They propose a framework for distilling LLMs’ text summarization abilities into a compact, local model using a curriculum learning strategy that evolves from simple to complex tasks. |
| Outcome: | The proposed framework outperforms baseline models on CNN/DailyMail, XSum, and ClinicalTrial, and improves interpretability by providing insights into the summarization rationale. |
Copied to clipboard
| Challenge: | Despite of significant achievements in improving instruction-following capabilities of large language models, the ability to process multiple potentially entangled or conflicting instructions remains a considerable challenge. |
| Approach: | They construct multi-turn instruction with 1.1K high-quality multi-turned conversations using the human-in-the-loop approach and examine their capabilities. |
| Outcome: | The proposed model shows that it is difficult to integrate multiple turns and balance competing objectives when instructions intersect or conflict. |
Copied to clipboard
| Challenge: | Existing approaches to integrate the recommendation function and dialog generation function smoothly are lacking. |
| Approach: | They propose to integrate dialog context for recommendation and dialog generation better using a pre-trained language model and an item metadata encoder to integrate the recommendation and dialogue generation. |
| Outcome: | The proposed architecture improves the integration of recommendation and dialog generation functions. |
Copied to clipboard
| Challenge: | Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), but its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content and poor OCR quality. |
| Approach: | They propose a corruption-robust training paradigm that surpasses existing strategies for mitigating the effects of corrupted data. |
| Outcome: | The proposed training paradigm surpasses existing strategies for mitigating the effects of corrupted data. |
Copied to clipboard
| Challenge: | Recent supervised ED approaches have achieved promising performance but require large number of manually annotated event data. |
| Approach: | They propose to overfit the trigger confounder of the context and the result . they propose to intervene on the context via backdoor adjustment during training . |
| Outcome: | The proposed method significantly improves the FSED on ACE05 and MAVEN datasets. |
Copied to clipboard
| Challenge: | Existing event extraction models have been limited to the sentence level . this formulation signifies a misalignment between the information seeking behavior and the informative seeking behavior. |
| Approach: | They propose a document-level neural event argument extraction model by formulating the task as conditional generation following event templates. |
| Outcome: | The proposed model achieves 7.6% F1 and 5.7% F1 over the best baseline on the document-level event extraction dataset WikiEvents and 9.3% F1 on the informative argument extraction task. |
Copied to clipboard
| Challenge: | Existing CSRL parsers struggle to handle conversational structural information. |
| Approach: | They propose a conversational semantic role labeling task which explicitly encodes speaker dependent information and proposes a multi-task learning method to further improve the model. |
| Outcome: | The proposed model outperforms baselines on benchmark datasets on conversation-based tasks. |
Copied to clipboard
| Challenge: | Existing methods to counter trolling in online communities are not yet available to address the diversity of trolling behaviors. |
| Approach: | They propose a method for generating counter-responses to trolls by aligning these strategies with human preferences across different trolled contexts. |
| Outcome: | The proposed approach reduces negative effects of trolling and improves the online community environment. |
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) reasoning has improved the performance of large language models (LLMs) however, the detailed reasoning process in CoT often incurs long generation times and high computational costs due to the inclusion of unnecessary steps. |
| Approach: | They propose a method to identify critical reasoning steps using perplexity as a measure of their importance. |
| Outcome: | The proposed method achieves a better balance between reasoning accuracy and efficiency of CoT. |
Copied to clipboard
| Challenge: | null |
| Approach: | null |
| Outcome: | null |
Copied to clipboard
| Challenge: | Existing models that require task labels or performance trade-offs are susceptible to catastrophic forgetting. |
| Approach: | They propose a representation-aware model merging framework for continual learning without access to historical data. |
| Outcome: | The proposed framework outperforms baselines in knowledge retention and generalization across five NLP tasks and multiple continual learning scenarios. |
Copied to clipboard
| Challenge: | OpenAI introduces deliberative alignment (DA) to enhance safety of its o-series models, but effectiveness of this approach in open-source LLMs is understudied. |
| Approach: | They propose a case-augmented deliberative alignment method for large language models . they propose to use reinforcement learning on self-generated safety reasoning chains . |
| Outcome: | The proposed method avoids narrowly enumerated rules and allows broader adaptability. |
Copied to clipboard
| Challenge: | LongLeader aims to assess different LLMs' long-context comprehension abilities . long-constext comprehension is a key bottleneck for many use cases . |
| Approach: | They propose a leaderboard to assess different LLMs' long-context comprehension abilities . they offer open-source access to the benchmarks and maintain a dedicated website . |
| Outcome: | The proposed model assesses different LLMs on selected benchmarks and provides open-source access to the benchmarks. |
Copied to clipboard
| Challenge: | Pre-trained language models have shown remarkable memory formation, but vanilla networks without pre-training suffer catastrophic forgetting problem. |
| Approach: | They conduct experiments to investigate the retentive-forgetful contradiction between vanilla and pre-trained language models by controlling the target knowledge types, learning strategies and learning schedules. |
| Outcome: | The results show that pre-trained language models are forgetful and pre-training leads to retentive models . |
Copied to clipboard
| Challenge: | Hate speech detection is challenging when there are insufficient lexical cues. |
| Approach: | They propose a contrastive learning method that pulls an implication and its corresponding posts close in representation space. |
| Outcome: | The proposed method improves on BERT and HateBERT benchmarks on three implicit hate speech benchmarks. |
Copied to clipboard
| Challenge: | Small language models struggle with tool-use tasks, particularly in selecting appropriate tools and identifying correct parameters. |
| Approach: | They propose a training-free method that leverages peakedness to align schemas with pretraining knowledge to rename tool components. |
| Outcome: | Experiments on MetaTool and RoTBench show that PA-Tool significantly improves tool-use accuracy without retraining. |
Copied to clipboard
| Challenge: | Current retrieval-augmented generation systems struggle when retrieval models fail to rank the most relevant documents . existing extractive methods reduce latency but rely on independent, non-adaptive sentence selection . |
| Approach: | They introduce an extractive context compression framework that enhances retrieval-augmented generation in question answering. |
| Outcome: | EXIT surpasses existing compression methods and uncompressed baselines in QA accuracy . the framework reduces inference time and token count while preserving contextual dependencies . |
Copied to clipboard
| Challenge: | Existing approaches to multimodal learning assume a complete input modality setting, i.e., each modality is either complete or completely missing in both training and test sets. |
| Approach: | They propose an alignment dynamics learning module based on the theory of optimal transport for missing data imputation and a denoising training algorithm to enhance the quality of iputation and accuracy of model predictions. |
| Outcome: | The proposed method performs faster and more accurate inferences under different missing conditions and alleviates the overfitting issue. |
Copied to clipboard
| Challenge: | Existing methods for understanding long videos are limited due to the sparsity of visual evidence relevant to a given query. |
| Approach: | They propose a framework that enables VideoLLMs to reason over long videos and refine their predictions through executable programs. |
| Outcome: | The proposed framework outperforms existing methods across long-video understanding benchmarks. |
Copied to clipboard
| Challenge: | Prior work has successfully applied Reinforcement Learning (RL) to mathematical reasoning, but generalization to broader domains remains challenging due to limited data and lack of verifiable rewards for unstructured domains. |
| Approach: | They propose a framework that integrates multi-domain corpora into RL training to improve generalization across diverse reasoning tasks. |
| Outcome: | The proposed framework improves generalization across diverse reasoning tasks. |
Copied to clipboard
| Challenge: | Existing RAG paradigms suffer from the impact of flawed information introduced during the retrieval phrase, thereby diminishing the reliability and correctness of the generated output. |
| Approach: | They propose a framework that empowers models to discern and process information based on its credibility. |
| Outcome: | The proposed framework outperforms existing models with retrieval augmentation and exhibits robustness despite increasing noise in the context. |
Copied to clipboard
| Challenge: | Existing relation extraction methods rely on exact matching with human-annotated reference relations, while GRE methods produce diverse and semantically accurate relations. |
| Approach: | They propose a multi-dimensional assessment of relation extraction methods using human-annotated reference relations. |
| Outcome: | The proposed method is consistent with human preferences for RE quality. |
Copied to clipboard
| Challenge: | OpenNRE provides a framework to implement neural relation extraction (RE) . the toolkit provides various functional modules based on TensorFlow and PyTorch . |
| Approach: | OpenNRE is an open-source framework to implement neural relation extraction models. they also release an online system to meet real-time extraction without any training and deployment. |
| Outcome: | OpenNRE provides a framework to implement neural models for relation extraction (RE) the toolkit also includes an online system to meet real-time extraction without training and deployment . |
Copied to clipboard
| Challenge: | Existing quantization methods are compromising performance of large language models (LLMs) despite their high computational intensity, LLMs are still demanding intensive computation. |
| Approach: | They propose to generate the KV cache of pivot tokens losslessly from the full-precision model. |
| Outcome: | The proposed method generates the KV cache of pivot tokens losslessly from the full-precision model with no extra inference overhead. |
Copied to clipboard
| Challenge: | Existing methods for tuning pre-trained language models ignore the running cost and only optimize the terminal cost. |
| Approach: | They propose to use stochastic bridges to regularize intermediate states and use regularization as running cost of PETs. |
| Outcome: | The proposed methods can be used to tune large pre-trained language models . they can be compared to full-parameter fine-tuning by tuning a small number of parameters . |
Copied to clipboard
| Challenge: | Product review summarization aims to generate a concise summary based on product reviews . factual accuracy, aspect comprehensiveness, and content relevance are challenges . |
| Approach: | They propose an FB-Thinker framework to improve product review summarization ability . they propose two Chinese product review summary datasets for instruction-tuning and evaluation . |
| Outcome: | The proposed framework improves product review summarization with forward reasoning and backward refinement. |
Copied to clipboard
| Challenge: | Language models such as GPT and Llama have shown remarkable ability on diverse natural language tasks, yet their performance on complex table tasks is suboptimal. |
| Approach: | They propose a generator-validator paradigm to iteratively generate-then-validate training data from language models to fine-tune stronger Table-Specialist models that can specialize in a given task, without using manually-labeled data. |
| Outcome: | The proposed model outperforms vanilla language models on diverse table tasks and can match or surpass GPT-4 level quality. |
Copied to clipboard
| Challenge: | Existing methods to enhance textual entity prediction neglect the need for external knowledge or encounter high redundancy in the retrieved knowledge. |
| Approach: | They propose a framework that leverages ChatGPT as an implicit knowledge base and heuristically generates auxiliary knowledge for more efficient entity prediction. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on two classic datasets and exhibits a stronger robustness and generalization capability. |
Copied to clipboard
| Challenge: | Existing monotonic scaling methods for large reasoning models are not reliable. |
| Approach: | They propose a universal framework for modulating reasoning progress in large reasoning models at test time. |
| Outcome: | The proposed framework unifies and generalizes existing monotonic scaling methods and enables flexible and dense slow-to-fast reasoning modulation. |
Copied to clipboard
| Challenge: | Existing rankers assign inconsistent scores to functionally equivalent SQL queries . ranking cannot recover when the correct SQL is absent from the pool. |
| Approach: | They propose a Text-to-SQL framework that rewards ranking and resampling . it first groups candidates by execution result and ranks groups for consistency . |
| Outcome: | The proposed framework achieves 75.03 execution accuracy on BIRD-dev, a new state of the art among methods using models with disclosed sizes. |
Copied to clipboard
| Challenge: | Pretrained language models (LMs) are a powerful transfer learning approach for knowledge graph (KG) completion. |
| Approach: | They propose a parameter-lite transfer learning approach for pretrained language models for knowledge graph (KG) completion. |
| Outcome: | The proposed model outperforms the state-of-the-art models on a knowledge graph completion benchmark by tuning 1% of the parameters. |
Copied to clipboard
| Challenge: | Existing open-domain question answering systems only select one source to generate answer or conduct reasoning on structured information. |
| Approach: | They propose a Document-Entity Heterogeneous Graph Network to integrate different sources of information and conduct reasoning on heterogeneous information. |
| Outcome: | The proposed model outperforms the state-of-the-art methods on a HybirdQA dataset. |
Copied to clipboard
| Challenge: | Current research on large language models with retrieval-augmented code generation (RACG) has focused on single-language settings, leaving their cross-lingual effectiveness underexplored. |
| Approach: | They construct a dataset covering 13 PLs with nearly 14K instances to study cross-lingual code knowledge transfer in RACG. |
| Outcome: | The proposed model shows unequal cross-lingual knowledge transfer even with direct injection and shows limited reliance on natural language information embedded in code when equipped with a code-specific retriever. |
Copied to clipboard
| Challenge: | Existing approaches to knowledge graph question answering (KGQA) rely on Large Language Model (LLM) agents for graph traversal and retrieval. |
| Approach: | They propose a framework that synergizes Large Language Models with specialized graph retrieval tools to enhance KGQA. |
| Outcome: | The proposed framework outperforms the second-best graph retrieval method by 4.5% points while showing better generalization to custom KGs. |
Copied to clipboard
| Challenge: | Existing methods to train neural machine translation models are data-hungry and low-resource . et al., 2018; Radford e.t., 2019; Yang ee.,2019) proposes a new pre-training method for NMT . |
| Approach: | They propose a new pre-training method which randomly replaces some words in the input sentence with their translation words in target language. |
| Outcome: | The proposed method improves on unsupervised and supervised NMT models by making full use of monolingual corpora. |
Copied to clipboard
| Challenge: | Existing methods for open relation extraction give sub-optimal results on specific topics. |
| Approach: | They propose a method that leverages the built-in knowledge of large language models to maintain a dynamic seed relation dictionary for the topic. |
| Outcome: | The proposed approach empowers better topic-oriented control over the generated relations and improves ORE performance along the five dimensions, especially on specialized and narrow topics. |
Copied to clipboard
| Challenge: | Current approaches to regular expression denial of service (ReDoS) are hampered by a trade-off. |
| Approach: | They propose a framework to harness LLM generalization while enforcing reliability by localizing the vulnerable subpattern and generating a semanticallyequivalent fix for this isolated segment. |
| Outcome: | The proposed framework improves repair rates by 15.4%p over the state-of-the-art. |
Copied to clipboard
| Challenge: | Existing literature primarily addresses this problem through external interventions such as retrieval augmentation and prompt engineering at the input or output level. |
| Approach: | They find that LLMs can still produce hallucinated outputs when using structured external knowledge. |
| Outcome: | The proposed models fail to ground the provided knowledge, causing the model to revert to parametric memory. |
Copied to clipboard
| Challenge: | Existing approaches to automatic text dating ignore diachronic change of words, which may affect the efforts of text modeling. |
| Approach: | They propose a time-aware language model to learn temporal word representations by transferring language models of general domains to those of time-specific ones and build a hierarchical modeling approach to represent diachronic documents by encoding them with temporal representations. |
| Outcome: | The proposed model outperforms state-of-the-art approaches in historical text dating and other NLP tasks. |
Copied to clipboard
| Challenge: | Existing frameworks for commonsense generation are lacking for pre-trained models. |
| Approach: | They propose a framework that uses concept matching to retrieve prototype sentences and trainable sentence retriever to enhance pre-training and fine-tuning. |
| Outcome: | The proposed framework achieves state-of-the-art on the large-scale Common-Gen benchmark. |
Copied to clipboard
| Challenge: | Recent work studies watermarking under benign prompts, but its behavior under jailbreaking prompts remains underexplored. |
| Approach: | They evaluate six methods on four LLMs using two jailbreak benchmarks and three settings: Static, AutoDAN, and DSN. |
| Outcome: | The proposed methods inflate judge-based attack success rate under jailbreaking, but not harmful-goal compliance. |
Copied to clipboard
| Challenge: | Existing models that ignore the temporal relatedness of documents are time-agnostic and therefore fail to perform in automatic text dating. |
| Approach: | They propose a supervised fine-tuning model for automatic text dating that captures temporal semantic information and uses a contrastive learning-based approach to model two types of temporal relations of diachronic documents. |
| Outcome: | The proposed model outperforms state-of-the-art models on two diachronic corpora and captures temporal semantic information. |
Copied to clipboard
| Challenge: | Recent studies suggest data augmentation approaches to resolve the low-resource problem in natural language processing tasks. |
| Approach: | They propose to use slot information to augment sentences using a set of injective relations between a sentence’s semantics and its syntactical structure to augment the dataset. |
| Outcome: | The proposed approach outperforms all other data augmentation methods by 19.38%. |
Copied to clipboard
| Challenge: | Loki is an open-source fact-checking tool designed to address the growing problem of misinformation. |
| Approach: | They propose a tool that breaks down the fact-checking task into five steps . they propose LOKI, which offers a semiautomated, human-in-the-loop approach . |
| Outcome: | a new open-source tool is designed to address the growing problem of misinformation . the tool breaks down the fact-checking task into five steps to assist human judgment . |
Copied to clipboard
| Challenge: | Clinical notes are an extensive repository of information specific to individual patients. |
| Approach: | They create synthetic large-scale clinical notes using publicly available case reports extracted from biomedical literature and train a clinical large language model, Asclepius. |
| Outcome: | The proposed model outperforms several other models and is supported by detailed evaluations conducted by GPT-4 and medical professionals. |
Copied to clipboard
| Challenge: | Logical table-to-text generation requires models to derive logical-level facts from table records via logical inference. |
| Approach: | They propose a pretrained logical form generator framework to improve generation fidelity . they use a dataset to test the logical inference accuracy of the framework . |
| Outcome: | The proposed framework outperforms baselines on LOGICNLG and CONTLOG on two benchmarks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for human-machine interactions are static and can be misleading. |
| Approach: | They propose to use a LLM-based user agent to assess an assistant's API call capability without human involvement. |
| Outcome: | The proposed method mirrors real human conversation patterns in human-machine interactions, and shows that it aligns more closely with human assessment. |
Copied to clipboard
| Challenge: | Large language models with instruction-following capabilities are not suitable for long-tail ad hoc extraction use cases for non-expert users. |
| Approach: | They propose a task that follows instructions to extract the desired content from the associated text and present it in a structured tabular format. |
| Outcome: | The proposed paradigm outperforms existing open-source models of similar size in terms of information extraction. |
Copied to clipboard
| Challenge: | Medical images are widely used in clinical decision-making, where writing radiology reports can be enhanced by automatic solutions to alleviate physicians’ workload. |
| Approach: | They propose an approach with reinforcement learning over a cross-modal memory to better align visual and textual features for radiology report generation. |
| Outcome: | The proposed approach improves cross-modal alignment on two English radiology report datasets and human evaluation confirms the results. |
Copied to clipboard
| Challenge: | Large language models exhibit behavior that deviates from the boundaries of their knowledge during response generation. |
| Approach: | They propose a framework that allows large language models to explore their knowledge boundaries and self-correct generation behavior through fine-grained feedback signals. |
| Outcome: | The proposed framework enables LLMs to explore their knowledge boundaries and self-correct generation behavior through fine-grained feedback signals. |
Copied to clipboard
| Challenge: | Existing methods for unknown intent detection are limited by prior knowledge of class labels. |
| Approach: | They propose to use a Gaussian mixture model to model utterance embeddings with a distribution and inject dynamic class semantic information into Gausssian means. |
| Outcome: | The proposed model performs well on three real task-oriented dialogue datasets in two languages. |
Copied to clipboard
| Challenge: | In-context learning is an inductive bias for compositional generalization, but many deep neural architectures struggle with this ability. |
| Approach: | They propose to force a causal Transformer to in-context learn to promote compositional generalization by using earlier examples to generalize to later ones. |
| Outcome: | The proposed model can solve 'ordinary' learning problems by utilizing earlier examples to generalize to later ones, i.e., in-context learning. |
Copied to clipboard
| Challenge: | Existing methods for developing compact and efficient large language models lack token-level dependencies and linguistic diversity. |
| Approach: | They propose a logits-based fine-tuning framework that integrates supervised learning and knowledge distillation to build enriched training targets using teacher logits and ground truth labels. |
| Outcome: | The proposed method outperforms existing methods on a large-scale logits dataset and a series of science-focused models. |
Copied to clipboard
| Challenge: | Prior studies have examined the impact of structured output on LLMs’ generation quality, often presenting one-way findings. |
| Approach: | They propose to derive five potential causal structures characterizing the influence of structured output on LLMs’ generation using one assumed and two guaranteed constraints. |
| Outcome: | The proposed pipeline can be extended to other modules and is not limited to structured output but can be used in industrial applications. |
Copied to clipboard
| Challenge: | Semantic parsers rely on accurate and high-coverage lexicons, but they often use annotated logical forms to learn the lexic. |
| Approach: | They propose a semi-supervised learning framework that makes use of large text corpora and lexical resources. |
| Outcome: | The proposed framework improves on two benchmarks: Webquestions and Free917. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have enabled them to process increasingly longer sequences, ranging from 2K to 2M tokens and even beyond. |
| Approach: | They propose a synthetic dataset in the financial domain that integrates Chain-of-Thought reasoning into LLMs in a supervised manner to facilitate effective long-context understanding. |
| Outcome: | The proposed model outperforms standard GPT-4o-mini on the Loong benchmark and fine tunes LLaMA-3.1-8B-Instruct on the model, achieving a 28.0% gain on the financial subset. |
Copied to clipboard
| Challenge: | Visual speech processing requires context modeling due to the ambiguous nature of lip movements. |
| Approach: | They propose a framework to maximize the context modeling capability by bringing the power of LLMs. |
| Outcome: | The proposed framework maximizes the power of visual speech processing by bringing it to the forefront of the field. |
Copied to clipboard
| Challenge: | Conventional "closed-world" information extraction methods rely on human ontologies to define scope for extraction. |
| Approach: | They propose a type abstraction approach where models are prompted to generalize and name the type . they use the similarity between inferred names to induce clusters . |
| Outcome: | The proposed method is complementary to token representations on relation extraction and event extraction datasets. |
Copied to clipboard
| Challenge: | Despite advances in large language models, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification. |
| Approach: | They propose a multi-agent debate framework that reformulates misinformation detection as a structured adversarial debate based on fact-checking workflows . |
| Outcome: | The proposed framework enables iterative refinement of evidence while improving decision transparency. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but the complexity of emerging tasks and higher performance demands highlight the need for continuous improvement. |
| Approach: | They propose a method that refines evaluation results and characterizes model profiles at the knowledge component level. |
| Outcome: | The proposed method improves performance across multiple benchmarks and academic exams. |
Copied to clipboard
| Challenge: | Aspect Sentiment Triplet Extraction (ASTE) is a thriving research area . current code-switching methods suffer from term boundary detection issues and out-of-dictionary problems. |
| Approach: | They propose a test-time code-switching framework which bridges the gap between bilingual training and monolingual test- time prediction. |
| Outcome: | The proposed framework achieves an average improvement of 3.7% on four cross-lingual datasets. |
Copied to clipboard
| Challenge: | Recent advances establish "SFT-then-RL" as the defacto paradigm for enhancing large reasoning mod- els on automatically verifiable tasks. |
| Approach: | They propose an entropy-preserving SFT method to enhance exploration capabilities through intrinsic curiosity. |
| Outcome: | The proposed method outperforms the vanilla method on reasoning tasks by 2.5 points . it also outperformed the vanilla SFT by 2.9 points on out-of-distribution tasks . |
Copied to clipboard
| Challenge: | Existing methods focus on designing efficient multimodal fusion frameworks to bridge the semantic gap between images and texts. |
| Approach: | They propose a covariance matrix-driven image channel allocation method that expands the number of original channel maps and assigns importance scores to the expanded channel maps. |
| Outcome: | The proposed method achieves state-of-the-art on three public multimodal fake news detection benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to reducing group bias do not account for correlations between author demographics and linguistic variables, limiting their effectiveness. |
| Approach: | They extend a method for countering group bias using balanced training by balancing each demographic group in training and using protected attributes as input. |
| Outcome: | The proposed model outperforms all other methods when combined with balanced training. |
Copied to clipboard
| Challenge: | Open Domain Question Answering (ODQA) is a longstanding task in Natural Language Processing that involves generating an answer solely based on a given question. |
| Approach: | They propose a novel approach that executes sentence selection on the encoded passages to enhance the inference speed while reducing the context length required for generating answers. |
| Outcome: | The proposed approach can increase inference speed by **2.3X-5.7X** while maintaining the model’s performance. |
Copied to clipboard
| Challenge: | Existing approaches to training LLMs at ultra-low precisions suffer from convergence instability and substantial training costs. |
| Approach: | They propose a progressive QAT framework with outlier channel splitting to address these issues . they use nested structure of integer quantization grids to enable a "train once, deploy any precision" paradigm . |
| Outcome: | The proposed framework outperforms baselines on both Llama2/3 and W2A16, with an 11 speedup over BF16. |
Copied to clipboard
| Challenge: | Existing hierarchical text classification methods make local decisions regarding labels or ignore hierarchy information during inference. |
| Approach: | They propose to learn a Label Assignment Policy via deep reinforcement learning to determine where to place an object and when to stop the assignment process. |
| Outcome: | The proposed method outperforms state-of-the-art methods on five datasets and four base models and achieves an average improvement of 33.4% over flat classifiers. |
Copied to clipboard
| Challenge: | We present a new information extraction system that can construct temporal event graphs from news documents. |
| Approach: | They propose a temporal event graph extraction system that can extract news documents . they extend the system from sentence-level event extraction to cross-document cross-media event extraction . |
| Outcome: | The proposed system can extract temporal event graphs from news documents in multiple languages and multiple data modalities. |
Copied to clipboard
| Challenge: | Existing approaches to QA using retrieval-augmented knowledge are limited by limited coverage and noisy information. |
| Approach: | They propose an induction-augmented generation framework that utilizes inductive knowledge along with retrieved documents for implicit reasoning. |
| Outcome: | The proposed framework outperforms RAG and ChatGPT on two Open-Domain QA tasks. |
Copied to clipboard
| Challenge: | Tabular data analysis is performed everyday across various domains. |
| Approach: | They propose to use a dataset of 467k tables with supervision labels for four types of field metadata. |
| Outcome: | The proposed framework improves the understanding capability of tabular models by incorporating distribution and knowledge information. |
Copied to clipboard
| Challenge: | Existing methods for interlingual homograph recognition require linguistic knowledge and massive annotation work. |
| Approach: | They propose an automatic interlingual homograph recognition method based on cross-lingual word embedding similarity and co-occurrence of form-identical words in parallel sentences. |
| Outcome: | The proposed method can make accurate predictions across languages. |
Copied to clipboard
| Challenge: | Existing reward models assume a global reward function, limiting personalization and pluralistic alignment. |
| Approach: | They propose a framework that leverages binary preference datasets to enhance personalized preference learning. |
| Outcome: | The proposed framework captures diverse human preferences without fine-grained annotations and significantly improves personalized preference learning on downstream tasks. |
Copied to clipboard
| Challenge: | Existing methods for integrating spatial layouts with text have limitations . existing methods produce overly long text sequences or lack autoregressive traits of LLMs . |
| Approach: | They introduce Interleaving Layout and Text in a Large Language Model (LayTextLLM) they use OCR-derived text and spatial layouts to integrate with LLMs for document understanding . |
| Outcome: | The proposed model shows an increase in performance in KIE and VQA tasks. |
Copied to clipboard
| Challenge: | Existing black-box jailbreak methods often rely on model feedback . existing methods may be intercepted by content moderators during the search process . |
| Approach: | They propose a method that guides malicious prompt construction by local training a mirror model of the target black-box model through benign data distillation. |
| Outcome: | The proposed method achieves a 92% attack success rate and 80% stealth rate on a subset of AdvBench. |
Copied to clipboard
| Challenge: | Retrieval Augmented Generation (RAG) is a non-parametric approach for large language models. |
| Approach: | They propose a framework that shifts from triples to context-rich propositions and introduces an efficient, LLM-free online beam search over proposition paths to discover multi-step reasoning chains. |
| Outcome: | The proposed framework achieves state-of-the-art zero-shot Recall@5 and F1 scores on 2Wiki, HotpotQA, and MuSiQue. |
Copied to clipboard
| Challenge: | Existing language model agents excel in planning and reasoning, but lack creativity in unfamiliar environments. |
| Approach: | They propose a benchmark suite of room escape game environments to challenge agents with creative reasoning, unconventional tool use and iterative problem-solving to uncover implicit goals. |
| Outcome: | The proposed framework can perform with 40% fewer steps and hints and performs robustly across difficulty levels. |
Copied to clipboard
| Challenge: | Existing fact-checking methods that use large language models often generate subtle factual errors. |
| Approach: | They propose a fact-checking framework that uses extracted knowledge graphs to enhance text representation. |
| Outcome: | GraphCheck outperforms existing specialized fact-checkers on seven benchmarks spanning general and medical domains . Graph Neural Networks process extracted knowledge graphs as a soft prompt, enabling efficient fact- checking in a single inference call. |
Copied to clipboard
| Challenge: | Existing large language models are pre-trained on unstructured data, which leads to poor performance when dealing with structured data. |
| Approach: | They propose a framework to train large language models to act as verifier modules and to apply iterative corrections offline. |
| Outcome: | The proposed framework improves graph-based generative capability of large language models by iterating corrective instructions on three graph-derived datasets. |
Copied to clipboard
| Challenge: | Existing gradient-based attribution methods are inapplicable to adversarial attacks . et al.: Targeted neuron tuning improves model robustness against jailbreak attacks despite the model's vulnerability to jailbreak. |
| Approach: | They propose a gradient-based method to identify key neurons sensitive to adversarial behaviors in open-ended generation tasks. |
| Outcome: | The proposed method detects key neurons sensitive to adversarial behaviors in open-ended tasks. |
Copied to clipboard
| Challenge: | Existing methods to learn new relations with limited labeled data are prone to catastrophic forgetting and overfitting. |
| Approach: | They propose a framework that uses prompts to acquire more generalized knowledge . they propose CFRE to continuously learn new relations while retaining knowledge of old ones . |
| Outcome: | The proposed method outperforms state-of-the-art methods by a large margin and significantly mitigates catastrophic forgetting and overfitting in low-resource scenarios. |
Copied to clipboard
| Challenge: | Code large language models (LLMs) are becoming tool-interactive agents . quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data . et al.: a new approach to scale trajectory diversity improves tool-use generalization . |
| Approach: | They propose a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. |
| Outcome: | Experiments on general tool-use benchmarks and code agent tasks show that TDScaling improves tool-user generalization and inherent coding proficiency. |
Copied to clipboard
| Challenge: | Speculative decoding is a widely used method that accelerates the generation process of large language models (LLMs) drafting efficiency has become a bottleneck in the final speedup of speculative drafting, therefore generating longer drafts at less cost can lead to better speedup. |
| Approach: | They propose a method that uses existing model to drafting and target LLM to verify draft in a low-cost parallel manner. |
| Outcome: | The proposed method can achieve speedups of up to 2.4 over speculative decoding and 3.9 over vanilla decoding without fine-tuning draft and target models. |
Copied to clipboard
| Challenge: | Existing studies on pre-trained Transformers show that they learn fine-grained neuron functions. |
| Approach: | They examine the presence of modularity in pre-trained Transformers . they focus on Mixture-of-Experts, a promising candidate for modularity . |
| Outcome: | The proposed structure stabilizes at the early stage, which is faster than neuron stabilization. |
Copied to clipboard
| Challenge: | Language models (LMs) often struggle to pay enough attention to the input context, and generate texts that are unfaithful or contain hallucinations. |
| Approach: | They propose a context-aware decoding technique that amplifies the difference between the output probabilities when a model is used with and without context. |
| Outcome: | The proposed model significantly improves faithfulness of different LM families including OPT, GPT, LLaMA, and FLAN-T5 for summarization tasks. |
Copied to clipboard
| Challenge: | Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world . |
| Approach: | They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions. |
| Outcome: | The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents. |
Copied to clipboard
| Challenge: | Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents. |
| Approach: | They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments. |
| Outcome: | The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning. |
Copied to clipboard
| Challenge: | Concept editing aims to control specific concepts in large language models (LLMs) however, there is a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods. |
| Approach: | They propose a paradigm where conceptual injection is aligned at the neuron level. |
| Outcome: | The proposed paradigm offers a clear framework and valuable insights for advancing interpretability and controlled generation in large language models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. |
| Approach: | They investigate the existence of code-switching in the pre-training corpus and categorize it into four types within two quadrants. |
| Outcome: | The proposed approach improves performance across benchmarks and representation space. |
Copied to clipboard
| Challenge: | Recent advances in summarization are driven by the availability of large datasets such as the CNN-DailyMail corpus and the New York Times corpus. |
| Approach: | They propose a method for fine-tuning pretrained models for summarization in unsupervised manner . they use Wikipedia data to produce pseudo-summaries which contain characteristics of target dataset . |
| Outcome: | The proposed method achieves state-of-the-art, zero-shot abstractive summarization performance on CNN-DailyMail dataset and compares with other methods on other datasets. |
Copied to clipboard
| Challenge: | Experimental results show that instruction tuning improves zero-shot generalization across various tasks and improves performance of specific tasks. |
| Approach: | They propose a task selection method that leverages instruction information alone to identify relevant tasks and optimize instruction tuning for specific tasks. |
| Outcome: | The proposed method is significantly more efficient than traditional approaches, which require complex measurements of pairwise transferability between tasks or the creation of data samples for the target task. |
Copied to clipboard
| Challenge: | Existing methods to fix erroneous knowledge in Pre-trained Language models experience a performance decline when the number of edits increases. |
| Approach: | They propose a framework that leverages factual information to enhance editing generalization and guide the identification of edits by retrieving related facts from the fact-patch memory. |
| Outcome: | The proposed framework can improve model generalization and accuracy even with thousands of edits. |
Copied to clipboard
| Challenge: | Existing approaches to provide token-level rewards fail to account for varying degrees of preference inherent to each token. |
| Approach: | They propose a reward model that uses a discriminator to assign token-based continuous rewards to each token considering the context. |
| Outcome: | Extensive experiments show that the proposed reward model improves on open-ended language generation benchmarks. |
Copied to clipboard
| Challenge: | Existing studies either overlook temporal shifts or hardly capture rich shifting patterns of both semantic and knowledge. |
| Approach: | They develop a temporal adaptive learning framework that captures temporal shifts . they use medical ontology and other knowledge sources to integrate temporal adaptation . |
| Outcome: | The proposed framework improves classification tasks across multiple domains and domains with knowledge integration. |
Copied to clipboard
| Challenge: | Entity set expansion and synonym discovery are two critical NLP tasks that are often performed separately, without exploring their interdependencies. |
| Approach: | They propose a framework that enables two tasks to mutually enhance each other by including popular entities’ infrequent synonyms into the set, which boosts set expansion recall. |
| Outcome: | The proposed framework can be used to enhance two NLP tasks by including popular entities’ infrequent synonyms into the set, which boosts set expansion recall. |
Copied to clipboard
| Challenge: | Using the structure of a radiology report, we propose a co-training approach to train two machine learning models using the dual views of MRI and CT data. |
| Approach: | They propose a co-training approach where two machine learning models are built upon the Findings and Impression sections and use each other's information to boost performance with massive unlabeled data in a semi-supervised manner. |
| Outcome: | The proposed model outperforms supervised and semi-supervised methods in a public health surveillance study and outperformed existing methods. |
Copied to clipboard
| Challenge: | Complex news events require swift responses from government and society, authors say . relying on historical events to project the future is insufficient, they say - a simulator for complex news events is needed . |
| Approach: | They propose a controllable complex news event simulator guided by event schema and user-provided assumptions . they incorporate a geo-diverse commonsense and cultural norm-aware knowledge enhancement component . |
| Outcome: | The proposed simulator achieves higher coherence and appropriateness than existing models. |
Copied to clipboard
| Challenge: | Existing approaches to address class imbalance and data difficulty have been used to train models. |
| Approach: | They propose a framework that prioritizes challenging samples and minority classes over hard examples and their semantically similar neighbors to address class imbalance. |
| Outcome: | The proposed framework outperforms baselines on six imbalanced datasets and achieves substantial improvements for minority classes. |
Copied to clipboard
| Challenge: | Existing supervised fine-tuning datasets are composed of general instructions without userspecified constraints. |
| Approach: | They propose a data augmentation method incorporating multiple constraints into the original data samples according to predefined rules to create new training tasks. |
| Outcome: | The proposed method improves LLM controllability while maintaining general instruction-following capabilities. |
Copied to clipboard
| Challenge: | Mandarinograd is a corpus of Winograd Schemas in Mandarin Chinese . WS are hard to collect and few datasets are publicly available . |
| Approach: | They introduce a corpus of Winograd Schemas in Mandarin Chinese . they describe the difficulties faced when building the corpus and explain how they overcome the anomalies. |
| Outcome: | The proposed corpus of Winograd Schemas in Mandarin Chinese is hard to build and resistant to statistical methods. |
Copied to clipboard
| Challenge: | Existing approaches to parsing are greedy transition-based and globally optimized . however, the decision-making process is based on local information, causing error propagation to subsequent steps. |
| Approach: | They propose hierarchical pointer network parsers and apply them to dependency and sentence-level discourse parsing tasks. |
| Outcome: | The proposed method outperforms existing methods and sets new state-of-the-art methods on benchmark datasets. |
Copied to clipboard
| Challenge: | PersLEARN is a tool designed to facilitate the cultivation of scientific perspectives . junior researchers struggle to identify the perspectives reflected in the literature and struggle to develop their own viewpoints. |
| Approach: | They propose a tool to facilitate the cultivation of scientific perspectives by interacting with a prompt-based model and allowing students to develop their own perspectives explicitly. |
| Outcome: | The proposed tool outperforms baseline approaches across multiple domains of literature from different perspectives. |
Copied to clipboard
| Challenge: | Existing methods to improve the robustness of text classification models are token-, sentence-, and hiddenlevel augmentation. |
| Approach: | They propose an interpolation-based data augmentation approach called DoubleMix to improve the robustness of text classification models by learning the “shifted” features in hidden space. |
| Outcome: | The proposed approach outperforms several popular methods on six text classification benchmark datasets and visual analysis shows that the model features are highly interpretable. |
Copied to clipboard
| Challenge: | Existing studies on story generation focus on coarse-grained control of the story, neglecting the details of the narrative. |
| Approach: | They propose a model for fine-grained control on the story that allows the generation of customized stories with characters, corresponding actions and emotions arbitrarily assigned. |
| Outcome: | The proposed method has strong controllability to generate customized stories according to the fine-grained personalized guidance. |
Copied to clipboard
| Challenge: | Existing methods for event prediction are incomplete and noisy. |
| Approach: | They propose to use news-related event schemas to extract newsworthy events . they build a demo website and include a video demonstrating the framework . |
| Outcome: | The proposed framework can be applied to a wide variety of newsworthy scenarios. |
Copied to clipboard
| Challenge: | Existing evaluation methods focus on fluency and factual reliability, while neglecting figurative quality. |
| Approach: | They propose a set of human evaluation metrics focused on the translation of figurative language and a parallel metaphor corpus generated by post-editing. |
| Outcome: | The proposed evaluation protocol estimates four aspects of MT: Metaphorical Equivalence, Emotion, Authenticity, and Quality. |
Copied to clipboard
| Challenge: | Existing resources for training neural models to finely classify mental-health stigma are limited, relying primarily on social media or synthetic data without theoretical underpinnings. |
| Approach: | They propose to use an expert-annotated corpus of human-chatbot interviews to finely classify mental-health stigma. |
| Outcome: | The proposed corpus can facilitate research on computationally detecting, neutralizing, and counteracting mental-health stigma. |
Copied to clipboard
| Challenge: | Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning algorithm for large-scale language models. |
| Approach: | They conduct a systematic study of Low-Rank Adaptation (LoRA) on diverse tasks and rich resources with different learning capacities. |
| Outcome: | The proposed algorithm can achieve remarkable performance in high-resource and multi-task scenarios, even comparable to full fine-tuning. |
Copied to clipboard
| Challenge: | Empirical results show that even the most competitive few-shot learning models struggle on this task, especially as compared with humans. |
| Approach: | They propose a Few-Shot Relation Classification Dataset consisting of 70, 000 sentences on 100 relations derived from Wikipedia and annotated by crowdworkers. |
| Outcome: | The proposed methods perform well on the most competitive few-shot learning models, especially as compared with humans. |
Copied to clipboard
| Challenge: | Existing systems treat this task as a pipeline of two separate subtasks, i.e., event extraction and temporal relation classification. |
| Approach: | They propose a joint event and temporal relation extraction model with shared representation learning and structured prediction. |
| Outcome: | The proposed method improves both event extraction and temporal relation extraction over state-of-the-art systems. |
Copied to clipboard
| Challenge: | Existing studies have shown that pre-trained langauge models tend to memorize and regenerate segments of their pre-training corpus when prompted appropriately. |
| Approach: | They conduct the first comprehensive analysis to explore language models’ memorization during fine-tuning across tasks. |
| Outcome: | The proposed analysis shows that memorization presents a strong disparity among different fine-tuning tasks. |
Copied to clipboard
| Challenge: | Large language models have shown impressive capabilities on code generation tasks. |
| Approach: | They propose a metric that combines functional correctness and syntactic similarity to measure the productivity gains generated by large language models. |
| Outcome: | The proposed model achieves a 14% stronger correlation with value and better represents real-world gains when evaluating and comparing models. |
Copied to clipboard
| Challenge: | Existing methods for relation extraction only implicitly learn to model relevant contexts and entity types while being trained for RE. |
| Approach: | They propose to explicitly teach the model to capture relevant contexts and entity types by supervising and augmenting intermediate steps (SAIS) for RE. |
| Outcome: | The proposed method outperforms the runner-up method on three benchmarks by 5.04% . textual contexts and entity types are the major information sources that lead to the success of previous approaches. |
Copied to clipboard
| Challenge: | Neural network pruning disrupts LLMs’ internal activation features crucial for lie detection . layer-wise pruning sparsity inadvertently removes crucial weights, failing to improve lie detection performance despite its reliance on the most crucial LLM layer. |
| Approach: | They propose a pruning approach that places greater emphasis on layers with more activation outliers and stronger discriminative features simultaneously. |
| Outcome: | The proposed approach improves the hallucination detection for pruned LLMs (achieving 88% accuracy at 50% sparsity) and enhances their performance on TruthfulQA. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated strong reasoning capabilities across various tasks. |
| Approach: | They propose a data-centric approach that enhances LLMs’ awareness of symmetry in query variations and propose syMmetry-ENhanceD (MEND) data augmentation. |
| Outcome: | Extensive experiments on logical and arithmetic reasoning tasks show that the proposed approach improves model robustness at the knowledge extraction stage through query augmentation. |
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts. |
| Approach: | They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
| Outcome: | The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
Copied to clipboard
| Challenge: | Large Speech Language Models (LSLMs) typically operate at high token rates to ensure acoustic fidelity, yet this results in sequence lengths that exceed the underlying semantic content, incurring prohibitive inference costs. |
| Approach: | They propose a token-based token merging mechanism that uses a training-free token pooling mechanism to reduce prefilling FLOPs by 27.48% while maintaining competitive accuracy. |
| Outcome: | The proposed method reduces prefilling FLOPs by 27.48% while maintaining competitive accuracy. |
Copied to clipboard
| Challenge: | Existing work neither proves that pre-trained models successfully learn the injected factual knowledge nor proves there is a causal relation between injected knowledge and downstream performance improvements. |
| Approach: | They propose a counterfactual-based analysis framework to explore the causal effects of factual knowledge injection on the performance of language models within pretrain-finetune paradigm. |
| Outcome: | The proposed framework shows that factual knowledge injection is successful but correctness of injected knowledge only has limited effect on the models’ downstream performance. |
Copied to clipboard
| Challenge: | Existing methods to generalize from seen intents to unseen intents are not effective . Xian et al., 2019: a novel approach to generalized zero-shot intent detection is needed . |
| Approach: | They propose a pairwise prompt-based tuning model with parameter efficient fast adaptation . they leverage hybrid contrastive learning in discriminant space and masked language modeling . |
| Outcome: | The proposed model can generalize to unseen intents with the help of seen intents . the proposed model is based on a pairwise prompt-based tuning model with fast adaptation . |
Copied to clipboard
| Challenge: | Existing work on semantic role labels ignores the semantic connection between the two tasks . et al. (2010) defined two types of semantic roles: core roles and non-core roles. |
| Approach: | They propose to use machine reading comprehension to bridge the gap between these two tasks . they formalize predicate disambiguation as multiple-choice machine reading understanding . |
| Outcome: | The proposed framework achieves state-of-the-art or comparable results to previous work . it uses the descriptions of candidate senses of a given predicate as options to select the correct sense . |
Copied to clipboard
| Challenge: | Recent studies for relation extraction (RE) leverage the dependency tree of the input sentence to improve performance. |
| Approach: | They propose to use a graph convolutional network to build a context graph without dependency parsers. |
| Outcome: | The proposed approach improves neural RE methods without dependency parsers on English benchmark datasets. |
Copied to clipboard
| Challenge: | Existing methods such as GRPO often break down when task difficulty exceeds the model’s capacity, resulting in sparse rewards and inefficient training. |
| Approach: | They propose to measure the compatibility between external guidance and a model's intrinsic policy by introducing an adaptive framework to enhance reasoning performance while explicitly preserving high Affinity. |
| Outcome: | The proposed framework outperforms baseline models while maintaining high Affinity. |
Copied to clipboard
| Challenge: | Experiments on GPT and other 23 LLMs indicate that tokens widely exist while GPT’s vocabulary behaves the worst: more than 23% long Chinese tokens (i.e., a token with more than two Chinese characters) are either porn or online gambling. |
| Approach: | They propose to locate Polluted Chinese (PoC) tokens in LLMs and build a PoC token detector to label them in vocabularies by considering each token’s semantics and related contents from the search engines. |
| Outcome: | The proposed method predicts that the ratio of “*” related webpages in GPT-4o's training data is around 0.5%. |
Copied to clipboard
| Challenge: | generative (tokenby-token) inference is memory-bound and requires a large amount of memory to perform. |
| Approach: | They propose a lookup table engine for weight-quantized large language models that uses offline restructuring of the quantized weight matrix to minimize bit manipulations associated with unpacking. |
| Outcome: | The proposed kernel can be 2-4x faster than existing GEMM kernels while achieving performance gains of 1.5 to 2 times. |
Copied to clipboard
| Challenge: | Existing evaluation methods are inadequate to evaluate large language models (LLMs). |
| Approach: | They propose a fine-grained generative LLM evaluator with instance-level customazable evaluation criteria that can be used to evaluate large language models. |
| Outcome: | The proposed model outperforms existing LLM evaluators and instruction-tuned LLMs on multiple benchmarks and sets new SOTA results. |
Copied to clipboard
| Challenge: | Minority languages in China face significant challenges due to their unique writing systems, which differ from international standards. |
| Approach: | They propose a dataset specifically curated for headline generation tasks for minority languages in China . they propose 50,000 entries each for Uyghur and Mongolian, and a test set annotated by native speakers . |
| Outcome: | The proposed dataset will help improve headline generation in minority languages . it includes 100,000 entries for Tibetan, 50,000 entries each for Uyghur and Mongolian . |
Copied to clipboard
| Challenge: | Recent studies show that character substitutions in toxic Chinese text can confuse state-of-the-art LLMs. |
| Approach: | They propose a taxonomy of 3 perturbation strategies and 8 specific approaches in Chinese text to assess if they can detect perturbed Chinese toxic contents. |
| Outcome: | The proposed model can detect perturbed Chinese text with 8 different approaches . the proposed model is compared with 9 other LLMs from the US and China . |
Copied to clipboard
| Challenge: | Retrieval-augmented generation (RAG) is a powerful technique to facilitate language model generation with proprietary and private data, where data privacy is . a privacy issue that is currently under-explored, is posed by RAG. |
| Approach: | They propose to use retrieval-augmented generation (RAG) to facilitate language model generation with proprietary and private data where data privacy is a pivotal concern. |
| Outcome: | The proposed attack methods demonstrate that RAG can mitigate the old risks, i.e., leakage of the LLMs’ training data. |
Copied to clipboard
| Challenge: | Existing studies in retrieval-augmented generation (RAG) do not sufficiently address the design of complex engineering solutions. |
| Approach: | They propose a retrieval-augmented generation system that leverages tree-based exploration and bi-point thinking mechanism to generate reliable solutions. |
| Outcome: | Experiments show that the proposed system achieves state-of-the-art (SOTA) performance on the SolutionBench, highlighting its potential to enhance the automation and reliability of complex engineering solution design in real-world applications. |
Copied to clipboard
| Challenge: | Medical large language models exhibit high domain specificity and condensed semantics, making them vulnerable to diagnostic errors in real-world clinical settings. |
| Approach: | They propose a framework for modeling and inducing multi-turn medical semantic jailbreaks in clinical dialogues. |
| Outcome: | Experiments on chest X-ray-based multimodal medical dialogues show that MSIA outperforms existing jailbreak methods with an average success rate of 76.67%. |
Copied to clipboard
| Challenge: | Existing approaches to generating reward models rely on voting-based mechanisms to evaluate CoT outputs. |
| Approach: | They propose an efficient generative reward modeling framework grounded in model-internal uncertainty. |
| Outcome: | The proposed framework reduces inference cost while improving answer accuracy. |
Copied to clipboard
| Challenge: | MLLMs lack visual grounding mechanism to read text embedded in images, or rely on parametric shortcuts . despite strong OCR capabilities, models suffer performance degradation of 12.7% in the VQ setting . |
| Approach: | They propose a plug-and-play training strategy that invalidates shortcuts in text prompts . they propose 'vq' setting where text queries are rendered directly onto images . |
| Outcome: | The proposed training strategy surpasses the base model by 5.4% and GRPO based on original images by 2.7% on four representative OOD benchmarks. |
Copied to clipboard
| Challenge: | Existing studies highlight that large language models are receptive to external information that contradicts their parametric knowledge, but little research has been conducted on the direct impact of instruction-tuning on this phenomenon. |
| Approach: | They examine how instruction-tuning influences LLMs' susceptibility to misinformation, particularly in knowledge conflict situations. |
| Outcome: | The proposed model is more user-oriented and more likely to accept misinformation when it is presented by the user. |
Copied to clipboard
| Challenge: | Large language models (LLMs) often produce unnecessarily long explanations that reduce efficiency. |
| Approach: | They propose a length-aware reward that selectively penalizes insignificance tokens . they also propose 'dynamic length control' that encourages more detailed reasoning . |
| Outcome: | The proposed method reduces response length while maintaining correctness, the authors show . it selectively penalizes insignificance tokens while maintaining accuracy . |
Copied to clipboard
| Challenge: | Existing EE methods do not model event characteristics from large unsupervised data. |
| Approach: | They propose a contrastive pre-training framework for event extraction to better learn event knowledge from large unsupervised data and their semantic structures. |
| Outcome: | The proposed framework improves on ACE 2005 and MAVEN datasets on event extraction tasks. |
Copied to clipboard
| Challenge: | Existing methods to enhance performance of large language models (LLMs) on Text-to-SQL tasks rely on execution-based or LLM-based reward models. |
| Approach: | They propose a reward model framework for RL-based Text-to-SQL that employs the GMNScore outcome reward model. |
| Outcome: | The proposed reward model outperforms existing reward models on standard benchmarks including Spider and BIRD. |
Copied to clipboard
| Challenge: | Existing methods for few-shot intent detection are limited due to data scarcity and lack of information for unseen domains. |
| Approach: | They propose to enhance utterance representations with label synset augmentation and refine prototypes by distilling coarse domain knowledge from a universal teacher model. |
| Outcome: | The proposed approach outperforms existing methods in terms of accuracy and generalization across domains. |
Copied to clipboard
| Challenge: | Recent multimodal large language models lack robust audio-visual integration ability and performance on DeafTest is highly correlated with AV-Odyssey accuracy. |
| Approach: | They propose a benchmarking tool that integrates audio-visual reasoning with audio-video cues to infer solutions. |
| Outcome: | The proposed model performs well on DeafTest, but lacks audio perception in simple audio tasks. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) is effective for aligning Large Language Models with human preferences, but its complex process limits its ability to continually learn human feedback. |
| Approach: | They propose a non-RL offline method to convert historical optimal policies into optimization constraints when continually learning new preferences. |
| Outcome: | The proposed method outperforms strong CL baselines in terms of reward-based evaluations and human assessment. |
Copied to clipboard
| Challenge: | Recent approaches to sequence labeling have been based on statistical models but a challenge is from the data sparsity problem. |
| Approach: | They propose to use local context reconstruction to implicitly incorporate contextual information into their representations. |
| Outcome: | The proposed model outperforms all previous methods on multiple benchmark datasets and achieves new start-of-the-art results. |
Copied to clipboard
| Challenge: | a central challenge remains balancing text quality against detection robustness. |
| Approach: | They propose a framework that aligns watermark strength with linguistic degrees of freedom . they use part-of-speech models to weaken the signal in grammatically constrained contexts . |
| Outcome: | The proposed framework outperforms existing methods in linguistic indeterminacy tests on languages . it weakens the watermark strength in grammatically constrained contexts and strengthens it in contexts with greater linguistic flexibility. |
Copied to clipboard
| Challenge: | Large language models have achieved remarkable success across a wide range of tasks, yet their performance remains heavily biased toward high-resource languages. |
| Approach: | They propose a pipeline for advancing Tibetan language modeling through multilingual continual pre-training with Tibetan, Chinese, and English. |
| Outcome: | The proposed model outperforms open-source and Tibetan-focused models on diverse tasks. |
Copied to clipboard
| Challenge: | Existing work on slot filling uses labeled data from source domains to train a model for target domains. |
| Approach: | They propose a model-agnostic Slot Transferability Measure (STM) to evaluate the transferability from a source slot to a target slot. |
| Outcome: | The proposed method outperforms state-of-the-art models on multiple datasets and models. |
Copied to clipboard
| Challenge: | In-context learning (ICL) has emerged as a capability of large language models (LLMs) but there is limited understanding of its vulnerability against data poisoning attacks. |
| Approach: | They propose an attack method that exploits ICL’s unique learning mechanisms by identifying discrete text perturbations that influence LLM hidden states. |
| Outcome: | The proposed attack method exploits ICL’s learning mechanisms by identifying discrete text perturbations that influence LLM hidden states. |
Copied to clipboard
| Challenge: | Existing supervised neural methods are underexplored for coreference resolution, especially in incremental clustering. |
| Approach: | They propose a dual-threshold incremental clustering approach based on a lightweight Transformer. |
| Outcome: | Experiments on common benchmarks show that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints. |
Copied to clipboard
| Challenge: | MultiPL is a special case of multiple natural languages and requires limited computational resources to generate multilingual code. |
| Approach: | They propose to extend LLMs by combining two paired experts to optimize expert selection at token and segment levels. |
| Outcome: | The proposed extension improves the performance of the base LLMs while retaining the most popular ones using limited computational resources. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding unexplored. |
| Approach: | They introduce a benchmark specifically designed to assess OSC performance in T2V models. |
| Outcome: | The proposed benchmark assesses the performance of open-source and proprietary T2V models on object state change (OSC) in the context of novel and compositional scenarios. |
Copied to clipboard
| Challenge: | Existing methods to generalize hate speech detection models have been limited by the labeling criteria between datasets. |
| Approach: | They propose a framework that uses the concept of multi-agent for hate speech detection that uses a set of labeling criteria to create multiple agents based on the induced labeling of given datasets. |
| Outcome: | The proposed framework achieves superior cross-evaluation performance compared to methods that focus on specific labeling criteria or majority voting methods. |
Copied to clipboard
| Challenge: | Prompt tuning is a technique for adapting large-scale pretrained language models for downstream tasks. |
| Approach: | They propose to condition a frozen pretrained language model with soft prompts from data . they propose to use a domain adaptation technique to regularize the decision boundary . |
| Outcome: | The proposed method outperforms full-model tuning in data-scarce settings by a large margin. |
Copied to clipboard
| Challenge: | Recent research points to knowledge distillation as a potential solution for NLU tasks. |
| Approach: | They propose a training approach that distills large finetuned LMs into a small network using unlabeled training examples. |
| Outcome: | The proposed approach outperforms BERT training approaches while using 300 times fewer parameters. |
Copied to clipboard
| Challenge: | Large language models are deployed in long-horizon tasks that require agents to track interleaved goals, resolve references to prior information, and coordinate actions over extended trajectories. |
| Approach: | They propose an agentic memory system that indexes each trajectory step with a structured retrieval cue, contextual intent, and retrieves history by matching the current step’s intent. |
| Outcome: | The proposed system outperforms the strongest benchmark by 35.6%, with the largest gains as trajectory length increases. |
Copied to clipboard
| Challenge: | Long-context understanding is a critical capability for large language models . evaluating this capability requires extensive human annotation, which is time-consuming and costly. |
| Approach: | They propose a benchmark to assess citation-grounded long-context reasoning in academic writing. |
| Outcome: | The proposed benchmark compares state-of-the-art models with human experts on two tasks . human experts achieve 90% accuracy, but most models struggle with the cloze-style task . |
Copied to clipboard
| Challenge: | Speculative decoding is a novel method to expedite inference in autoregressive (large) language models. |
| Approach: | They propose to use a smaller model as a draft model to speculate a block of tokens, which the target model then evaluates for acceptance. |
| Outcome: | The proposed method can be used to accelerate inference in autoregressive (large) language models by using smaller models as draft models to speculate tokens for multiple inference steps. |
Copied to clipboard
| Challenge: | Existing relevance models rely on query-keyword pairs but keywords are usually short texts with scarce semantic information, which may not accurately reflect the underlying advertising purposes. |
| Approach: | They propose a bidding-graph augmented triple-based relevance model with three towers to deeply fuse the bidding graphs and semantic textual data. |
| Outcome: | The proposed model outperforms existing models on a large industry dataset and consistently outperformed existing models. |
Copied to clipboard
| Challenge: | Existing retrieval methods aim to gather relevant passages but fail to prioritize consistent and useful information for the reader. |
| Approach: | They propose a novel method which re-ranks passages based on the reader's prediction probability distribution and clusters passage according to the predicted answers. |
| Outcome: | The proposed method improves the quality of evidence passages under zero-shot scenarios. |
Copied to clipboard
| Challenge: | Existing methods for manipulation detection and grounding focus on manipulator type classification under result-oriented supervision. |
| Approach: | They propose a reasoning-driven framework that shifts learning from outcome fitting to process modeling. |
| Outcome: | The proposed framework achieves state-of-the-art with superior generalization on large-scale datasets. |
Copied to clipboard
| Challenge: | Clinical texts contain important temporal information, such as medication start and end dates, appointment dates, and diagnosis dates. |
| Approach: | They propose to use prompt-based learning and fine-tuning to classify temporal relations between treatments and hospitalisation periods in discharge summaries. |
| Outcome: | The proposed method identifies whether a treatment was administered between the time of admission and discharge from the hospital. |
Copied to clipboard
| Challenge: | Abstractive conversation summarization systems rely on large-scale annotated summaries, but collecting and annotating these conversations can be time-consuming and labor-intensive. |
| Approach: | They propose a method for generating diverse and high-quality pairs of conversations and summaries by extracting conversation structures and organizing meaningful conversation snippets. |
| Outcome: | The proposed method outperforms baseline methods on SAMSum and DialogSum datasets and achieves a 10% increase in ROUGE scores with limited data. |
Copied to clipboard
| Challenge: | Existing models cannot capture consistency and diversity of relation patterns in different languages. |
| Approach: | They propose an adversarial multi-lingual neural relation extraction model which considers consistency and diversity among languages. |
| Outcome: | The proposed model outperforms the state-of-the-art models on real-world datasets. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have improved Text-to-SQL methods . however, they still face challenges such as complex multi-stage pipelines and poor robustness to noisy schema information. |
| Approach: | They propose a single-stage SFT framework that optimizes schema linking and SQL generation via a unified loss. |
| Outcome: | Experiments on the Spider and BIRD benchmarks show that JOLT-SQL achieves state-of-the-art execution accuracy among comparable-size open-source models. |
Copied to clipboard
| Challenge: | Question decomposition has been found to improve large language models’ (LLMs) performance on complex question answering (QA) however, performance on the task remains dominated by supervised approaches, suggesting room for making LLMs better decomposers. |
| Approach: | They propose to generate synthetic decomposition data with only five annotated examples by extending recent advances in using LLM-as-judge and for reranking in novel ways. |
| Outcome: | The proposed approach generates synthetic decomposition data with only five examples over two benchmark datasets. |
Copied to clipboard
| Challenge: | Existing approaches to inference have been based on stochastic decoding but they sacrifice output quality due to randomness. |
| Approach: | They propose a deterministic decoding scheme, local temperature beam search, which reduces repetition while maintaining the level of coherence as in beam search. |
| Outcome: | The proposed inference scheme reduces repetition while maintaining coherence as in beam search. |
Copied to clipboard
| Challenge: | Existing studies focus on learning global or local correspondence, but lack fine-grained local-global alignment. |
| Approach: | They propose a High Order Semantic Alignment (HOSA) model that can provide complementary and comprehensive semantic clues to infer correlation scores. |
| Outcome: | The proposed model outperforms state-of-the-art models in retrieving the most relevant results. |
Copied to clipboard
| Challenge: | Existing solutions for document QA fail to provide personalized and up-to-date information efficiently. |
| Approach: | They propose to deploy a self-evolving, efficient LLM system that can offer personalized research services, maintaining a real-time updated database. |
| Outcome: | The proposed system saves 69.92% of time after efficient deployment. |
Copied to clipboard
| Challenge: | In the last decade, the maturity achieved by NLP has turned the spotlight to the possibilities offered by Nlp for a variety of novel applications. |
| Approach: | They propose to use NLP to raise the benefits of BPM at different levels . they propose to focus on the daily tasks that an organization must perform . |
| Outcome: | The proposed approach could be applied to a variety of business processes in a scalable fashion. |
Copied to clipboard
| Challenge: | Using mixture-of-memory augmenting to augment language models improves model generalization but with diminishing return. |
| Approach: | They develop a mechanism that augments language models with mixture-of-memory Augmentation (MoMA) they augment strong T5-based retrievers with the option to "plug in" unseen memory at inference time. |
| Outcome: | The proposed model outperforms methods with larger model sizes on the BEIR benchmark and achieves comparable or even better performance than methods relying on target-specific pretraining. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) is a fundamental step in scientific literature analysis to build AI-driven systems for molecular discovery, synthetic strategy designing, and manufacturing. |
| Approach: | They propose an ontology-guided method for fine-grained named entity recognition (NER) it leverages the chemistry type ontologies to generate distant labels with flexible KB-matching . |
| Outcome: | The proposed method significantly outperforms the state-of-the-art methods with a .25 absolute F1 improvement. |
Copied to clipboard
| Challenge: | Existing systems that use long-context modeling incur computational and memory overhead. |
| Approach: | They propose a visual memory framework that pre-rendered text into structured images and stored as visual notes for agentic systems. |
| Outcome: | The proposed system reduces token consumption while preserving effective long-term memory recall. |
Copied to clipboard
| Challenge: | Multi-turn dialogues pose a greater risk than single prompts, but existing safety benchmarks do not account for this situation. |
| Approach: | They propose a benchmark that features dialogues of varying lengths generated from harmful queries accompanied by images. |
| Outcome: | The proposed model reduces multi-turn Attack Success Rate (ASR) compared to existing guard models. |
Copied to clipboard
| Challenge: | . - (EN) |
| Approach: | . - (EN) |
| Outcome: | . - (EN) |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) struggle with capturing long-distance dependencies within sequences to deeply understand semantics. |
| Approach: | They propose a system that captures relevant information within a fixed window size and provides precise answers to queries. |
| Outcome: | The proposed system can read Harry Potter within 30s and accurately answer the questions. |
Copied to clipboard
| Challenge: | a method that extracts experimental procedures from human language into actionable sequences in robotics language is challenging given the complexity of the instructions and context-dependent nature of the instruction. |
| Approach: | They propose a method that converts actions written in natural language into Python code that can be easily translated into robotics language. |
| Outcome: | The proposed method can extract experimental procedures from human language into actionable sequences in robotics language. |
Copied to clipboard
| Challenge: | Existing locate-and-edit knowledge editing methods suffer from two limitations: they are infeasible for large scale KE in practice and require long run-time. |
| Approach: | They propose to use parametric fine-tuning techniques to update obsolete knowledge and induce new knowledge into LLMs. |
| Outcome: | The proposed methods improve the performance of KE and knowledge update in a temporal dataset with knowledge update and knowledge injection examples. |
Copied to clipboard
| Challenge: | Existing information extraction (IE) tasks rely on in-context learning with large language models. |
| Approach: | They propose a Bayesian-based in-context learning framework that refines label representations across IE tasks using particle filtering and Bayes updates. |
| Outcome: | The proposed framework improves performance over existing methods (up to 30%) it underperforms one-shot prompting by a substantial margin on NER tasks and CodeIE fails on RE tasks with near-zero micro-F1. |
Copied to clipboard
| Challenge: | Scientific data visualization is an essential process in research, but its use of large language models remains unexplored. |
| Approach: | They propose a model-agnostic LLM agent framework to automate scientific data visualization tasks. |
| Outcome: | The proposed framework improves performance of commercial and open-source models. |
Copied to clipboard
| Challenge: | Word composition models do not take ambiguity of words and context into consideration for learning representations, and thus suffer from the inaccurate representation of semantics. |
| Approach: | They propose a model-free context-aware word composition model which takes the latent semantic information as global context for learning representations. |
| Outcome: | The proposed model improves on existing composition models at different granularities and shows that it can be used to learn semantics. |
Copied to clipboard
| Challenge: | Existing methods for entity prediction cannot predict when an event will occur . there are many facts not related to the query that can confuse the model . |
| Approach: | They propose a temporal knowledge Graph reasoning model based on Graph Hawkes Transformer . the model captures instantaneous structural and temporal evolution information . |
| Outcome: | The proposed model performs much better under long-term evolution scenarios. |
Copied to clipboard
| Challenge: | Legal documents have complex document layouts involving multiple nested sections and lengthy footnotes that make question answering challenging. |
| Approach: | They propose a question answering system that parses document layouts while isolating sections and footnotes and linking them appropriately. |
| Outcome: | The proposed system can parse complex document layouts while isolating sections and footnotes and linking them appropriately. |
Copied to clipboard
| Challenge: | Instruction Tuning has the potential to stimulate or enhance specific capabilities of large language models. |
| Approach: | They propose a mixture-of-LoRAs architecture which is a parameter-efficient tuning method designed for multi-task learning with LLMs. |
| Outcome: | The proposed method can be iteratively adapted to a new domain, enabling quick domain-specific adaptation. |
Copied to clipboard
| Challenge: | Language-Model-as-a-Services (LMaaSs) support a variety of user tasks through in-context learning from prompts. |
| Approach: | They propose a lightweight automatic prompt generation method that meta-trains a prompt generation model to enable robust learning from the contexts created by the generated prompts. |
| Outcome: | The proposed method improves performance on unseen tasks by 19.4% compared to the state-of-the-art prompt generation method. |
Copied to clipboard
| Challenge: | a new method to detect political bias in news articles overcomes this domain dependency . partisan bias exists in various social issues, including the 2016 presidential election . |
| Approach: | They propose a multi-head hierarchical attention model that encodes the structure of long documents through a diverse ensemble of attention heads. |
| Outcome: | The proposed model outperforms existing methods for detecting political bias in news articles. |
Copied to clipboard
| Challenge: | Existing methods for rewriting text-to-image models require specialized vocabulary . a new approach uses large vision language models to optimize text-based models . |
| Approach: | They propose a prompt optimization framework that rephrases a user prompt into a text-to-image model by using large vision language models as solver and reward model. |
| Outcome: | The proposed model outperforms existing models on two popular datasets. |
Copied to clipboard
| Challenge: | Empirical evaluations show that Mixture of Expert Prompt Tuning outperforms state-of-the-art parameter efficient baselines on SuperGLUE. |
| Approach: | They propose a pretrain-then-fine-tune paradigm for manifold mapping using multiple prompt experts. |
| Outcome: | Empirical results show that the proposed approach outperforms state-of-the-art methods on SuperGLUE while reducing activated prompts by 79.25%. |
Copied to clipboard
| Challenge: | prevailing taxonomies neglect robustness and honesty, yielding safer-on-paper but less useful systems. |
| Approach: | They propose a soft-gating pipeline where a guardian predicts a binary risk label plus a concise explanation and prepends this advice to the original query for re-inference. |
| Outcome: | The proposed model maintains safety while reducing over-refusal. |
Copied to clipboard
| Challenge: | Several methods have been proposed to mitigate bias in training on biased datasets. |
| Approach: | They propose to examine the effect of target class imbalance and stereotyping on model performance by analyzing binary classification, profession prediction and regression tasks. |
| Outcome: | The proposed methods show that data conditions have a strong influence on relative model performance. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are the foundation of modern natural language processing, powering applications across diverse domains. |
| Approach: | They propose a model-agnostic defense framework which aggregates and evaluates the outputs of a knowledge-injected LLM, a base LLM and a dedicated judge model to enhance resistance against membership inference attacks. |
| Outcome: | The proposed framework reduces MIA success by up to 27.8% for SFT and 526.3% for RAG compared to inference-time baseline while maintaining answer quality. |
Copied to clipboard
| Challenge: | a framework for model merging is proposed without additional training . task vectors from fine-tuned models exhibit a limited number of dominant singular values . |
| Approach: | They propose a framework for model merging based on low-rank estimation of task vectors without access to the base model. |
| Outcome: | The proposed framework improves models without additional training without additional inputs. |
Copied to clipboard
| Challenge: | Existing methods to update large language models (LLMs) without expensive retraining are fragile under single-edit evaluation protocols. |
| Approach: | They propose a framework that characterizes activation-based editing as a constrained intervention on intermediate representations. |
| Outcome: | The proposed method reveals local knowledge conflicts invisible to existing benchmarks. |
Copied to clipboard
| Challenge: | Structured chemical reaction information is a vital tool for chemists engaged in laboratory work and advanced endeavors such as computer-aided drug design. |
| Approach: | They propose a method which utilizes frequent patterns within the text as linguistic cues to identify specific characteristics of chemical reactions. |
| Outcome: | The proposed model outperforms baselines and outperformed existing models. |
Copied to clipboard
| Challenge: | Recent dynamic computation methods show that not all components are required for inference, enabling a training-free pipeline. |
| Approach: | They propose a token-position aware layer skipping framework to save 1.5x times operations efficiently while maintaining performance. |
| Outcome: | The proposed algorithm achieves 1.5x speedup on large language models with no retraining and with comparable performance on the GSM8K and BBH benchmarks. |
Copied to clipboard
| Challenge: | Existing methods often apply coarse-grained constraints over entire reasoning trajectories . Existing approaches often apply unsafe constraints, causing unsafe outputs . |
| Approach: | They propose a trajectory-level training framework that mitigates Self-Jailbreak . they propose 'chain-of-guardrail' to mitigate self-jailbreak by targeting step-level interventions . |
| Outcome: | The proposed framework mitigates Self-Jailbreak by targeting step-level interventions while maintaining reasoning ability. |
Copied to clipboard
| Challenge: | Existing approaches to multiple intent detection and slot filling focus on task-specific components to capture the relationships between intents and slots. |
| Approach: | They propose a Unified Generative framework that captures the relationships between intents and slots in an utterance and formulates the task as a question-answering problem. |
| Outcome: | The proposed framework surpasses baselines on full-data and multi-intent benchmarks on 5-shot and 10-shot scenarios. |
Copied to clipboard
| Challenge: | Recent work on code summarization relies on structural information from the abstract syntax tree (AST) of source codes. |
| Approach: | They propose a program dependency graph (PDG) that represents the structure of a code more effectively. |
| Outcome: | The proposed model improves the performance of an out-of-domain benchmark dataset and the measure SBERT score. |
Copied to clipboard
| Challenge: | Existing models ignore ability to skip irrelevant snapshots according to entity-related relations in query . TKGC is difficult and even large-scale pre-trained language models such as gist ignore explicit temporal information. |
| Approach: | They propose a model that leverages explicit temporal embedding as input to skip unnecessary information for prediction. |
| Outcome: | The proposed model outperforms all state-of-the-art models on six datasets . it incorporates skip information flow after each timestamp to skip unnecessary information . |
Copied to clipboard
| Challenge: | Existing studies on relation extraction only take into account intrasentence relationships that contain pairs of entities. |
| Approach: | They propose to capture omitted arguments in relation extraction given a proper knowledge base for entities of interest. |
| Outcome: | The proposed method improves relation extraction quality by capturing omitted arguments in sentences. |
Copied to clipboard
| Challenge: | Existing prompt engineering methods exploit database content and execution feedback to improve text-to-sql performance. |
| Approach: | They propose a framework for large language model-based text-to-sql task that exploits database content and execution feedback to improve execution accuracy. |
| Outcome: | The proposed framework improves execution accuracy and usability by 12.41% and 5.38% on four widely used benchmarks. |
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) struggle with physics problem solving due to difficulties in decoding implicit constraints and maintaining physical consistency. |
| Approach: | They propose a Generative PRM that treats evaluation as a generative task . it produces fine-grained diagnoses comprising critiques, final judgments, and specific error types . |
| Outcome: | The proposed model improves performance across seven benchmarks in Best-of-N and critique refinement strategies. |
Copied to clipboard
| Challenge: | Existing evaluation methods for floor plan generation rely on statistical metrics like FID, GED, and PSNR, which fail to evaluate using domain knowledge. |
| Approach: | They propose to use a first floor plan dataset to train a floor plan generation model based on a multi-dimensional preference score and a textual analysis to integrate architects’ professional expertise and preferences. |
| Outcome: | The proposed model outperforms baseline models in text-conditional and class-condition tasks and is more rational and aligns better with human preferences. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for long-form speech are limited to limited domains, creating a significant gap with the diverse downstream applications. |
| Approach: | They propose a benchmark that decomposes "long-form speech quality" into specific, disentangled dimensions. |
| Outcome: | The proposed benchmark decomposes “long-form speech quality” into specific, disentangled dimensions. |
Copied to clipboard
| Challenge: | InfiMM is a multimodal large language model that adapts to complex vision-language tasks. |
| Approach: | They present a Multimodal Large Language Model that adapts to intricate vision-language tasks using large-scale training data and comprehensive training strategies. |
| Outcome: | Empirical evaluations across a variety of benchmarks underscore InfiMM’s remarkable capability in multimodal understanding. |
Copied to clipboard
| Challenge: | Multi-modal large language models (MLLMs) generate plausible but incorrect content, resulting in hallucinations . recent advances in MLLM technology have demonstrated their outstanding performance in a variety of visual tasks, such as object detection. |
| Approach: | They propose a plug-and-play method which leverages MLLMs’ internal representations to mitigate hallucinations by analyzing input and output tokens. |
| Outcome: | The proposed method exploits MLLMs’ internal representations to mitigate hallucinations. |
Copied to clipboard
| Challenge: | Event information is a type of common sense knowledge that helps people understand how stories evolve and provides predictive hints for future events. |
| Approach: | They propose a temporal event understanding pipeline that integrates state-of-the-art components. |
| Outcome: | The proposed pipeline can be easily adapted to other domains, including biomedical domains. |
Copied to clipboard
| Challenge: | Existing methods for table-text retrieval are limited due to the need to bridge structured tables and unstructured passages. |
| Approach: | They propose a table-text retrieval system that combines the strengths of both approaches . they propose bipartite subgraph retrieval and query-relevant node expansion . |
| Outcome: | The proposed method outperforms state-of-the-art models with a 42.6% and 39.9% improvement on the OTT-QA benchmark. |
Copied to clipboard
| Challenge: | Existing methods for video question answering align visual or textual features directly with large language models, limiting the deep semantic association between modalities and hindering a comprehensive understanding of interactions within spatial and temporal contexts. |
| Approach: | They propose a temporal-aware framework for multi-modal video question answering that aligns videos and questions at fine-grained levels. |
| Outcome: | The proposed framework improves reasoning ability and accuracy of videoQA by aligning videos and questions at fine-grained levels. |
Copied to clipboard
| Challenge: | Current “sample and select” methods rely on majority voting to score answers . however, when tasks have many distinct and valid answers, selection by voting requires a large number of samples. |
| Approach: | They introduce a method that replaces SC's discontinuous scoring with a continuous score computed from model likelihoods to increase selection even when actions are sparsely distributed. |
| Outcome: | The proposed method improves performance and efficiency on long-horizon interactive tasks by replacing SC’s discontinuous scoring with a continuous score computed from model likelihoods. |
Copied to clipboard
| Challenge: | Structured dropout approaches have been investigated to regularize the multi-head attention mechanism in Transformers. |
| Approach: | They propose a new regularization scheme based on token-level rather than structure-level to reduce overfitting by manipulating the connections between tokens in the multi-head attention via masking. |
| Outcome: | The proposed regularization scheme outperforms attention dropout and DropHead on 18 datasets and can establish a new record on the data-to-text benchmark Rotowire (18.93 BLEU). |
Copied to clipboard
| Challenge: | Low-shot relation extraction (RE) aims to recognize novel relations with very few or even no samples. |
| Approach: | They propose a method that leverages triplet paraphrase to pre-train zero-shot label matching ability and uses meta-learning paradigm to learn few-shot instance summarizing ability. |
| Outcome: | The proposed method outperforms strong baselines and achieves the best performance on few-shot RE leaderboard. |
Copied to clipboard
| Challenge: | Existing approaches to optimize Register-Transfer Level (RTL) code fail to simultaneously optimize functional correctness and hardware efficiency metrics such as Power, Performance, and Area (PPA). |
| Approach: | They propose a hierarchical reward based reinforcement learning framework that integrates direct feedback from EDA simulators and synthesis tools into a reward mechanism. |
| Outcome: | The proposed framework integrates direct feedback from EDA simulators and synthesis tools into a hierarchical reward based reinforcement learning framework. |
Copied to clipboard
| Challenge: | Existing models for grounding are unable to understand modified color expressions, such as “light blue”. |
| Approach: | They propose a model that learns more complex transformations in RGB space and a hard ensemble model that selects a color space depending on the modifier-color pair. |
| Outcome: | The proposed model performs better in the HSV color space than the state-of-the-art model. |
Copied to clipboard
| Challenge: | Recent years, pre-trained language models (PLMs) have achieved promising results on various NLP tasks. |
| Approach: | They propose an open-source toolkit for big model inference and tuning which can support big model tuning at extremely low computation cost. |
| Outcome: | The proposed toolkit can support big model inference and tuning at extremely low computation cost. |
Copied to clipboard
| Challenge: | Current outcome-centric verification paradigms neglect potential errors in the derivation process. |
| Approach: | They propose a process-aware RLVR training paradigm utilizing verifiers selected via **PRIME**. |
| Outcome: | The proposed approach outperforms the baseline verification paradigm on AIME24, AIME25, and Beyond-AIME models. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing studies have focused on re-modeling the given NEs and thus lead to inferior results when NE is sometimes ambiguous. |
| Approach: | They propose a relation extraction model with two training stages that uses adversarial multi-task learning to recover the given NEs. |
| Outcome: | The proposed model improves on two English benchmark datasets and shows state-of-the-art performance. |
Copied to clipboard
| Challenge: | Existing methods for MT have problems with translating homographs, as it is difficult to select the correct translation based on the context. |
| Approach: | They propose to model the context of the input word with context-aware word embeddings that help to differentiate the word sense before feeding it into the encoder. |
| Outcome: | The proposed models improve translation accuracy and BLEU score on three language pairs. |
Copied to clipboard
| Challenge: | Existing approaches for local citation recommendation map or translate a query to citation-worthy research papers. |
| Approach: | They propose a local citation recommendation task that uses latent evidence spans to recommend papers . proposed system retrieves ranked lists of evidence span and recommended paper pairs . |
| Outcome: | The proposed system retrieves ranked lists of evidence span and recommended paper pairs based on evidence from the existing literature. |
Copied to clipboard
| Challenge: | Existing vision-language models are not equipped to read diverse languages and scripts found in historical materials. |
| Approach: | They propose to train an open-weight vision-language model for historical text recognition on CHURRO-DS, the largest historical text-recognition dataset to date. |
| Outcome: | The proposed model outperforms existing vision-language models on CHURRO-DS, the largest historical text recognition dataset to date. |
Copied to clipboard
| Challenge: | Existing image captioning models have a lack of diversity between sentences . current models have limited their effectiveness due to repetitive paragraphs . |
| Approach: | They propose to apply sequence-level training to image paragraph captioning models . they find that standard self-critical training produces poor results . |
| Outcome: | The proposed training improves on the Visual Genome dataset with no architectural changes. |
Copied to clipboard
| Challenge: | Existing methods to extract text snippets from input text to support model predictions without explicit rationale annotation have limited their ability to capture meaningful internal correlations between aspects. |
| Approach: | They propose a multi-aspect rationale extractor that extracts text snippets to support model predictions without explicit rationale annotation. |
| Outcome: | The proposed method achieves state-of-the-art on two unsupervised rationale extraction benchmarks. |
Copied to clipboard
| Challenge: | Prompt tuning for pre-trained language models has shown remarkable performance . however, prompt tuning is still not fully explored . |
| Approach: | They propose to pre-train prompts by adding soft prompts into the pre-training stage to obtain a better initialization. |
| Outcome: | The proposed framework outperforms full-model tuning under full-data and few-shot learning settings. |
Copied to clipboard
| Challenge: | Text analysis of tabular data relies on two core operations: summarization for corpus-level theme extraction and tagging for row-level labeling. |
| Approach: | They propose a framework that enhances output stability by constraining the model’s latent reasoning trajectory. |
| Outcome: | The proposed framework improves stability by constraining the model's latent reasoning trajectory. |
Copied to clipboard
| Challenge: | Linguistic steganography studies how to hide secret messages in natural language cover texts. |
| Approach: | They propose a method which encodes secret messages using self-adjusting arithmetic coding based on a neural language model. |
| Outcome: | The proposed method outperforms the state-of-the-art methods on four datasets by 15.3% and 38.9% in terms of bits/word and KL metrics. |
Copied to clipboard
| Challenge: | a new framework to digest relevant biomedical knowledge is needed to combat COVID-19 . quantity of research results is a bottleneck, and false information promoted in publications . |
| Approach: | a team of researchers has developed a framework to extract multimedia knowledge elements from scientific literature to combat COVID-19. |
| Outcome: | a new framework extracts fine-grained multimedia knowledge elements from scientific literature . it provides detailed contextual sentences, subfigures, and knowledge subgraphs as evidence . the framework is based on a case study of drug repurposing . |
Copied to clipboard
| Challenge: | Existing work on online abusive language detection focused on detecting a single abusive language problem in a domain, like Twitter, but none of them was successfully transferable to general ALD in different online communities. |
| Approach: | They propose a generic ALD framework that can address multiple types of ALD tasks across different domains and use a textual graph embedding to analyse the user’s linguistic behaviour. |
| Outcome: | The proposed framework surpasses the current state-of-the-art ALD algorithms across seven datasets covering multiple aspects of abusive language and different online community domains. |
Copied to clipboard
| Challenge: | Existing ODQA datasets consist mainly of Wikipedia corpus, and are insufficient to study models’ generalizability across diverse domains. |
| Approach: | They propose a benchmark to evaluate ODQA's domain robustness using Wikipedia corpus . they annotate QA pairs in retrieval datasets with rigorous quality control . |
| Outcome: | The proposed benchmark improves model performance on annotated QA pairs in retrieval datasets with rigorous quality control. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a promising foundation for literature surveys, but guiding them to generate accurate, reliable content remains a fundamental challenge. |
| Approach: | They propose a feedback-driven framework that incorporates feedback across three dimensions: outline feedback for structural clarity, citation feedback for evidence validation, and content feedback for readability and analytical depth. |
| Outcome: | The proposed framework significantly improves both citation and content quality, demonstrating feedback as the critical mechanism for automatic survey generation. |
Copied to clipboard
| Challenge: | Large Language Models excel at understanding the semantic relationships between queries and documents, even with lengthy and complex long-tail queries. |
| Approach: | They propose an efficient label generation pipeline and novel sLLM training methods for both encoder and decoder models. |
| Outcome: | The proposed method improves re-ranking for long-tail queries on a Korean-based search platform. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown great ability in solving traditional natural language tasks and elementary reasoning tasks with appropriate prompting techniques. |
| Approach: | They propose a collaborative multi-agent, multi-reasoning-path prompting framework that prompts LLMs to play different roles in a problem-solving team and encourages different role-play agents to collaboratively solve the target task. |
| Outcome: | The proposed framework is applied to two college-level science problems over competitive baselines. |
Copied to clipboard
| Challenge: | Existing methods for long-video inference use compression or sparse attention . existing methods restrict LMMs from handling longer, more complex videos . |
| Approach: | They propose a sequence-parallel framework with optimized attention that accelerates long-video inference across multiple GPUs. |
| Outcome: | The proposed framework delivers speedups of 12.72x, 1.70x, and 1.18x over FlashAttn, ZigZagRing, and APB without significant performance loss. |
Copied to clipboard
| Challenge: | Few-shot domain adaptation and NOTA detection are two real-world challenges for few-shot relation classification models. |
| Approach: | They propose a task to investigate two aspects of few-shot relation classification models . they build upon the FewRel dataset by adding a new test set in a different domain . |
| Outcome: | The proposed task can evaluate few-shot domain adaptation and few- shot none-of-the-above detection on a new domain and NOTA relation choice. |
Copied to clipboard
| Challenge: | Existing approaches for non-factoid question answering can be categorized into representation and interaction focused approaches. |
| Approach: | They propose a novel approach which derives contextualized uni-gram representation from n-grams. |
| Outcome: | The proposed approach achieves state-of-the-art in two public non-factoid question answering datasets. |
Copied to clipboard
| Challenge: | Recent advances on prompting and post-training have enabled LLMs to perform step-wise reasoning tasks, but they tend to explore unproductive solution paths without effective backtracking or strategy adjustment. |
| Approach: | They propose a framework that empowers LLMs to “think about how to think” and dynamically adapts reasoning strategies in real-time. |
| Outcome: | The proposed framework outperforms previous SOTA methods by 9-12% in accuracy while reducing inference time by 28-35% under the same compute budget. |
Copied to clipboard
| Challenge: | Existing efficiency-oriented methods attempt to shorten or mix reasoning strategies, yet often degrade reasoning capability. |
| Approach: | They propose a token-level dual-process framework that explicitly decouples efficiency and correctness signals during training. |
| Outcome: | The proposed framework reduces inference cost while maintaining strong reasoning ability across multiple benchmarks. |
Copied to clipboard
| Challenge: | Digital media platforms often contribute to cognitive-behavioral fixation, a phenomenon in which users exhibit sustained and repetitive engagement with narrow content domains. |
| Approach: | They propose a multimodal topic extraction module and a cognitive-behavioral fixation quantification module that collaboratively enable adaptive, hierarchical, and interpretable assessment of user behavior. |
| Outcome: | The proposed framework lays the groundwork for scalable computational analysis of cognitive fixation. |
Copied to clipboard
| Challenge: | Aphasia is a language disorder caused by brain damage affecting speech functions . a detailed diagnosis of aphasia type is imperative for effective treatment . but, little attention has been paid to developing methods to detect different types of sphasis . |
| Approach: | They propose a multimodal graph neural network for aphasia type detection using co-speech gestures and corresponding speech and gesture patterns. |
| Outcome: | The proposed model outperforms existing methods in F1 and 84.2% of cases. |
Copied to clipboard
| Challenge: | Existing search-augmented approaches rely on indiscriminate whole-image retrieval and lack deep iterative reflection, limiting their effectiveness on complex visual queries. |
| Approach: | They propose a fully autonomous framework that shifts from passive perception to active visual planning and introduces a Selective Gaze mechanism that dynamically chooses whether to glance at global context or gaze into high-value regions. |
| Outcome: | Experiments across six benchmarks demonstrate state-of-the-art performance. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language and symbolic notations. |
| Approach: | They analyze the current LLM learning paradigms to tackle four critical evaluation dimensions that have emerged as critical dimensions in recent studies. |
| Outcome: | The proposed models are able to interact with chemical spaces through natural language and symbolic notations, and have emerging extensions to incorporate multi-modal inputs. |
Copied to clipboard
| Challenge: | Existing approaches to automatically generate commit messages are repetitive or redundant. |
| Approach: | They propose a retrieval-augmented neural commit message generation method which treats the retrieved similar commit as an exemplar and leverages it to generate an accurate commit message. |
| Outcome: | The proposed method outperforms baselines on a large dataset with five programming languages and can boost existing Seq2Seq models in commit message generation. |
Copied to clipboard
| Challenge: | Experimental results show that our approach can effectively improve the performance of both the policy model and the reward model. |
| Approach: | They propose to use Monte Carlo Tree Search for both policy model improvement and reward model improvement to bridge it to more subtle open-domain question answering. |
| Outcome: | The proposed approach surpasses existing methods for annotation and training data with fewer data points and achieves better performance in test-time scaling strategies. |
Copied to clipboard
| Challenge: | a new dataset is being developed to improve the capabilities of mobile GUI-control agents. |
| Approach: | They propose a dataset designed for generalist mobile GUI-control agents . they use screenshots from popular mobile applications to create a detailed GUI-annotated dataset . |
| Outcome: | The Android Multi-annotation EXpo (AMEX) is a large-scale dataset for generalist mobile GUI-control agents . it includes screenshots from popular mobile applications, which are annotated at multiple levels . |
Copied to clipboard
| Challenge: | Academic paper search often struggles to match underlying academic concepts between queries and documents. |
| Approach: | They propose a framework that extracts key concepts from papers and organizes them as a semantic index guided by an academic taxonomy. |
| Outcome: | The proposed framework can be flexibly employed to enhance existing retrieval frameworks. |
Copied to clipboard
| Challenge: | Existing efforts to generate static visualizations focus on static charts and interactive dashboards. |
| Approach: | They propose a dashboard2code task that requires a model to explore an interactive dashboard, acquire feedback from its own interactions and generate code that reproduces the target dashboard. |
| Outcome: | The proposed task is based on 180 carefully designed and manually verified dashboard–code pairs spanning three difficulty levels and covering eight common real-world interaction patterns. |
Copied to clipboard
| Challenge: | Existing text VQA systems generate an answer by selecting from optical character recognition (OCR) texts or a fixed vocabulary. |
| Approach: | They propose a localization-aware answer prediction network that generates the answer and predicts a bounding box as evidence of the generated answer. |
| Outcome: | The proposed network outperforms existing methods on three benchmark datasets for the text VQA task by a noticeable margin. |
Copied to clipboard
| Challenge: | Existing studies on cross-lingual entity alignment under adversarial attacks have not been conducted. |
| Approach: | They propose to use adversarial attack techniques to perturb cross-lingual entity alignment under adversarials. |
| Outcome: | The proposed model hides the attacked entities in dense regions in two KGs, and reduces the gradient vanishing issues in the process of adversarial attacks for further improving the attack effectiveness. |
Copied to clipboard
| Challenge: | State-of-the-art large language models (LLMs) are vulnerable to jailbreak attacks, such as GCG and AutoDAN. |
| Approach: | They propose to take the advances of online In-Context Learning and an offline defensive suffix and optimize them using an iterative algorithm and an online stochastic random search to identify the most effective ICL demonstrations. |
| Outcome: | The proposed method reduces attack success rate to nearly *0% while maintaining the model’s utility on benign tasks and incurring only *negligible* computational overhead. |
Copied to clipboard
| Challenge: | Existing work assumes main task labels and protected attributes are available in the dataset, but protected labels are often unavailable or only available in limited numbers. |
| Approach: | They propose a method which uses only a small volume of protected labels to train adversarial models using a dataset with a discriminator. |
| Outcome: | The proposed method can be used to transfer private-labelled instances from one dataset to another without requiring large amounts of protected labels. |
Copied to clipboard
| Challenge: | Existing methods to expand query use pseudo relevance feedback (PRF) but they are under-equipped to evaluate the relevance of information pieces used for expansion. |
| Approach: | They propose a query expansion model that leverages the BERT model to select relevant document chunks for expansion. |
| Outcome: | The proposed model significantly outperforms existing models on the TREC Robust04 and GOV2 test collections. |
Copied to clipboard
| Challenge: | Recent studies also use large language models (LLMs) for query understanding, but these methods lack grounding in corpus-specific knowledge and may generate unreliable or unfaithful content. |
| Approach: | They propose a paper retrieval framework that combines large language models (LLMs) with a concept-based semantic index to capture scientific concepts. |
| Outcome: | The proposed framework improves the performance of various base retrievers, surpasses strong existing LLM-based baselines, and remains highly efficient. |
Copied to clipboard
| Challenge: | Reinforcement learning with verifiable rewards (RLVR) has emerged as a paradigm for enhancing the reasoning capabilities of large language models. |
| Approach: | They propose a positive-advantage reweighting approach that regulates model entropy by adjusting the loss weights assigned to tokens with positive advantages during RLVR training. |
| Outcome: | The proposed approach regulates model entropy by adjusting loss weights assigned to tokens with positive advantages during RLVR training while maintaining competitive performance. |
Copied to clipboard
| Challenge: | Experimental results show that by applying our framework, we can easily learn effective FGET models for low-resource languages. |
| Approach: | They propose a cross-lingual contrastive learning framework to learn FGET models for low-resource languages. |
| Outcome: | The proposed framework can learn effective FGET models for low-resource languages even without human-labeled data. |
Copied to clipboard
| Challenge: | Existing studies on storytelling with sound have focused on visuals and sounds, but little attention has been given to sound. |
| Approach: | They propose to establish a new component called background sound which is story context-based audio without any linguistic information. |
| Outcome: | The proposed dataset is the largest well-curated dataset for storytelling with sound . it contains 27,354 stories with 19.6 images per story and 984 hours of speech-decoupled audio . |
Copied to clipboard
| Challenge: | Existing OpenRE methods cast different relation types in isolation without considering their hierarchical dependency. |
| Approach: | They propose a framework to establish bidirectional connections between OpenRE and relation hierarchies by integrating hierarchy information into relation representations. |
| Outcome: | The proposed framework outperforms state-of-the-art models on relation clustering and hierarchy expansion. |
Copied to clipboard
| Challenge: | Existing benchmarks rely on human annotations that are vulnerable to value-related biases. |
| Approach: | They propose a value portrait benchmark that uses items that capture real-life user-LLM interactions and a rated item based on its similarity to their own thoughts to determine reliability. |
| Outcome: | The proposed framework improves the relevance of assessment results to real-world LLM usage by allowing human subjects to rate items with similarity to their own thoughts and derived correlations between these ratings and the subjects’ actual value scores. |
Copied to clipboard
| Challenge: | Existing embedding approaches for temporal knowledge graphs typically learn entity representations and their dynamic evolution in the Euclidean space. |
| Approach: | They propose a non-Euclidean embedding approach that learns evolving entity representations in a product of Riemannian manifolds. |
| Outcome: | The proposed model improves on three real-world datasets showing that the embeddings on Riemannian manifolds can capture the evolution of temporal KGs. |
Copied to clipboard
| Challenge: | Existing detectors rely on stylistic cues to distinguish between surface-level language refinement and genuine content generation. |
| Approach: | They propose a content-based detection paradigm to detect substantive AI-generation . they propose 'CoCoDet' detector that can detect surface-level language refinement . |
| Outcome: | The proposed detector achieves a macro F1 score of 98.24% on permissible machine-polished reviews and maintains 3.89% false positive rate on real-world reviews. |
Copied to clipboard
| Challenge: | Existing black-box large language models (LLMs) have excellent performance in task-oriented dialogue (TOD) tasks, but obtaining suitable prompts for specific tasks is challenging. |
| Approach: | They propose a black-box large language model that generates domain and slot information in the belief state, which serves as prior knowledge for subsequent prompt generation. |
| Outcome: | The proposed framework outperforms existing prompting methods on the MultiWOZ 2.0 dataset. |
Copied to clipboard
| Challenge: | Existing context condensing methods cannot accurately understand the full context, as there is a considerable amount of information loss in the condensed process. |
| Approach: | They propose a framework to extend the fixed context length of any decoder-only LLM by distilling crucial information from long sequences. |
| Outcome: | The proposed framework extends the fixed context length of any decoder-only LLM, allowing it to focus on relevant information from very long sequences. |
Copied to clipboard
| Challenge: | Existing methods for benchmarking the uncertainty of large language models face challenges . existing methods require internal model access, additional training, or high computational costs . |
| Approach: | They propose a new benchmark for evaluating the uncertainty of large language models based on confidence intervals . UBench encompasses 11,978 multiple choice questions spanning knowledge, language, understanding, and reasoning capabilities. |
| Outcome: | The proposed method outperforms existing methods for benchmarking the uncertainty of large language models. |
Copied to clipboard
| Challenge: | Existing reasoning methods for sparse KGs are incomplete and lack of evidential paths to target entities makes multi-hop reasoning difficult. |
| Approach: | They propose a multi-hop reasoning model over sparse KGs to solve this problem . they use latent prediction of embedding-based models to make the model perform more potential path search over sparses . |
| Outcome: | The proposed method outperforms state-of-the-art models on five datasets from Freebase, NELL and Wikidata. |
Copied to clipboard
| Challenge: | Claims are often nuanced and cannot be clearly labeled as “true” or “false” . however, a claim can be dissected into integral aspects and sub-aspects that are individually easier to validate . |
| Approach: | They propose a retrieval-augmented generation-based framework for deconstructing nuanced claims . claim can be dissected into integral aspects and sub-aspects, which are easier to validate . |
| Outcome: | The proposed framework can be easily deconstructed into integral aspects and sub-aspects, which are easier to validate. |
Copied to clipboard
| Challenge: | Recent studies have discovered notable disparities in their performance across different languages. |
| Approach: | They conduct a systematic investigation into the behaviors of large language models across 27 different languages on 3 different scenarios and reveals a Linguistic Map correlates with the richness of available resources and linguistic family relations. |
| Outcome: | The proposed model demonstrates that there are significant disparities in performance across languages across 27 different languages on 3 different scenarios. |
Copied to clipboard
| Challenge: | Current large language models often give away solutions directly, making them ineffective instructors. |
| Approach: | They propose to use a state space-based planning algorithm to build a question tree based on a student's knowledge state to help students independently identify and resolve errors. |
| Outcome: | The proposed model is able to debug code efficiently with minimal turns and highly Socratic questioning. |
Copied to clipboard
| Challenge: | Multilingual pretrained language models (mPLMs) have shown their effectiveness in multilingual word alignment induction, but these methods usually start from mBERT or XLM-R. |
| Approach: | They propose to fine tune multilingual sentence Transformer LaBSE for alignment induction using parallel corpus and a parallel corpora model. |
| Outcome: | The proposed model outperforms existing models on seven language pairs and achieves new state-of-the-art on zero-shot language pairs. |
Copied to clipboard
| Challenge: | Recent large-scale Visual-Language Generative Models (VLGMs) generate toxic content, e.g., offensive text and pornography images, raising significant ethical risks. |
| Approach: | They propose a bottleneck-based detoxification method to reduce toxicity while maintaining comparable generation quality. |
| Outcome: | The proposed method could reduce toxicity while maintaining comparable generation quality. |
Copied to clipboard
| Challenge: | Probabilistic finite automata (PFAs) are statistical language models used in natural language processing. |
| Approach: | They develop an asymptotic algorithm to compute the infix probabilities of each prefix of a string from streaming data. |
| Outcome: | The proposed algorithm improves the infix probabilities of a weighted automata from streaming data. |
Copied to clipboard
| Challenge: | Speculative sampling is an efficient way to accelerate the auto-regressive generation process of large language models. |
| Approach: | They propose a frequency-ranked speculative sampling framework that optimizes draft candidate selection through vocabulary space compression. |
| Outcome: | Experiments show that FR-Spec reduces LM Head computation overhead by 75% while ensuring the equivalence of the final output distribution. |
Copied to clipboard
| Challenge: | Existing methods for aligning large language models rely on preference-based approaches that require both positive and negative feedback as a pair. |
| Approach: | They propose a binary classifier optimization technique that trains a classifier using only binary feedback and a reward shift technique which minimizes the DPO loss. |
| Outcome: | The proposed method performs on a paired preference dataset and on 'likert-5 scale annotation dataset' it consistently demonstrates effective and robust alignment across four base LLMs and three different datasets, showcasing the strength of the proposed technique. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on scalar rewards to capture user preferences. |
| Approach: | They propose a framework that integrates multi-objective reward modeling to represent diverse preference profiles. |
| Outcome: | The proposed method improves performance across reward objectives and targets. |
Copied to clipboard
| Challenge: | a limited number of human annotations are required to evaluate multilingual summarization evaluation metrics. |
| Approach: | They propose a multilingual meta-evaluation framework that uses machine translation systems to transform a monolingual metaevaluations dataset into multilingual versions. |
| Outcome: | The proposed framework outperforms classical text-matching-based metrics in non-English languages. |
Copied to clipboard
| Challenge: | Existing syntactically-controlled paraphrase generation models perform well with human-annotated or well-chosen syntaktic templates. |
| Approach: | They propose a quality-based Syntactic Template Retriever to retrieve templates based on the quality of the to-be-generated paraphrases. |
| Outcome: | The proposed algorithm can generate high-quality paraphrases without sacrificing quality. |
Copied to clipboard
| Challenge: | Existing datasets only cover limited relation types at once, which prevents models from taking full advantage of relation interactions. |
| Approach: | They construct a large-scale human-annotated ERE dataset with improved annotation schemes to address these drawbacks. |
| Outcome: | The proposed dataset is larger than existing datasets of all the ERE tasks by at least an order of magnitude. |
Copied to clipboard
| Challenge: | Document Structured Extraction (DSE) is a field of document structure analysis that aims to extract structured content from raw documents. |
| Approach: | They propose a benchmark to evaluate document structured extraction systems by converting unstructured PDFs into semantically rich Markdown. |
| Outcome: | The proposed benchmark is based on 3,576 diverse and real-world documents from arXiv, GitHub, and Zenodo. |
Copied to clipboard
| Challenge: | Experimental results show that events can greatly improve the quality of KG embeddings on multiple downstream tasks. |
| Approach: | They propose an event-enhanced KG embedding model that incorporates events into KGs . they first incorporate event nodes by building a heterogeneous network with event argument links . |
| Outcome: | The proposed model incorporates event nodes into the original knowledge graphs . it can be used to fuse event information into the KG embeddings on multiple tasks . |
Copied to clipboard
| Challenge: | Recent studies have discussed its capability to assist language models for various applications. |
| Approach: | They propose a structure to organize arguments using the **Hi**erarchical **Ar**gumentation **G**raph (Hi-ArG) and propose two approaches to exploit Hi-AarG, including a text-graph multi-modal model GreaseArR and a framework augmented with graph information. |
| Outcome: | The proposed structure supersedes existing language models on two argumentation tasks while incorporating graph information during further training improves vanilla language models. |
Copied to clipboard
| Challenge: | Efficient reproduction of research papers requires deep domain expertise. |
| Approach: | They propose a framework that systematically mines implicit knowledge from the cited literature to reproduce experimental code in a complete, end-to-end manner. |
| Outcome: | The proposed framework surpasses baselines across all metrics and reproduces experimental code in a complete, end-to-end manner. |
Copied to clipboard
| Challenge: | Existing paradigm to fine-tune parameters of pre-trained language models poses problems in data-scarce and resource-limited scenarios. |
| Approach: | They propose a parameter-efficient fine-tuning method HiFi that fine-tails only the highly informative and strongly correlated attention heads for the specific task. |
| Outcome: | The proposed method obtains state-of-the-art over the prior benchmarks on the GLUE benchmark. |
Copied to clipboard
| Challenge: | Existing efforts to create benchmarks that move beyond superficial pattern recognition to delve into the profound reasoning skills required for problemsolving face challenges such as insufficient interpretability, performance saturation or data contamination. |
| Approach: | They propose a gaming arena designed for rigorous assessment of LLM reasoning capabilities. |
| Outcome: | The proposed framework decomposes complex reasoning into predefined modular subproblems and generates ground truth for these subproblem types. |
Copied to clipboard
| Challenge: | Current TIMT studies focus on providing translations for all text within an image, neglecting to provide bounding boxes and covering limited scenarios. |
| Approach: | They extend traditional TIMT into position-aware TIMt to support fine-grained translation . they introduce an Adaptive Image OCR Refinement Pipeline to refine results . |
| Outcome: | The proposed model supports fine-grained and layout-preserving translation . the experimental data highlight the scalability and generalizability of the model. |
Copied to clipboard
| Challenge: | a growing number of LLMs have been used to provide reasoning, writing, text-editing capabilities. |
| Approach: | They propose a method to inject imperceptible phantom tokens into LLMs to deceive users . the technique generates outputs that appear plausible to users but are in fact incorrect . |
| Outcome: | a new method injects imperceptible tokens into documents to deceive users . the proposed framework is compared to baselines to show its effectiveness . |
Copied to clipboard
| Challenge: | Existing approaches to stream learning NLP tasks suffer from catastrophic forgetting and are exacerbated when the previous task’s pseudo data is insufficient. |
| Approach: | They propose to use a new data format to train pseudo questions of previous tasks to stream learning NLP tasks while retaining knowledge of previous ones. |
| Outcome: | The proposed model is more robust to sufficient and insufficient pseudo-data when the task boundary is both clear and unclear. |
Copied to clipboard
| Challenge: | Experimental results show that PREMISE achieves promising performance with less computational cost. |
| Approach: | They propose a new architecture for matching-based learning in multimodal fields for the MRHP task. |
| Outcome: | The proposed architecture significantly boosts performance on multimodal tasks with less computational cost compared to the state-of-the-art fusion-based methods. |
Copied to clipboard
| Challenge: | Impact assessment is an evolving area of research that aims at measuring and predicting the potential effects of projects or programs. |
| Approach: | They propose a framework for automatically assessing the impact of scientific research by identifying pertinent sections in project reports that indicate potential impacts. |
| Outcome: | The proposed method achieves accuracy scores up to 0.81 and is generalizable to scientific research from different domains and languages. |
Copied to clipboard
| Challenge: | a dataset evaluating harmful capabilities in large language models is available at https://github.com/Libr-AI/do-not-answer. |
| Approach: | They collect an open-source dataset to evaluate the safeguards in large language models . they find that simple BERT-style classifiers can achieve results comparable to GPT-4 . |
| Outcome: | The proposed dataset compares the safety of six popular LLMs to GPT-4 on automatic safety evaluation. |
Copied to clipboard
| Challenge: | vocab expansion scaling laws are well-established for high-resource languages, but they remain unverified in low-resourced settings. |
| Approach: | They propose to scale trilingual vocabulary for languages with 140 to 195,000 tokens . they find that BBPE follows a "decline-then-rise" pattern, whereas BPE improves monotonically . |
| Outcome: | The proposed configuration reduces pre-training duration by over 71% across 1.5B to 8B models while improving downstream performance. |
Copied to clipboard
| Challenge: | Diffusion large language models (dLLMs) offer bidirectional attention and parallel generation . fixed anchors can enforce constraints, but they often impose rigid spans, leading to truncated reasoning . |
| Approach: | They propose a method that dynamically estimates end-anchor positions to adjust generation length before iterative infilling. |
| Outcome: | The proposed method improves format compliance and answer accuracy on GSM8K and MATH. |
Copied to clipboard
| Challenge: | a new study analyzes the political slants of user comments on partisan media in Korea . the classifiers detect political leaning on conservative and liberal news outlets . |
| Approach: | They built a BERT-based classifier to detect political leaning of short comments . they found a high presence of conservative bias on conservative and liberal news outlets . |
| Outcome: | The proposed classifier produced an F1 score of 0.83 for 21.6K comments . it shows that more liberals comment on stories resonating with their political perspectives . |
Copied to clipboard
| Challenge: | Existing IE tools for atomic events are limited when applied to such complex events. |
| Approach: | They propose to use event schemas to guide the organization of complex events and to edit hierarchical graphs. |
| Outcome: | The proposed tool outperforms existing IE visualization tools in both IE result analysis and general model improvements. |
Copied to clipboard
| Challenge: | Current OpenRE models are often trained on the datasets generated from distant supervision, which often results in instability and makes the model easily collapsed. |
| Approach: | They propose to use a causal model to identify relation instances referring to the same relation . they propose to perform Element Interventions on context and entities respectively . |
| Outcome: | The proposed method outperforms existing methods and is robust across datasets. |
Copied to clipboard
| Challenge: | Large language models have demonstrated impressive reasoning capabilities across multiple languages, but the relationship between capabilities in different languages is less explored. |
| Approach: | They decompose the process of reasoning tasks into two separate components: knowledge retrieval and knowledge-free reasoning. |
| Outcome: | The proposed model can be transferred across source-target languages despite secondary impact of resource in some specific target languages, while cross-lingual knowledge retrieval significantly hinders the transfer. |
Copied to clipboard
| Challenge: | Existing methods that ignore contextual knowledge fail to reliably fall back to parametric knowledge when presented with irrelevant context. |
| Approach: | They propose to use contextual knowledge to update and correct LLMs' knowledge by in-context editing instead of retraining. |
| Outcome: | The proposed method outperforms current state-of-the-art methods by a large margin on a dataset that contains irrelevant questions. |
Copied to clipboard
| Challenge: | Recent research shows that LLM Agents can generate “believable” human behaviors via prompt-only methods, leaving open questions of whether they can accurately generate step-by-step actions in multi-turn interaction tasks. |
| Approach: | They propose to use shopping data to evaluate LLMs' ability to accurately generate step-by-step actions in a multi-turn interaction task. |
| Outcome: | The proposed model achieves 17.26% action generation accuracy and 33.86% F1 score on final purchase prediction, representing improvements of 5.4% and 13.85% over baselines. |
Copied to clipboard
| Challenge: | Existing models for visually rich document understanding do not account for the diverse carriers of document versions and their associated noises. |
| Approach: | They propose a multimodal, multi-task, multiteacher joint-grained knowledge distillation model for visually-rich form document understanding. |
| Outcome: | The proposed model outperforms baselines on a comprehensive evaluation of public datasets showing it can handle complex structures and content of visually-rich forms. |
Copied to clipboard
| Challenge: | Existing auto-regressive pre-trained language models are challenged by recent emerging numerical reasoning datasets due to the error-prone implicit calculation. |
| Approach: | They propose a pre-computation tool to pre-compute aggregation/arithmetic results for the table in advance, so they are handy and readily available for PLMs to answer numerical reasoning questions. |
| Outcome: | The proposed model improves on TAT-QA and T5 and BART-large on multiple benchmarks. |
Copied to clipboard
| Challenge: | Existing evaluation methods for mobile GUI agents rely on static frame assessments or offline static apps. |
| Approach: | They propose an evaluation system that leverages large language models as reward models to verify task completion and process achievement. |
| Outcome: | The proposed system addresses the limitations of traditional function based evaluation methods on online dynamic apps. |
Copied to clipboard
| Challenge: | Sequence-to-Sequence (S2S) models have been successful on text generation tasks . however, learning complex structures with S2S models remains challenging . |
| Approach: | They propose to use constrained decoding to model part-of-speech tagging, named entity recognition, constituency, and dependency parsing tasks with 3 lexically diverse linearization schemas and corresponding constrained coding methods. |
| Outcome: | The proposed methods outperform the state-of-the-art on four core tasks. |
Copied to clipboard
| Challenge: | Factuality evaluation aims to detect factual errors produced by language models and guide the development of more factual models. |
| Approach: | They propose a framework that leverages FenCE to improve the factuality of LM generators by constructing training data. |
| Outcome: | The proposed framework improves the factuality of LM generators by enhancing their training data. |
Copied to clipboard
| Challenge: | toxicity detection in French remains underdeveloped due to the lack of culturally relevant, human-annotated, large-scale datasets. |
| Approach: | They propose a method that generalizes French online comments using a semi-automated annotation pipeline that reduces manual labeling to only 10% through high-confidence LLM-based pre-annotation and human verification. |
| Outcome: | The proposed model outperforms GPT-4o and DeepSeek-R1 on the benchmark while maintaining cross-lingual capabilities. |
Copied to clipboard
| Challenge: | Existing methods to extract event records from text decompose complex structure prediction task into multiple subtasks. |
| Approach: | They propose a sequence-to-structure generation paradigm that can extract events from text . they propose unified event extraction, constrained decoding algorithm and curriculum learning algorithm . |
| Outcome: | The proposed method can achieve competitive performance using record-level annotations in both supervised learning and transfer learning settings. |
Copied to clipboard
| Challenge: | Existing research on machine translation tools has not revealed how users perceive MT errors and how they evolve through interaction. |
| Approach: | They propose a framework where users accept MT output or request professional re-translation to answer questions based on information presented in a foreign language. |
| Outcome: | The proposed framework can predict where the system is likely to be wrong and how it evolves through interaction. |
Copied to clipboard
| Challenge: | Task-oriented dialog (TOD) is one of the central objectives, hallmarks, and applications of machine intelligence. |
| Approach: | They propose a multilingual, multi-domain, multiparallele ToD dataset that offers culturally adapted dialogs in 4 languages for training and evaluation of multilingual and cross-lingual systems. |
| Outcome: | The proposed dataset is large-scale and culturally adapted to enable training and evaluation of multilingual and cross-lingual ToD systems. |
Copied to clipboard
| Challenge: | TrickCatcher generates test cases that pass existing tests yet contain bugs . a recent study found that tricky bugs are not detected by test suites . |
| Approach: | They propose an LLM-powered approach to generating test cases for uncovering bugs in plausible programs . they use a PUT and specification to generate program variants, an input generator and an Llm to construct test inputs . |
| Outcome: | The proposed approach achieves recall, precision, and F1 scores that are 1.80, 2.65, and 1.66 . trickCatcher generates program variants based on the program under test and its specification . |
Copied to clipboard
| Challenge: | State-of-the-art automatic event detection struggles with interpretability and adaptability to evolving large-scale key events. |
| Approach: | They propose a task which identifies episodes within a news corpus of key event articles. |
| Outcome: | The proposed framework achieves 59.2% gain across all metrics compared to baselines. |
Copied to clipboard
| Challenge: | Existing backdoor attacks on Multimodal Large Language Models are less applicable to open-ended conversations with users. |
| Approach: | They propose a shadow-activated backdoor attack scenario where attackers inject malicious content into the responses of MLLMs when the responses explicitly relate to the shadowed object. |
| Outcome: | The proposed framework achieves the desired behaviors by constructing a poisoned dataset and implementing an attention-regularized tuning strategy. |
Copied to clipboard
| Challenge: | Existing methods to extract relational facts from open domain corpora are time-consuming and human-intensive. |
| Approach: | They propose a framework to learn similarity metrics of relations from labeled data . they propose to transfer relational knowledge to identify novel relations in unlabeled data. |
| Outcome: | Experiments on two real-world datasets show that the proposed framework improves compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing NL2SQL systems rely on in-context learning with only correct examples . current test-time scaling methods often decompose questions arbitrarily, resulting in poor performance . |
| Approach: | They propose a structured decomposition and experience-aware self-correction framework for NL2SQL . they build a dynamic memory of successful queries and historical error–fix pairs . |
| Outcome: | The proposed framework achieves 68.5% execution accuracy on BIRD, setting new state of the art among open, zero-fine-tuning methods. |
Copied to clipboard
| Challenge: | Existing knowledge graphs lack the ability to integrate structural information into LLMs and output predictions deterministically. |
| Approach: | They propose a method which encodes structural information of KGs and merges it with LLMs to enhance KGC performance. |
| Outcome: | The proposed method improves the performance of KG Completion datasets on KGs by integrating structural information with LLMs. |
Copied to clipboard
| Challenge: | Existing approaches to contrastive learning are heavily affected by superficial features like sentence length and syntax. |
| Approach: | They propose a semantic-aware contrastive learning framework for sentence embeddings that explores the pseudo-token space representation of a sentence while eliminating the impact of superficial features such as sentence length and syntax. |
| Outcome: | The proposed framework outperforms the state-of-the-art on six standard semantic textual similarity tasks while maintaining an additional queue to store the representation of sentence embeddings. |
Copied to clipboard
| Challenge: | Existing systems that generate *flashbacks* are monotonic and lack explicit guidance on how to insert them. |
| Approach: | They propose to use event temporal orders to encode events as temporal prompts . they leverage a Plan-and-Write framework enhanced by reinforcement learning to generate storylines . |
| Outcome: | The proposed method generates more interesting stories with *flashbacks* while maintaining textual diversity, fluency, and temporal coherence. |
Copied to clipboard
| Challenge: | Existing methods for generating SQL queries using natural language questions produce inconsistent NLQ-SQL pairs. |
| Approach: | They propose a text-to-SQL data synthesis framework that generates domain-relevant questions . they synthesize NLQ-SqL pairs that are domain-specific and intent-consistent . |
| Outcome: | The proposed method outperforms closed-source LLMs on the Text-to-SQL task. |
Copied to clipboard
| Challenge: | Mixture of Experts (MoE) models use homogeneous experts with diverse capacities, resulting in a lack of expert specialization and parameter utilization. |
| Approach: | They propose a framework where experts differ in size and possess diverse capacities . they propose HMoE to encourage frequent activation of smaller experts . |
| Outcome: | The proposed framework outperforms homogeneous homogenous MoE models on evaluation benchmarks and achieves lower loss rate with fewer activated parameters. |
Copied to clipboard
| Challenge: | philology requires years of professional training in extensive knowledge memorization and manual textual retrieval. |
| Approach: | They curated the PhiloCorpus-ZH, a rich collec-tion of ancient Chinese texts spanning a millennium with 30 diverse topics, including firsthand folk copies. |
| Outcome: | The PhiloCorpus-ZH corpus facilitated the development of the first LLM tailored for discovering ancient Chinese manuscripts. |
Copied to clipboard
| Challenge: | Existing methods for continual learning (CL) are designed to mitigate catastrophic forgetting while neglecting knowledge sharing across tasks. |
| Approach: | They propose a framework that facilitates knowledge transfer while mitigating catastrophic forgetting by assigning task-specific parameter subspaces to new tasks . they then leverage attribution scores to evaluate task similarity and employ soft orthogonality between task- specific subspace . |
| Outcome: | The proposed framework facilitates knowledge transfer while mitigating catastrophic forgetting. |
Copied to clipboard
| Challenge: | Existing methods for quantizing large language models focus on breaking down the problem into layer-wise sub-problems and minimizing per-layer error, but this approach lacks theoretical justification and the metrics employed may be sub-optimal. |
| Approach: | They propose a "linearity theorem" establishing a direct relationship between the layer-wise reconstruction error and the model perplexity increase due to quantization. |
| Outcome: | The proposed method outperforms previous data-free methods and improves accuracy-compression trade-offs on Llama-family models. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) is an attractive solution for task-oriented dialog systems . but extending RL-based systems to handle new intents and slots requires a system redesign . |
| Approach: | They propose a teacher-student framework to extend RL-based dialog systems . they propose to specify constraints held in the new dialog manager . |
| Outcome: | The proposed framework makes no assumption about unsupported intents and slots, making it possible to improve RL-based systems incrementally. |
Copied to clipboard
| Challenge: | Existing long-context training data is scarce and requires substantial GPU resources for training. |
| Approach: | They propose a training-free plug-and-play method to enhance long-context understanding in existing large language models. |
| Outcome: | The proposed method outperforms existing LLMs on various tasks and surpasses baseline methods. |
Copied to clipboard
| Challenge: | Out-of-distribution (OOD) detection is essential for multimodal learning systems . a novel scoring framework is proposed to efficiently detect OOD in multi-round long dialogues . |
| Approach: | They propose a scoring framework that integrates visual language models with a score framework that detects OOD in two key scenarios. |
| Outcome: | The proposed framework detects OOD in two key scenarios: mismatches between dialogue and image input pair and previously unseen labels. |
Copied to clipboard
| Challenge: | Top-view perspective is a typical way in which humans read and reason over different types of maps, but spatial reasoning capabilities of modern VLMs in this setup remain unattested and underexplored. |
| Approach: | They introduce a top-view spatial reasoning dataset and use it to evaluate VLMs across 4 perception and reasoning tasks with different levels of complexity. |
| Outcome: | The proposed model can understand and reason over spatial relations from the top view and can be controlled at different granularities of spatial reasoning. |
Copied to clipboard
| Challenge: | Existing approaches struggle with temporal-spatial challenges in capturing subtle linguistic shifts across different disease stages. |
| Approach: | They propose a large language model-driven T-S fusion framework that integrates multilingual LLMs, contrastive learning and interpretable marker discovery to revolutionize late onset AD detection. |
| Outcome: | The proposed framework achieves state-of-the-art performance in late onset AD detection while enabling cross-linguistic diagnostics. |
Copied to clipboard
| Challenge: | Xu et al., 2021: conversational semantic role labeling is under-explored in non-Chinese languages due to the lack of multilingual CSRL annotations for the parser training. |
| Approach: | They propose a model that implicitly learns conversational structure-aware representations with hierarchical encoders and elaborately designed pre-training objectives. |
| Outcome: | The proposed model outperforms baselines on English CSRL tests by large margins . it will facilitate the research of non-Chinese dialogue tasks which suffer from ellipsis and anaphora . |
Copied to clipboard
| Challenge: | Existing methods to enhance reasoning capabilities of large language models incur significant overhead in token usage, leading to increased costs. |
| Approach: | They propose a token-budget-aware LLM reasoning framework that adjusts the number of reasoning tokens based on the reasoning complexity of each problem. |
| Outcome: | The proposed method reduces token costs in CoT reasoning with only a slight performance reduction. |
Copied to clipboard
| Challenge: | Existing studies have shown that cross-lingual knowledge distillation can improve the performance of pre-trained models for cross-linguistic similarity matching tasks. |
| Approach: | They propose a multi-stage distillation framework for constructing a small-size but high-performance cross-lingual model using contrastive learning, bottleneck, and parameter recurrent strategies. |
| Outcome: | The proposed model can compress the size of XLM-R and MiniLM by more than 50% while the performance is only reduced by about 1%. |
Copied to clipboard
| Challenge: | Vision-Language-Action models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic variation remains poorly understood. |
| Approach: | They propose a step-wise inference-time intervention that aligns representations according to step language sensitivity, significantly improving performance under linguistic variation. |
| Outcome: | The proposed model significantly improves performance under linguistic variation under non-English instructions under language-agnostic steps. |
Copied to clipboard
| Challenge: | Existing LLMs are opaque and difficult to interpret, resulting in limited interpretability. |
| Approach: | They propose an interaction-aware profile generator that jointly produces user and item profiles conditioned on both user history and item evidence. |
| Outcome: | The proposed model outperforms baselines on three real-world datasets. |
Copied to clipboard
| Challenge: | Reinforcement learning from human feedback (RLHF) is the primary method for aligning large language models with human preferences. |
| Approach: | They propose to train an Absolute-Rating Multi-Objective Reward Model with multi-dimensional absolute-rating data. |
| Outcome: | The proposed model outperforms the LLM-as-a-judge method on RewardBench . it achieves state-of-the-art performance on the benchmark . |
Copied to clipboard
| Challenge: | Recent studies on single-document summarization (SDS) benefit from advances in neural sequence learning, but they produce unsatisfactory results on multi-document summary (MDS). |
| Approach: | They propose a neural sequence learning method that unifies advanced neural SDS methods and statistical measures used in classical MDS. |
| Outcome: | The proposed method achieves state-of-the-art performance on benchmark MDS datasets. |
Copied to clipboard
| Challenge: | Existing methods for text classification use human annotations or a set of class seed words for supervision, which can be costly, especially in emerging domains. |
| Approach: | They propose a weakly-supervised method that leverages mutually-enhancing text granularities to learn a contextualized document representation that captures the most discriminative class indicators. |
| Outcome: | Extensive experiments on seven benchmark datasets show that MEGClass outperforms other weakly and extremely weakly supervised methods. |
Copied to clipboard
| Challenge: | Existing benchmarks explore aspects of threedimensional spatial reasoning and visual-language reasoning in dynamic environments, but they are unable to perform well on 3D spatial deformation reasoning. |
| Approach: | They propose to use a ladder competition format to assess the model's spatial deformation reasoning abilities to determine its performance. |
| Outcome: | The proposed framework assesses the performance of Vision-Language Models in spatial deformation reasoning tasks. |
Copied to clipboard
| Challenge: | Recent advances in automatic quality estimation for machine translation focus on written language, leaving the speech modality underexplored. |
| Approach: | They propose a new quality estimation system based on cascaded and end-to-end architectures. |
| Outcome: | The proposed system is better suited to estimating the quality of direct speech translation than existing systems designed for text translation. |
Copied to clipboard
| Challenge: | Existing studies on how SAEs derive most fine-grained latent features for safety remain unexplored. |
| Approach: | They propose a framework for interpreting SAE features in safety-critical domains . they train a suite of SAEs with human-readable explanations and systematic evaluations based on pornography, politics, violence, and terror . |
| Outcome: | The proposed framework reduces interpretation cost by 55% and improves safety-critical features. |
Copied to clipboard
| Challenge: | Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions. |
| Approach: | They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks. |
| Outcome: | The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces. |
Copied to clipboard
| Challenge: | Recent studies have shown that alignment of large language models with human values and preferences requires substantial data and computation resources. |
| Approach: | They propose a method to extract and isolate superficial knowledge from aligned models by focusing on the shallow modifications to the final token selection process. |
| Outcome: | The proposed method extracts and isolates superficial knowledge from aligned models, focusing on the shallow modifications to the final token selection process. |
Copied to clipboard
| Challenge: | Existing methods for XMC struggle with the growing set of labels due to their static label assumptions, and embedding-based methods struggle with complex mapping relationships due to late interaction paradigm. |
| Approach: | They propose a large language model (LLM) powered agent framework for extreme multi-label classification, XMC-Agent, which can effectively learn, manage and predict the extremely large and dynamically increasing set of labels. |
| Outcome: | The proposed framework can learn, manage and predict the extremely large and dynamically growing set of labels and achieves state-of-the-art performance on three standard datasets. |
Copied to clipboard
| Challenge: | Experimental results show that our approach significantly outperforms the supervised counterparts, and can even achieve competitive performance to supervised state-of-the-art (SoA) model. |
| Approach: | They propose a syntactic and semantic-driven learning approach that can learn open IE models without human-labelled data by leveraging syntakic and semantic knowledge as noisier, higher-level supervision. |
| Outcome: | The proposed approach outperforms supervised counterparts and can achieve competitive performance to supervised state-of-the-art models. |
Copied to clipboard
| Challenge: | Structured pruning can reduce model size but results in significant accuracy degradation . quantization and pruning increase the difficulty of fine-tuning, requiring a more refined quantization scheme. |
| Approach: | They propose a structured pruning framework followed by a layer-wise mixed-precision quantization scheme to reduce model memory consumption during fine-tuning and inference. |
| Outcome: | Experiments on benchmark datasets show that QPruner outperforms existing methods in memory savings while maintaining or improving model performance. |
Copied to clipboard
| Challenge: | Visual storytelling is the task of generating a story paragraph that describes a given image sequence. |
| Approach: | They propose 3 evaluation metrics sets that analyze which aspects we would look for in a good story . they compare their correlation with human judgement scores on a sample of machine stories . |
| Outcome: | The proposed evaluation metrics outperform other metrics on human correlation on a sample of machine stories from state-of-the-art models. |
Copied to clipboard
| Challenge: | Recent advances in large language models have push NLP into a new era, moving away from traditional task-specific pre-train finetuning paradigm. |
| Approach: | They provide a comprehensive analysis of declarative and procedural knowledge for large language models and evaluate their effectiveness. |
| Outcome: | The proposed model can perform better with both kinds of knowledge, but at different speeds. |
Copied to clipboard
| Challenge: | Autoregressive large language models suffer from high inference latency due to memorybandwidth constraints. |
| Approach: | They propose a method that decouples generation and verification by decoupling tokens and a lightweight draft model. |
| Outcome: | The proposed method delivers consistent and significant speedups over state-of-the-art baselines while preserving generation quality across diverse benchmarks. |
Copied to clipboard
| Challenge: | Recent studies show that LLMs’ intrinsic self-correction fails without oracle labels as feedback. |
| Approach: | They propose to use one simple task and three complex tasks with state-of-the-art LLMs like ChatGPT, Llama, and DeepSeek to interpret LLM's intrinsic self-correction. |
| Outcome: | The proposed methods reveal the dark side of LLMs’ intrinsic self-correction for different tasks, especially for those failure cases. |
Copied to clipboard
| Challenge: | Existing video benchmarks often resemble image-based questions with scans of only a few key frames, without deep temporal reasoning. |
| Approach: | They propose a video benchmark to assess whether large vision-language models can genuinely think with videos rather than perform superficial frame-level analysis. |
| Outcome: | The proposed benchmark consists of 3,269 videos and over 4,342 highly visual-centric questions across 11 categories, including Trajectory Analysis, Temporal Reasoning, and Forensics Detection. |
Copied to clipboard
| Challenge: | Existing methods for retrieving medical textual knowledge Graphs struggle to perform well, a study finds . existing methods struggle to provide accurate answers to complex questions, he says . |
| Approach: | They synthesize user queries integrating diverse topological structures, relational information, and complex textual descriptions. |
| Outcome: | a new dataset for medical textual knowledge graphs shows that existing methods struggle to perform well . main bottlenecks lie in the scarcity of existing medical TKGs and the limited expressiveness of their topological structures . |
Copied to clipboard
| Challenge: | Recent work on code search proposes data augmentation of queries for contrastive learning. |
| Approach: | They propose to augment query-code pairs with key words to preserve key words . they use keyDAC to fine-tune various pre-trained language models . |
| Outcome: | The proposed approach outperforms the current state-of-the-art in code search and question answering tasks. |
Copied to clipboard
| Challenge: | Existing methods for labeling relational facts require significant expert labor to write relation-specific patterns, which makes them too sophisticated to generalize quickly. |
| Approach: | They propose a neural pattern diagnosis framework that can summarize and refine relation-specific patterns with human experts in the loop. |
| Outcome: | The proposed framework can summarize and refine high-quality relational patterns from noise data with human experts in the loop. |
Copied to clipboard
| Challenge: | FTibSuite provides an end-to-end training-and-evaluation workflow for vision–language models . Tibetan is underserved due to the lack of infrastructure for reproducible training and evaluation. |
| Approach: | They propose a resource-centric workflow for Tibetan VLMs that provides an end-to-end training-and-evaluation workflow and human-verified multimodal annotations. |
| Outcome: | FTibSuite provides an end-to-end training-and-evaluation workflow and human-verified multimodal annotations. |
Copied to clipboard
| Challenge: | NL2SQL provides a model-centric paradigm that simplifies database access for non-technical users . challenges such as inaccurate task decomposition and keyword extraction remain major bottlenecks . |
| Approach: | They propose a RAG-based NL2SQL pipeline that employs three modules for query understanding, entity retrieval, and generation to improve SQL generation accuracy. |
| Outcome: | The proposed pipeline improves the accuracy of query generation on BIRD and Spider datasets. |
Copied to clipboard
| Challenge: | Existing knowledge injection methods are not suitable for enhancing pre-trained language models with external knowledge bases. |
| Approach: | They propose a plug-and-play knowledge injection method where knowledge bases are injected into frozen existing downstream models by a knowledge plugin. |
| Outcome: | The proposed method improves the performance of knowledge injection on knowledge-driven tasks while keeping model parameters frozen. |
Copied to clipboard
| Challenge: | Existing vision-and-language navigation methods do not incorporate environmental feedback into their decision-making processes. |
| Approach: | They propose a framework that incorporates environmental feedback into decision-making and a 3D simulator that renders realistic scenarios using Unreal Engine 5. |
| Outcome: | The proposed framework outperforms existing vision-and-language navigation methods in a zero-shot multi-task setting by 28.1% on average. |
Copied to clipboard
| Challenge: | Existing defense methods rely on fine-tuning or input modification, which suffer from limited generalization and reduced utility. |
| Approach: | They propose a finetuning-free approach that improves the defensive capabilities against jailbreak attacks of LLMs via targeted attention modification. |
| Outcome: | The proposed approach outperforms baselines in jailbreak defense and exhibits robust generalization across attacks and models, maintaining its effectiveness even on in-the-wild jailbreak data. |
Copied to clipboard
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Copied to clipboard
| Challenge: | Recent advances in neural theorem-proving resort to large language models and tree searches. |
| Approach: | They propose a Dynamic-Tree Driven Theorem Solver to accommodate general theoremes by guiding the search procedure with state confidence and proof-level values. |
| Outcome: | The proposed method outperforms state-of-the-art methods on two popular theorem-proving datasets with a 6.65% improvement on average in terms of success rate. |
Copied to clipboard
| Challenge: | Existing approaches to self-reflection rely on heuristic prompting or unidirectional reasoning traces. |
| Approach: | They propose a structured reflection method that transforms the "from error to repair" process into a first-class, controllable, and trainable action. |
| Outcome: | The proposed method improves multi-turn tool-call success rates and error recovery while reducing redundant calls. |
Copied to clipboard
| Challenge: | In-context learning (ICL) has gained considerable attention due to its data efficiency and task adaptability. |
| Approach: | They propose to de-biase demonstration bias in in-context learning by focusing on semantic ambiguity induced by demonstrations and reducing the semantic hazard. |
| Outcome: | The proposed methods significantly improve performance on six datasets. |
Copied to clipboard
| Challenge: | Existing methods for question decomposition focus on unimodal language models, but question decomposing capability of Multimodal Large Language Models (MLLMs) has yet to be explored. |
| Approach: | They propose a finetuning dataset and a training objective for selective decomposition to enhance the model's question decomposing capability. |
| Outcome: | The proposed dataset shows that existing models struggle to produce high-quality sub-questions. |
Copied to clipboard
| Challenge: | Existing MWP solvers do not handle variants that can be derived via mathematical manipulation. |
| Approach: | They propose a non-autoregressive solver to present a solution expression and decode it from a given problem description. |
| Outcome: | The proposed solver is able to decode multiple expression variants and correct them . it is based on a unified tree structure and is available on Math23K and MAWPS. |
Copied to clipboard
| Challenge: | Existing SOTA models segment long texts into equal-length snippets, but they have new challenges of context fragmentation and generalizability due to sentence boundaries and varying text lengths. |
| Approach: | They propose a Length-Aware Multi-Kernel Transformer to encode long documents by transformers and vectorize text length by the kernels to promote model robustness over varying document lengths. |
| Outcome: | The proposed model outperforms existing models on five benchmarks from health and law domains up to an absolute 10.9% improvement. |
Copied to clipboard
| Challenge: | Existing approaches to code generation rely on rejection sampling to generate multiple code snippets then select the best. |
| Approach: | They propose a framework that prioritizes sampling on test problems that models can solve. |
| Outcome: | The proposed framework reduces sampling costs while maintaining comparable code generation performance. |
Copied to clipboard
| Challenge: | Empirical evaluations across various prominent LLMs and benchmarks show that key-favored allocations retain up to 98.3% accuracy compared to uniform allocations (e.g., 4-bit keys, 2-bit values). |
| Approach: | They propose two theorems that anchor mixed-precision KV quantization in the intrinsic geometry of Transformer models. |
| Outcome: | Empirical evaluations show that key-favored allocations retain up to 98.3% accuracy while conserving memory. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are widely deployed as zero-shot evaluators for answer grading, content moderation, and document ranking. |
| Approach: | They propose a system that trains LLMs with adapters to denoise embeddings and refocus attention. |
| Outcome: | The proposed model lifts adversarial accuracy from 5% to 95% a 90 percentage-point gain while reducing clean-data accuracy by just 8 percentage points. |
Copied to clipboard
| Challenge: | InsightBuddy-AI is a system for extracting medication mentions and their associated attributes. |
| Approach: | They propose a system for extracting medication mentions and their associated attributes . they use stacked and voting ensembles built upon pre-trained language models . |
| Outcome: | The proposed system outperforms fine-tuned models in the extraction of medication mentions and associated attributes. |
Copied to clipboard
| Challenge: | Existing extractive models generate texts through word-by-word decoding, causing factual inconsistencies and slow inference. |
| Approach: | They propose a framework that integrates the behavior of copying EDUs into generative models. |
| Outcome: | The proposed framework reduces the number of generated tokens significantly. |
Copied to clipboard
| Challenge: | Recent advances in large language models (LLMs) have yielded remarkable performance, but objective mismatch issues hinder RLHF learning. |
| Approach: | They propose a Reinforcement Learning framework enhanced with Label-sensitive reward to enhance LLMs' alignment and generation capabilities. |
| Outcome: | The proposed framework improves performance on five diverse models across eight tasks. |
Copied to clipboard
| Challenge: | Entity matching (EM) is a critical step in entity resolution (ER). |
| Approach: | They propose a method that incorporates record interactions from different perspectives. |
| Outcome: | The proposed framework improves on 8 ER datasets and 10 LLMs and achieves higher efficiency and effectiveness. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Large Multimodal Models have exceeded general human capabilities in various tasks. |
| Approach: | They present an Olympiad-level bilingual multimodal scientific benchmark featuring 8,476 problems from Olympiad level mathematics and physics competitions. |
| Outcome: | The best performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning. |
Copied to clipboard
| Challenge: | LongInsightBench is the first benchmark designed to assess models’ ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements. |
| Approach: | They propose a benchmark to assess models’ ability to understand long videos with a focus on human language, viewpoints, actions, and other contextual elements. |
| Outcome: | The proposed model excels in three key areas: a) long-duration, human-centric videos; b) diversifying and challenging task scenarios; c) quality assurance pipeline; and d) reliability. |
Copied to clipboard
| Challenge: | Existing methods for multimodal content detection fail to capture cross-modal semantic inconsistencies and ignore inherent noise in multimodal features. |
| Approach: | They propose a multimodal rumor detection method based on a frequency domain spectral selection method and entropy-guided uncertainty fusion method to capture cross-modal semantic inconsistencies. |
| Outcome: | The proposed method outperforms state-of-the-art methods in multimodal rumor detection . it shows stronger detection capability and robustness on multiple datasets . |
Copied to clipboard
| Challenge: | Southeast Asia is underrepresented in vision-language research . SEA-VL is an open-source initiative dedicated to developing culturally relevant datasets for SEA languages. |
| Approach: | They propose to use crowdsourced, automated image crawling and synthetic image generation to develop culturally relevant datasets for SEA languages. |
| Outcome: | The proposed datasets capture SEA cultural nuances and contexts better than existing datasets. |
Copied to clipboard
| Challenge: | Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the reasoning process. |
| Approach: | They propose a three-stage framework based on complex question decomposition that decomposes the complex question, then reads the sub-questions and then performs numerical comparison to get the final answer. |
| Outcome: | The proposed framework achieves state-of-the-art in the 2WikiMultiHopQA dataset, with a winning joint F1 score of 53.58 on the leaderboard. |
Copied to clipboard
| Challenge: | Existing evaluations of LLMs in finance are text-only, monolingual, and largely saturated by current models. |
| Approach: | They propose a multilingual and multimodal benchmark for evaluating LLMs in real financial contexts. |
| Outcome: | The first expert-annotated multilingual and multimodal benchmark is released . it evaluates 21 leading LLMs and shows they perform better in multilingual settings . |
Copied to clipboard
| Challenge: | Existing supervised neural methods for coreference resolution are underexplored . current methods rely on small language models, but their potential is underexploited . |
| Approach: | They propose a framework that integrates an enhanced supervised model with LLM-based reasoning. |
| Outcome: | The proposed method surpasses existing state-of-the-art methods in coreference resolution. |
Copied to clipboard
| Challenge: | Large language models (LLMs) follow maliciously crafted instructions to generate deceptive responses, posing safety challenges. |
| Approach: | They use Sparse Autoencoders to analyze LLM's internal representations to determine when and how they "flip" from truthful to deceptive under deceptively crafted instructions. |
| Outcome: | The proposed model's True/False output is predictable across all conditions based on the model''s representation, and the Deceptive instructions induce significant representational shifts compared to Truthful/Neutral representations. |
Copied to clipboard
| Challenge: | Existing methods to detect LLM-generated content use simple hashes of precedent tokens to partition vocabulary. |
| Approach: | They propose a semantics-based watermark framework to enhance the robustness against paraphrase. |
| Outcome: | The proposed framework is robust under different paraphrases and the semantic meaning of the sentences will be likely preserved under paraphrase. |
Copied to clipboard
| Challenge: | Existing summarization systems alter the political opinions and stances of news articles in more than 50% of summaries, misrepresenting the intent and perspectives of the authors. |
| Approach: | They propose a model-based summarization approach controlled by political perspective classifiers that preserves the political stance of a generated summary. |
| Outcome: | The proposed model outperforms state-of-the-art summarization systems and large language models by up to 13.7% in terms of success rate of stance preservation, with competitive performance on standard metrics of summarizing quality. |
Copied to clipboard
| Challenge: | Existing methods for obtaining task-specific labels require prior knowledge of clustering categories and uncontrollable clustering centers. |
| Approach: | They propose a framework for supervised clustering using a discrete process and a robust Contrastive Learning module. |
| Outcome: | The proposed framework outperforms state-of-the-art models on a real-world dataset with just one label per class . the proposed framework is based on k-means clustering and a robust Contrastive Learning module . |
Copied to clipboard
| Challenge: | Existing approaches to joint Information Extraction (IE) neglect cross-instance or cross-task dependencies. |
| Approach: | They propose a joint IE framework that formulates joint 'conditional random field' to model cross-instance interactions . they incorporate a high-order neural decoder that is unfolded from a mean-field variational inference method . |
| Outcome: | The proposed approach improves on three IE tasks compared with baseline and prior work. |
Copied to clipboard
| Challenge: | Prompt-based learning is an emerging paradigm for exploiting knowledge learned by a pretrained language model. |
| Approach: | They propose a method to automatically select label mappings for few-shot text classification with prompting. |
| Outcome: | The proposed method achieves competitive performance on the GLUE benchmark without human effort or external resources. |
Copied to clipboard
| Challenge: | Existing paradigms rely on unreliable prompting or rigid constrained decoding strategies to achieve aesthetic unity. |
| Approach: | They propose a framework to embed external constraints into the model’s intrinsic intuition and use it to generate open-ended creative texts. |
| Outcome: | The proposed framework surpasses baselines in both strict constraint adherence and literary aesthetics. |
Copied to clipboard
| Challenge: | Existing approaches to support academic rebuttal rely on off-the-shelf LLMs or simple pipelines that struggle with long-context understanding. |
| Approach: | They propose an agentic framework for automatic academic rebuttal generation that operates through four steps: Decompose reviews into atomic concerns, Retrieve relevant evidence from the paper, Plan refortations, and Generate responses accordingly. |
| Outcome: | The proposed framework outperforms existing rebuttal pipelines and achieves 98% accuracy beyond the average human level using only an 8B model. |
Copied to clipboard
| Challenge: | Existing frameworks for fine-grained few-shot entity extraction are difficult to implement in the chemical domain due to the information overload of scientific papers. |
| Approach: | They propose a sequence-to-sequence based few-shot entity extraction approach . it uses a seq2seq entity extractor and a self-validation module to reconstruct original input sentence . |
| Outcome: | The proposed framework achieves 8.26% and 6.84% performance gains on two datasets. |
Copied to clipboard
| Challenge: | Existing evaluation datasets lack cross-lingual alignment, leaving assessments of multilingual capabilities fragmented in both language and skill coverage. |
| Approach: | They propose to use multilingual consistency as a complementary metric to assess performance bottlenecks and guide model improvement. |
| Outcome: | The proposed model lacks cross-lingual alignment and language coverage gaps between state-of-the-art models. |
Copied to clipboard
| Challenge: | Current approaches to detect hallucination require many samples from the LLM generator . current methods require multiple samples, which is computationally infeasible . |
| Approach: | They propose a simple baseline for detecting hallucinations in long-form LLM generations . they show that LLM hidden states are highly predictive of factuality in long form natural language generation . |
| Outcome: | The proposed method is comparable to expensive multi-sample approaches while drawing only a single sample from the LLM generator. |
Copied to clipboard
| Challenge: | Existing efforts to compress medium-sized models for specific tasks have limited results. |
| Approach: | They propose a task-agnostic compression toolkit for big models that implements quantization, pruning, distillation and MoEfication methods. |
| Outcome: | The proposed tool improves performance on a model with 3 billion parameters by 12x . it also outperforms the original model on three typical NLP benchmarks. |
Copied to clipboard
| Challenge: | a new method for learning unsupervised sentence embeddings is proposed . unsup-SimCSE is biased because of the length information encoded into the sentence embeds . |
| Approach: | They propose a new unsupervised sentence embedding method that uses dropout to obtain positive pairs from a pre-trained Transformer encoder. |
| Outcome: | The proposed method outperforms the state-of-the-art unsup-SimCSE on a STS task. |
Copied to clipboard
| Challenge: | Existing methods for capturing dialogue data are expensive and limited in their application. |
| Approach: | They propose a domain-agnostic extractive question answering approach with shared weights across domains to disentangle complex domain information in ToDs. |
| Outcome: | The proposed model can efficiently leverage domain-agnostic QA datasets while being domain-scalable and open vocabulary in DST. |
Copied to clipboard
| Challenge: | LegalBench evaluated 20 LLMs in 162 legal tasks in 20 countries and jurisdictions. |
| Approach: | They present a comprehensive evaluation of 21 popular Large Language Models and the first comparative analysis of the empirical results. |
| Outcome: | The proposed benchmarks are based on the Bloom’s cognitive taxonomy and are compared to 21 popular LLMs. |
Copied to clipboard
| Challenge: | Recent studies have used prompt-based fine-tuning methods for text classification tasks . however, the difficulty and costs of manually selecting domain label terms for the verbalizer remain unexplored . |
| Approach: | They propose a framework to automatically retrieve scientific topic-related terms for low-resource text classification tasks. |
| Outcome: | The proposed method outperforms state-of-the-art methods on scientific text classification tasks under few and zero-shot settings. |
Copied to clipboard
| Challenge: | Symbolic regression is a powerful technique for discovering mathematical expressions that best fit observed data. |
| Approach: | They propose a syntax-aware retrieval-augmented mechanism that leverages syntactic structure of symbolic expressions to perform context-awful retrieval from a pre-constructed token datastore. |
| Outcome: | The proposed method outperforms representative baselines on symbolic regression benchmarks and is validated on multiple symbolic regression datasets. |
Copied to clipboard
| Challenge: | Multimodal Process Reward Models (MPRMs) have emerged as a pivotal framework for enhancing the reasoning capabilities of Multimodal Large Language Models. |
| Approach: | They propose a benchmark specifically designed to evaluate MPRMs’ proficiency in detecting erroneous reasoning steps across diverse error categories. |
| Outcome: | The proposed model achieves up to 4.8% performance improvement through test-time scaling. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning and reinforcement learning use only positive examples, limiting their efficiency in low-resource scenarios. |
| Approach: | They propose a method that leverages both successful and failed trajectories for fine-tuning, maximizing the utility of limited resources. |
| Outcome: | The proposed method surpasses existing methods, including SFT, DPO, and PPO, across various tasks. |
Copied to clipboard
| Challenge: | Existing datasets focus on sentence-level event extraction, but document-level EE is limited due to the lack of large-scale and practical training and evaluation datasets. |
| Approach: | They propose a document-level event extraction dataset with 27,000+ events and 180,000+ arguments. |
| Outcome: | The proposed dataset includes 27,000+ events, 180,000+ arguments and large-scale manual annotations, fine-grained argument types and application-oriented settings. |
Copied to clipboard
| Challenge: | Existing research on information extraction tasks focuses on one specific task, but in real-world scenarios, new data of different IE tasks and domains come in a stream over time. |
| Approach: | They propose a parameter- and deployment-efficient prompt tuning method to evaluate the UIE system under a “lifelong learning” setting. |
| Outcome: | The proposed method is able to learn new tasks without forgetting old ones and expand knowledge and functionalities without retraining the whole system. |
Copied to clipboard
| Challenge: | Existing methods for expanding seed entities with new entities belong to the same semantic class are difficult to implement and can lead to accumulative errors. |
| Approach: | They propose an iterative set expansion framework that leverages automatically generated class names to address the semantic drift issue. |
| Outcome: | The proposed framework generates high-quality class names and outperforms state-of-the-art methods significantly. |
Copied to clipboard
| Challenge: | Prior research focused on developing data generation methods, while insufficient attention has been paid to quality control mechanisms and often produces inaccurate and unhelpful data. |
| Approach: | They propose an algorithm that automatically generates high-quality preference data, eliminating manual annotation requirements. |
| Outcome: | The proposed algorithm outperforms baselines in human preference alignment and reward optimization. |
Copied to clipboard
| Challenge: | Existing frameworks for large language model (LLM) inference on CPUs overlook overhead of cross-NUMA memory access. |
| Approach: | They propose a lightweight LLM inference architecture designed from the ground up for many-core CPUs. |
| Outcome: | Experimental results show that ArcLight surpasses the performance ceiling of mainstream frameworks, achieving up to 46% higher inference throughput. |
Copied to clipboard
| Challenge: | Parameter-shared pre-trained language models (PLMs) have emerged as a successful approach in resource-constrained environments. |
| Approach: | They propose a method to enhance the inference efficiency of parameter-shared PLMs by pre-training models that can achieve even greater acceleration. |
| Outcome: | The proposed method improves inference efficiency on autoregressive and autoencoding models. |
Copied to clipboard
| Challenge: | Existing approaches to optimize tool-use policies are monolithic and prone to entangling behaviors. |
| Approach: | They propose a framework that decomposes agent’stool-use policy into four modules and improves them via three mechanisms. |
| Outcome: | The proposed framework outperforms strong baselines on bothGPT-4.1 and Qwen3-8B while maintaining superior efficiency and transferability. |
Copied to clipboard
| Challenge: | UCSMNLP submitted to WAT 2019 Translation Tasks focusing on Myanmar-English translation. |
| Approach: | They propose to use Name Entity Recognition corpus and bilingual dictionary to build phrase based statistical machine translation system using listwise reranking process and initial distortion weight is changed to improve translation quality. |
| Outcome: | The proposed system outperforms the baseline system in the Myanmar-English translation task. |
Copied to clipboard
| Challenge: | Uncertainty identification is an important semantic processing task, critical to the quality of information in terms of factuality in many NLP techniques and applications. |
| Approach: | They propose to annotate Chinese microblogs with an open uncertainty corpus . they propose to use contextual uncertain semantics rather than traditional cue-phrases to identify uncertainty . |
| Outcome: | The proposed corpus can be used to identify uncertainty in social media texts. |
Copied to clipboard
| Challenge: | Existing studies on table reasoning focus on flat tables and hierarchical tables . a new dataset, HiTab, aims to examine numerical reasoning over hierarchic tables based on hierarchically structured tables - a strong challenge for existing baselines and a valuable benchmark for future research. |
| Approach: | They propose a hierarchical question answering and natural language generation dataset to study hierarchic tables. |
| Outcome: | The proposed model shows that it is effective in QA and natural language generation over hierarchical tables. |
Copied to clipboard
| Challenge: | Existing methods for parameter-efficient language model tuning (PELT) match the performance of fine-tuning with fewer trainable parameters. |
| Approach: | They propose a framework which integrates different PELT methods as submodules and learns to activate the ones that best suit the current data or task setup via gating mechanism. |
| Outcome: | The proposed framework outperforms fine-tuning methods on the GLUE benchmark and achieves 14% gains over the best individual PELT method. |
Copied to clipboard
| Challenge: | Recent studies have identified significant redundancy in large language models . quantization and pruning are two methods that reduce computational resources . |
| Approach: | They propose simple pruning methods that prune redundant layers based on their BI scores. |
| Outcome: | The proposed pruning methods demonstrate superior performance over previous pruning methods. |
Copied to clipboard
| Challenge: | Current GQA configurations overlook how context length influences inference cost . |
| Approach: | They propose a recipe for deriving cost-optimal GQA configurations that decouple the total head size from the hidden size and allow more flexible control over attention FLOPs. |
| Outcome: | The proposed configurations reduce memory usage and FLOPs by more than 50% compared to Llama-3's GQA, with *no degradation in model capabilities*. |
Copied to clipboard
| Challenge: | Existing approaches to climate research are limited to simple Q A tasks . a lack of data and computational expertise has created bottlenecks . |
| Approach: | They propose a general-purpose autonomous framework to perform end-to-end climate research tasks across diverse climate sub-fields. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks in rigorousness and practicality. |
Copied to clipboard
| Challenge: | Existing domain-specific pre-trained language models lack domain knowledge in domain-focused training. |
| Approach: | They propose a unified domain language model development service to inject domain knowledge into the PLM fine-tuning stage. |
| Outcome: | Experiments on domain-specific text classification and QA tasks verify the effectiveness and generalizability of KnowledgeDA. |
Copied to clipboard
| Challenge: | Large Audio-Language Models (LALMs) are augmented with the ability to perceive audio, but their reliability when faced with conflicting inputs remains largely unexplored. |
| Approach: | They examine how LALMs prioritize information when presented with inconsistent audio-text pairs. |
| Outcome: | The proposed models display a significant bias toward textual input when presented with inconsistent audio-text pairs. |
Copied to clipboard
| Challenge: | Prior work to mitigate fairness issues often employs subjective demonstration selection, leading to low controllability and limited stability across different models and tasks. |
| Approach: | They propose to use in-context learning to insert social biases into large language models to create a structured and controllable representation of the relationship between sensitive attributes and predicted labels. |
| Outcome: | Extensive experiments show that Fair-CCD consistently improves fairness metrics without degrading task accuracy. |
Copied to clipboard
| Challenge: | Existing algorithms for post-training large datasets are requiring a large computational effort. |
| Approach: | They propose to model the changes at logits level during post-training using a separate neural network . they demonstrate that the value network can be seamlessly integrated with another pre-trained model . |
| Outcome: | The proposed model can be integrated with another pre-trained model during inference, enabling similar capability enhancements. |
Copied to clipboard
| Challenge: | Hierarchical multi-label text classification (HMTC) aims to assign each text document to a set of relevant classes from a taxonomy. |
| Approach: | They propose to conduct HMTC based on only class surface names as supervision signals to mimic human experts. |
| Outcome: | The proposed framework outperforms the best existing method by 25% on two challenging datasets. |
Copied to clipboard
| Challenge: | Existing FL frameworks require a trusted aggregator or require heavy-weight cryptographic primitives, which makes the performance significantly degraded. |
| Approach: | They propose a framework that is federated and efficient for NLP . they propose to eliminate the need for trusted entities and achieve better model accuracy . |
| Outcome: | The proposed framework achieves better model accuracy and model accuracy than existing FL frameworks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly integrated into real-world decision-making, but their ability to comprehend and reason about policy-related content remains underexplored. |
| Approach: | They propose a bilingual benchmark evaluating policy comprehension comprising 21K cases across a broad spectrum of policy areas. |
| Outcome: | The proposed model shows stronger performance on application-oriented policy tasks than on memorization or conceptual understanding, and yields the highest accuracy on structured reasoning tasks. |
Copied to clipboard
| Challenge: | Existing methods for code summarization do not capture rich information in ASTs . existing methods are labor-intensive and time-consuming to document code with good summaries manually. |
| Approach: | They propose a model that hierarchically splits and reconstructs ASTs by a neural network . they propose to use AST embeddings and a vanilla code token encoder to generate the model . |
| Outcome: | The proposed model splits and reconstructs ASTs into subtrees and then aggregates embeddings of subtreas to get the complete AST. |
Copied to clipboard
| Challenge: | Existing literature on dialog memory systems is inconsistent on their effectiveness . empirical findings on graph structures are difficult to attribute to specific design choices . |
| Approach: | They propose a framework that decomposes dialog memory systems into core components . they conduct stage-wise experiments on LongMemEval and HaluMeM, and compare implementation details . |
| Outcome: | The proposed framework compares graph-based and non-graph memory architectures on long-term dialog memory systems. |
Copied to clipboard
| Challenge: | Existing methods to improve code generation from natural language descriptions are difficult due to complex structure, subtle bugs, and lack of supplementary contents. |
| Approach: | They propose a framework that enhances complex code generation by online searching for more information with planned queries and correctness testing for code refinement. |
| Outcome: | The proposed framework improves the quality of complex code generation on the DS-1000 and ClassEval datasets. |
Copied to clipboard
| Challenge: | Existing approaches to multimodal speech emotion recognition and sentiment analysis have not improved results due to their relatively simple fusion mechanisms and lack of proper cross-modal pretraining. |
| Approach: | They propose a deep-fused audio-text bi-modal transformer with carefully designed cross-modal fusion mechanism and stage-wise cross-mod pretraining scheme to facilitate cross-modulation. |
| Outcome: | The proposed method exceeds benchmarks on public IEMOCAP emotion and CMU-MOSEI sentiment datasets by a large margin. |
Copied to clipboard
| Challenge: | Existing in-context learning methods for relation extraction often overlook entity relationships . Existing methods for RE prioritize language similarity over structural similarity . |
| Approach: | They propose an AMR-enhanced retrieval-based ICL method for relation extraction . their method retrieves in-context examples based on semantic structure similarity . |
| Outcome: | The proposed method outperforms baselines on four English RE datasets and in the more demanding unsupervised setting. |
Copied to clipboard
| Challenge: | Existing studies have focused on the effectiveness of contrastive learning in deep learning. |
| Approach: | They propose a method to improve sentence representation of unsupervised contrastive learning by examining the role of temperature in VRL and SRL. |
| Outcome: | The proposed method improves representation of unsupervised contrastive learning by cooling the temperature of the representation space. |
Copied to clipboard
| Challenge: | Existing scaling laws for language models are limited to a limited number of languages, but they can be applied to arbitrary number of different languages. |
| Approach: | They propose a scaling law for general-purpose decoder-only language models trained on multilingual data that shifts focus from individual languages to language families. |
| Outcome: | The proposed scaling law can be applied to models trained on multilingual data . it can be used to predict performance across multiple languages and models . |
Copied to clipboard
| Challenge: | Existing methods for product attribute value identification suffer from cascading errors and lack of generalization capability. |
| Approach: | They propose a multi-level retrieval scheme that uses products and attribute values as distinct hierarchical levels in PAVI domain. |
| Outcome: | The proposed method performs better than the state-of-the-art methods on a real-world industrial dataset. |
Copied to clipboard
| Challenge: | Recent advances in fine-tuning large language models have greatly enhanced their usage in domain-specific tasks. |
| Approach: | They propose a method which internalizes prompt knowledge during model fine-tuning to achieve efficient inference and save costs. |
| Outcome: | The proposed approach reduces input tokens by 90%, accelerates inference by 4.2 times, and reduces monetary inference costs by 88.3%. |
Copied to clipboard
| Challenge: | Recent explosion of performance of large language models (LLMs) has changed the field more abruptly and seismically than any other shift in the field’s 80 year history. |
| Approach: | They propose 20+ PhD-dissertation-worthy research directions to define a new NLP playground by combining theoretical analysis, new and challenging problems, learning paradigms and interdisciplinary applications. |
| Outcome: | The proposed research will cover theoretical analysis, new and challenging problems, learning paradigms and interdisciplinary applications. |
Copied to clipboard
| Challenge: | a promising baseline SimCSE has made notable breakthroughs in unsupervised SRL . however, there is still room for designing a novel contrastive framework specifically targeted for SRL. |
| Approach: | They propose an angle-based similarity function for a contrastive objective and propose a new approach for SRL. |
| Outcome: | The proposed approach shows better training dynamics on SRL than the standard cosine similarity function. |
Copied to clipboard
| Challenge: | Existing supervised fine-tuning (SFT) fails to address these issues, as it trains models on single gold-standard responses without modeling nuanced strategy trade-offs. |
| Approach: | They propose a two-stage framework that optimizes strategy selection preferences at each dialogue turn. |
| Outcome: | The proposed framework improves strategy selection preferences at each dialogue turn. |
Copied to clipboard
| Challenge: | Unstructured natural language explanations lack a comprehensive explanation mechanism to verify a model's true reasoning capabilities. |
| Approach: | They propose a reward engineering method which uses semi-structured explanations to verify a model's true reasoning capabilities. |
| Outcome: | The proposed method achieves new state-of-the-art on two semi-structured explanation generation benchmarks (ExplaGraph and COPA-SSE) . |
Copied to clipboard
| Challenge: | Money laundering (AML) is the process of transferring criminal and illegal proceeds into ostensibly legitimate assets. |
| Approach: | They propose a framework that uses deep learning to augment AML monitoring and investigation . money laundering is the process of transferring criminal and illegal proceeds into ostensibly legitimate assets . |
| Outcome: | The proposed framework reduces time and cost by 30% compared to existing methods . money laundering is the world's third largest "industry" |
Copied to clipboard
| Challenge: | a new framework for hate speech detection addresses implicit hate speech by tailoring the detection process to dataset-specific attributes. |
| Approach: | They propose a framework to account for the dataset-specific characteristics of hate speech datasets. |
| Outcome: | The proposed framework improves detection accuracy and provides interpretable insights into the distinctive features of each dataset. |
Copied to clipboard
| Challenge: | Existing methods for finding the optimal prompt for a task are difficult to optimize. |
| Approach: | They propose an efficient discrete prompt optimization approach with reinforcement learning that generates the optimal discrete stimulus after training with reward. |
| Outcome: | The proposed approach is based on a parameter-efficient policy network that generates the optimal discrete prompt after training with reward. |
Copied to clipboard
| Challenge: | balancing the training budget, downstream performance, and general capabilities of large language models remains a challenge in many applications. |
| Approach: | They propose a mixture of expert framework based on Soft LoRA and Identity Mixture . SLIM allows dynamic routing between LoRA adapters and identity layers . |
| Outcome: | The proposed framework reduces training cost while maintaining general capabilities . it can be open-sourced upon publication. |
Copied to clipboard
| Challenge: | Existing methods only conduct network growth in a single dimension, but compound growth operators are beneficial for multiple dimensions. |
| Approach: | They propose a method to train BERT progressively using a Transformer model and explore alternative growth operators in each dimension via controlled comparison. |
| Outcome: | The proposed method speeds up BERT pre-training by 73.6% and 82.2% for the base and large models respectively while achieving comparable performances. |
Copied to clipboard
| Challenge: | Existing Language Agents neglect the notion of uncertainty during interactions with external worlds. |
| Approach: | They propose a framework that orchestrates the interaction between the agent and the external world using uncertainty quantification. |
| Outcome: | The proposed framework improves performance on 3 representative tasks and lowers reliance on external world. |
Copied to clipboard
| Challenge: | NormLens is a visual-grounded framework for understanding commonsense norms . state-of-the-art models are not well-aligned with human annotation, we show . |
| Approach: | They propose a visual-grounded framework to study commonsense norms by NormLens . they find that models are not well-aligned with human annotation . |
| Outcome: | The proposed model judgments and explanations are not well-aligned with human annotations. |
Copied to clipboard
| Challenge: | Existing work on multimodal sentiment analysis relies on back-propagated task loss or geometric property of feature spaces to produce favorable fusion results. |
| Approach: | They propose a framework which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimod input to maintain task-related information through multimodal integration. |
| Outcome: | The proposed framework maximizes the Mutual Information (MI) in unimodal input pairs and between multimodal fusion result and unimodulated input to maintain task-related information through multimodal integration. |
Copied to clipboard
| Challenge: | Semantic frames are conceptual structures that describe specific types of situations or events. |
| Approach: | They propose to generate frame definitions from a set of frame-evoking words using a large language model. |
| Outcome: | The proposed task incorporates frame element reasoning as chain-of-thought to enhance the inclusion of correct frame elements in the generated definitions. |
Copied to clipboard
| Challenge: | Existing methods to induce grammars of multiple languages do not consider language similarity measures. |
| Approach: | They propose a universal grammar induction approach that captures similarity between languages . they use vector representations to capture similarity and softly tie grammar parameters . |
| Outcome: | The proposed approach performs well over monolingual and multilingual datasets. |
Copied to clipboard
| Challenge: | Knowledge distillation (KD) is a technique for transferring expertise from large teacher models to compact student models with reduced memory footprints and inference costs. |
| Approach: | They propose to transfer knowledge from large teacher models to compact student models by exploiting teacher-student capacity discrepancies to generate pseudo-preference pairs where teacher outputs are preferred over student outputs. |
| Outcome: | The proposed framework exploits teacher-student capacity discrepancy to generate pseudo-preference pairs where teacher outputs are preferred over student outputs. |
Copied to clipboard
| Challenge: | Existing multilingual grammar induction methods require external resources such as parallel corpora, word alignments or linguistic phylogenetic trees. |
| Approach: | They propose a framework in which the learning process of the grammar model of one language is influenced by knowledge from the model of another language. |
| Outcome: | The proposed method outperforms baselines on transfer grammar induction and bilingual grammar inducing on multiple languages. |
Copied to clipboard
| Challenge: | Existing methods for instruction tuning do not include associating instructions with existing datasets. |
| Approach: | They propose a dynamic growth paradigm for the automatic curation of instruction-tuning data . they use existing datasets to automatically construct instruction-uning datasets . |
| Outcome: | The proposed model reduces the API cost for generating instructions and provides high-quality data. |
Copied to clipboard
| Challenge: | Many essays are submitted to tutoring services by English learners on the Web every day . few systems provide focused suggestions on how to raise the level of proficiency. |
| Approach: | They propose a method for generating suggestions on a sentence for improving proficiency . they propose identifying grammatical elements and ranking related elements to provide suggestions . |
| Outcome: | The proposed method helps english learners improve their writing and reading skills. |
Copied to clipboard
| Challenge: | Existing studies on pretraining of LLMs on extensive web-based texts are insufficient for advanced scientific discovery, especially in chemistry. |
| Approach: | They outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs and conceptualize chemistry LLM agents using chemistry tools. |
| Outcome: | The proposed models are based on domain-specific chemistry knowledge and multi-modal information and are capable of accelerating scientific research. |
Copied to clipboard
| Challenge: | Experimental results show that stories outperform rules as the expression for retrieving commonsense from LLMs, exhibiting higher generation confidence and commonsensense accuracy. |
| Approach: | They investigate the commonsense ability of large language models expressed through stories and rules to retrieve commonsensing knowledge from LLMs. |
| Outcome: | The stories outperform rules as commonsense expressions on 28 commonsensense QA datasets, exhibiting higher generation confidence and commonsence accuracy. |
Copied to clipboard
| Challenge: | Existing studies show that training-based methods are ineffective to detect LLM generated texts from unseen tasks or topics which are not collected during training. |
| Approach: | They propose to train classification models to distinguish LLMs from human texts by a distribution shift caused by prompts, text lengths, topics, and language tasks. |
| Outcome: | The proposed methods can detect LLMs from black-box models, but they suffer from distribution shifts due to a wide range of factors, including prompts, text lengths, topics, and language tasks. |
Copied to clipboard
| Challenge: | Recent event-centric reading comprehension datasets focus mostly on event arguments or temporal relations. |
| Approach: | They propose a machine reading comprehension dataset that leverages natural language queries to reason about the five most common event semantic relations. |
| Outcome: | The proposed dataset shows that current SOTA systems achieve 22.1%, 63.3% and 83.5% for token-based exact-match, **F1** and event-based **HIT@1** scores. |
Copied to clipboard
| Challenge: | Existing methods for semi-supervised text classification have shown great performance in few-shot scenarios, where both labeled and unlabeled data are utilized. |
| Approach: | They propose a simple instance-adaptive self-training method for semi-supervised text classification that generates two augmented views for each unlabeled data and trains a meta learner to identify relative strength of augmentations based on the similarity between the original view and the augmented view. |
| Outcome: | The proposed method consistently shows competitive performance with varying sizes of labeled training data compared to existing semi-supervised learning methods. |
Copied to clipboard
| Challenge: | Code search is to search reusable code snippets from source code corpus based on natural languages queries. |
| Approach: | They propose a method to accelerate code search with deep hashing and code classification by using deep hashes and code hash. |
| Outcome: | The proposed method can save 90% of retrieval time while preserving at least 99% of retrievals accuracy. |
Copied to clipboard
| Challenge: | Existing studies focus on monolingual hypernymy detection on high-resource languages, but few investigate low-resourced scenarios. |
| Approach: | They propose to combine high-resource languages to solve low-resourced hypernymy detection problem . they extensively compare three joint training paradigms and propose meta learning . |
| Outcome: | The proposed method significantly improves performance of extremely low-resource languages by preventing over-fitting on small datasets. |
Copied to clipboard
| Challenge: | Existing methods for testing harmful information on social media rely on fixed parameters that fail to handle substantial semantic discrepancies . RLAT can be used to adapt to semantic variations while preventing overfitting from continuous tuning. |
| Approach: | They propose a reinforcement learning-guided adaptive tuning method for harmful text detection that optimizes consistency loss and applies word-level attention constraints to reduce over-reliance on local words. |
| Outcome: | The proposed method outperforms state-of-the-art models in cross-platform and cross-temporal scenarios across multiple public datasets. |
Copied to clipboard
| Challenge: | Virtual adversarial training (VAT) is a powerful approach to improving robustness and performance, leveraging both labeled and unlabeled data to compensate for the scarcity of labeles. |
| Approach: | They propose a Sparse Parse Adjustment algorithm which combines VAT and a graph-based dependency parsing model in an exact computational manner and enhances the dependency parsed with controllable and adjustable sparsity. |
| Outcome: | Empirical results show that the proposed algorithm outperforms other methods without sparsity regularization. |
Copied to clipboard
| Challenge: | Tabular data are crucial in many fields and their understanding by large language models (LLMs) under high parameter efficiency paradigm is important. |
| Approach: | They propose a module that uses 2D LoRA to encode low-rank information on cell positions to improve table serialization and representation of two-dimensional structured information within a one-dimensional sequence. |
| Outcome: | Experiments on four tabular-related datasets show that TableLoRA outperforms vanilla LoRA and surpasses table encoding methods tested in control. |
Copied to clipboard
| Challenge: | Prior work in ABSA has investigated opinion extraction as an important subtask, but these works only label concise, *explicitly*-stated opinion spans. |
| Approach: | They propose a new ABSA dataset with implicit opinion span annotations . they use paragraph-length inputs and prompted-LLM baselines to evaluate the dataset . |
| Outcome: | The proposed dataset presents significant challenges for fully-supervised models and LLMs. |
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) can generate coherent text, but they struggle to recognise user intent behind queries. |
| Approach: | They propose a novel approach leveraging multi-level intent, domain, and slot knowledge distillation for multi-turn NLU. |
| Outcome: | The proposed model improves multi-turn conversation understanding by integrating teacher teachers into a student model. |
Copied to clipboard
| Challenge: | Several studies have explored delta parameter properties via pruning, quantization, low-rank approximation, and extrapolation, but what properties of delta parameters are essential for maintaining performance? |
| Approach: | They propose to examine delta parameter properties along magnitude and sign . they propose to use a loss-based local surrogate analysis to examine editing effects . |
| Outcome: | The proposed analysis shows that delta parameters can be edited while maintaining performance. |
Copied to clipboard
| Challenge: | Domain Adaptive Continual Pretraining (DACP) is a method to mitigate performance degradation in small LLMs and enhance their effectiveness in target domains. |
| Approach: | They propose a continual pretraining methodology that optimizes sLLMs within service domains and enhances their effectiveness in target domains. |
| Outcome: | The proposed model achieves significant gains in target-domain performance while preserving general capabilities, offering a cost-efficient and scalable solution for enterprise-level deployment. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have been gaining popularity in multimodal tasks . a bilingual benchmark is available for MLLM users to evaluate their multimodal capabilities . |
| Approach: | They propose a bilingual multimodal ability norms benchmark that measures multimodality across nine tasks. |
| Outcome: | The proposed benchmark compared human performance against state-of-the-art MLLMs. |
Copied to clipboard
| Challenge: | Automated Alignment (ALM) is a set of algorithms designed to align Large Language Models (LLMs) with human intentions and values while minimizing manual intervention. |
| Approach: | They propose an open-source toolkit that integrates mainstream automated algorithms through a consistent interface and an accessible workflow supporting one-click execution for prompt synthesis and automatic alignment signal construction. |
| Outcome: | The proposed framework enables easy reproduction of existing results through extensive benchmarks and facilitates the development of novel approaches via modular components. |
Copied to clipboard
| Challenge: | In this paper, we consider mimicking fictional characters as a promising direction for building engaging conversation models. |
| Approach: | They propose a task where only a few utterances of each fictional character are available to generate responses mimicking them. |
| Outcome: | The proposed method generates responses better reflecting the style of fictional characters than baseline methods. |
Copied to clipboard
| Challenge: | Existing machine reading comprehension datasets lack an explainable evaluation of systems' reasoning capabilities. |
| Approach: | They propose a dataset with multi-choice questions that evaluates MRC systems' reasoning process . they use sentence-level relevant supporting facts, error reason of distractors to evaluate MRC . |
| Outcome: | The proposed dataset is more challenging and useful for identifying limitations of existing MRC systems in an explainable way. |
Copied to clipboard
| Challenge: | Bipolar Disorder (BD) is a mental disorder characterized by intense mood swings, ranging from depression to manic states. |
| Approach: | They propose to use social media data to identify BD risk in individuals misdiagnosed as MDD by multi-task learning. |
| Outcome: | The proposed approach outperforms state-of-the-art baselines and can provide insights into the impact of BD mood on future risk. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can be used in psychotherapy to overcome challenges such as shame, distrust, and resource scarcity. |
| Approach: | They propose a cognitive reframing therapy method that uses empathetic dialogue to address deep-rooted negative thoughts and fosters rational, balanced perspectives. |
| Outcome: | The proposed model outperforms other models in terms of empathy, guidance, and logical coherence, demonstrating its effectiveness and potential positive impact on psychotherapy. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can be used to broaden user experiences beyond established preferences and reinforce feedback loops. |
| Approach: | They propose a hierarchical approach that combines hierarchic planning with LLM inference-time scaling to improve recommendation relevancy without compromising novelty. |
| Outcome: | The proposed approach shows significant gains in both user satisfaction and exploration diversity. |
Copied to clipboard
| Challenge: | Existing linear transformers suffer from performance degradations on various tasks and corpus. |
| Approach: | They propose a new linear attention that replaces scaling with a normalization to stabilize gradients and confine attention to neighbouring tokens in early layers. |
| Outcome: | The proposed model outperforms vanilla transformers on the long-range arena benchmark while being significantly more space-time efficient. |
Copied to clipboard
| Challenge: | a growing demand for Large Language Models (LLMs) is requiring specialized models to augment customer service agents' skills. |
| Approach: | They propose a methodology for developing a specialized Telecommunications LLM . they use a dataset to evaluate customer service expertise in the telecommunications domain . |
| Outcome: | The proposed model improves the efficiency of customer service agents and reduces response times. |
Copied to clipboard
| Challenge: | Experimental evaluations show that RL methods favor outliers rather than truly informative samples under low-resource and class-imbalanced conditions. |
| Approach: | They propose a robust sample selection strategy using reinforcement learning to identify the most informative samples using a class imbalance approach. |
| Outcome: | The proposed strategy improves model transferability while maintaining robust performance under extreme class imbalance compared to traditional methods. |
Copied to clipboard
| Challenge: | Existing models that generate semantically correct regular expressions from NLs are not yet fully understood. |
| Approach: | They propose a model that rewards reinforcement learning based on the semantic equivalence between two regular expressions. |
| Outcome: | The proposed model reduces training time and produces state-of-the-art results on three benchmark datasets. |
Copied to clipboard
| Challenge: | Evaluation benchmarks based on predefined domains and human-labeled data face limitations in addressing evaluation needs for emerging domains. |
| Approach: | They propose an automated information retrieval benchmark based on predefined domains and human-labeled data . AIR-Bench is automated and Heterogeneous with three key features . |
| Outcome: | The proposed benchmarks are based on predefined domains and human-labeled data. |
Copied to clipboard
| Challenge: | Current language models lack the structured deliberation needed for high-stakes tasks such as healthcare and finance. |
| Approach: | They propose a decision-making framework that guides models to reason over structured representations of actions, attributes, and constraints. |
| Outcome: | The proposed framework achieves up to 30% accuracy gains over strong prompting baselines and enhances alignment in outcomes. |
Copied to clipboard
| Challenge: | EVIDENCEMINER is a web-based system that allows users to query a natural language statement and retrieve textual evidence from a background corpora for life sciences. |
| Approach: | They propose a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
| Outcome: | EVIDENCEMINER is a web-based system that lets users query a natural language statement and automatically retrieves textual evidence from a background corpora for life sciences. |
Copied to clipboard
| Challenge: | Long-form question answering requires two procedures: information retrieval and information synthesis. |
| Approach: | They propose a Chinese long-form question answering dataset called WebCPM . the dataset is based on a web search interface that engages with a search engine in real time . |
| Outcome: | The proposed dataset generates answers that are no worse than human-written ones . the dataset is the first Chinese LFQA dataset . |
Copied to clipboard
| Challenge: | InfiMM-WebMath-40B is a dataset of interleaved image-text documents . it consists of 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
| Approach: | InfiMM-WebMath-40B is a high-quality dataset of interleaved image-text documents . it contains 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
| Outcome: | InfiMM-WebMath-40B is a high-quality dataset of interleaved image-text documents . it consists of 24 million web pages, 85 million image URLs, and 40 billion text tokens . |
Copied to clipboard
| Challenge: | Prior studies have shown that sequence-to-sequence models learn to hallucinate when the conditioning data has poor correlation with the sequence being produced. |
| Approach: | They construct a dataset that pairs Knowledge Graphs (KG) and text together and compare their results to a cyclic evaluation model. |
| Outcome: | The proposed model performs better on cyclic generation of KGs than on KG-T, but less well on synchronization of KTs. |
Copied to clipboard
| Challenge: | Current reinforcement learning methods suffer from coarse-grained, trajectory-level rewards that provide insufficient learning signals for complex multi-turn interactions, leading to training stagnation. |
| Approach: | They propose a novel RL algorithm for training large language models for multi-turn tool-integrated reasoning (TIR) that incorporates three innovations: turn-level reward assignment that provides fine-grained feedback for individual turns, return-based advantage estimation where normalized discounted returns are calculated as advantages, and self-supervised reward shaping that exploits self-supervision signals from generated code to densify sparse binary outcome-based rewards. |
| Outcome: | The proposed algorithm outperforms GRPO by 3.0% across diverse math reasoning benchmarks and improves grepo by 3.9% on commonsense reasoning and program synthesis tasks. |
Copied to clipboard
| Challenge: | Existing models for text classification are based on encoder-only transformers and generative pre-trained transformers. |
| Approach: | They propose an uncertainty-aware contrastive sentence embedding approach that addresses language ambiguity and inter-class separability for a text classification task. |
| Outcome: | The proposed approach improves classification accuracy on public datasets compared with state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing studies treat each transformer encoding layer as a single artificial neuron . layer-level embeddings aggregate multiple types of contextual attention captured by multiple head modules . |
| Approach: | They propose to embed each transformer encoding layer as a single artificial neuron . they propose to couple those ANs with their biological-neuron counterparts in the human brain . |
| Outcome: | The proposed models can be used to link representations to brain activity, the authors say . their results show that the proposed models carry meaningful neurolinguistic information . |
Copied to clipboard
| Challenge: | Existing diffusion models for continuous-valued domains have not been adopted for text data. |
| Approach: | They propose a diffusion-based language model with two key design choices . semi-autoregressive model generates blocks of text and allows local context updates . they evaluate it on unconstrained text generation benchmarks . |
| Outcome: | The proposed model outperforms autoregressive models on unconstrained text generation benchmarks on uncontrolled text generation. |
Copied to clipboard
| Challenge: | Current evaluations measure functional correctness on well-formed inputs, but they filter out inputs that violate them. |
| Approach: | They propose a benchmark to evaluate whether generated code enforces preconditions . they use a neuro-symbolic pipeline to evaluate code with test cases . |
| Outcome: | The proposed benchmark aims to evaluate whether generated code enforces preconditions . it aims at achieving pass@k scores while ignoring those that violate them . |
Copied to clipboard
| Challenge: | Entity alignment (EA) methods identify the aligned entities based on cosine similarity, ignoring the semantics underlying the embeddings themselves. |
| Approach: | They propose to model entity alignment as a sequential decision-making task where an agent sequentially decides whether two entities are matched or mismatched based on representation vectors. |
| Outcome: | The proposed framework consistently advances the performance of several state-of-the-art methods, with a maximum improvement of 31.1% on Hits@1. |
Copied to clipboard
| Challenge: | Recent studies have achieved inspiring success in unsupervised grammar induction using masked language modeling (MLM) as the proxy task. |
| Approach: | They propose to regularize the parser with phrases extracted by an unsupervised phrase tagger to help the LM model quickly manage low-level structures. |
| Outcome: | The proposed method improves the identification of high-level structures using phrase-guided masking. |
Copied to clipboard
| Challenge: | Using the ABT structure, academic abstracts are structured to provide clear and concise prose, but a lack of clarity and logical coherence is a challenge for authors struggling with English proficiency or academic writing conventions. |
| Approach: | They propose a framework that identifies the key components of an abstract and reorients itself to properly reflect the ABT logical progression. |
| Outcome: | The proposed framework improves comprehensibility of academic writing, particularly for non-native English speakers, and is based on a human evaluation and automated metrics. |
Copied to clipboard
| Challenge: | Existing fingerprinting methods require impractical white-box access or introduce detectable statistical anomalies. |
| Approach: | They propose a gray-box fingerprinting framework that ensures stealthy and robust model provenance tracing. |
| Outcome: | The proposed framework is the first to repurpose Membership Inference Attacks (MIAs) for defensive use, embedding ownership signals via memorization instead of artificial trigger-output overfitting. |
Copied to clipboard
| Challenge: | Recent advances in multimodal recommenders lack explicit reasoning and self-awareness of uncertainty. |
| Approach: | They propose a reasoning-augmented multimodal agent structured around a three-stage explicit reasoning pipeline. |
| Outcome: | The proposed agent improves ranking metrics and performance on four standard recommendation tasks across five real-world datasets. |
Copied to clipboard
| Challenge: | Existing studies seek to enhance the graph reasoning capabilities of Large Language Models (LLMs) by specialized instruction tuning. |
| Approach: | They propose to evaluate LLM graph reasoning generalization using in-distribution settings . they propose to use three strategies to improve LLM generalization . |
| Outcome: | The proposed benchmark evaluates LLM graph reasoning generalization with in-distribution settings only . it shows that LLMs struggle to generalize across reasoning and real-world patterns . |
Copied to clipboard
| Challenge: | Existing automated evaluation metrics for machine translation are expensive and lack inter-rater reliability. |
| Approach: | They propose a task-oriented and human-centric evaluation framework for machine translation output based on professional post-e diting annotations. |
| Outcome: | The proposed framework improves translation quality and system performance and transparency . it is cost-effective, easy to use and faster to implement . |
Copied to clipboard
| Challenge: | Existing work on long document visual question answering is based on Retrieval-Augmented Generation (RAG) where textual or visual content is encoded into embeddings and relevance is determined by similarity scores with respect to the original query. |
| Approach: | They propose a framework that employs an agentic, vision-aware workflow to address long document visual question answering through iterative information discovery and synthesis. |
| Outcome: | The proposed framework outperforms existing RL systems by 10.4% on the MMLongbench-Doc benchmark and demonstrates superior training performance over GRPO. |
Copied to clipboard
| Challenge: | Current approaches to detect hate speech rely on contrastive learning to distinguish hate from non-hate sentences. |
| Approach: | They propose a novel approach to detect implicit hate speech by identifying explicit targets . they use a pretrained Named Entity Recognition model to capture explicit target information . |
| Outcome: | The proposed approach outperforms current methods and achieves faster convergence. |
Copied to clipboard
| Challenge: | Existing approaches to interpret black-box models to learn spurious correlations are not well understood. |
| Approach: | They propose a procedure that leverages model interpretations to update parameters towards a plausible interpretation rather than an interpretation that relies on spurious patterns in data. |
| Outcome: | The proposed procedure outperforms baseline methods that use adversarial training in a controlled setup. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are susceptible to a type of attack known as jailbreaking, which misleads LLMs to output harmful contents. |
| Approach: | They propose to leverage hidden representations into existing jailbreak targets to move the attacks along the acceptance direction. |
| Outcome: | The proposed methods are validated using the objective of existing jailbreak attacks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are increasingly used in socially complex, interaction-driven tasks, yet their ability to mirror human behavior in emotionally and strategically complex contexts remains underexplored. |
| Approach: | They examine alignment of personality-prompted Large Language Models in conflict dialogues that incorporate negotiation by simulating a five-factor personality profile. |
| Outcome: | The proposed model achieves the closest alignment with humans in linguistic style and emotional dynamics while Claude-3.7-Sonnet best reflects strategic behavior. |
Copied to clipboard
| Challenge: | Existing studies have explored textual graph descriptions and visual modalities for VLMs to understand graphs. |
| Approach: | They propose a unified framework that enhances both scalability and modality coordination in graph understanding by integrating textual and visual modalities. |
| Outcome: | GraphVista scales to large graphs, 200 larger than those used in existing benchmarks, and consistently outperforms existing textual, visual, and fusion-based methods. |
Copied to clipboard
| Challenge: | Reinforcement learning (RL) has shown strong promise for LLM-based machine translation . however, translation-oriented RL remains challenged by high-variance policy gradients induced by Monte Carlo baselines and large trajectory space that favors global exploration over fine-grained local optimization. |
| Approach: | They propose a two-stage RL framework that uses post-editing as an auxiliary task to stabilize training and guide overall optimization. |
| Outcome: | The proposed framework supports global exploration and fine-grained optimization while supporting global exploration. |
Copied to clipboard
| Challenge: | Visual persuasion uses visual elements to influence cognition and behaviors . lack of comprehensive data sets connect persuasiveness of images with personal information . |
| Approach: | They propose to use a dataset to connect persuasiveness with personal information . they find psychological characteristics enhance the generation and evaluation of persuasive images . |
| Outcome: | The proposed dataset provides persuasiveness scores of images evaluated by human annotators along with demographic and psychological characteristics. |
Copied to clipboard
| Challenge: | Distantly supervised relation extraction (RE) has attracted much attention in the past few years . previous methods to evaluate models manually or directly on autolabeled data have produced inaccurate evaluations . |
| Approach: | They propose to use distant supervision to generate large-scale autolabeled data . they build manually-annotated test sets for two DS-RE datasets and evaluate models . |
| Outcome: | The proposed method produces 53% wrong labels at the entity pair level in the popular NYT10 dataset. |
Copied to clipboard
| Challenge: | Existing multimodal foundation models suffer from serious factual inaccuracy in radiology report generation. |
| Approach: | They propose a fact-aware multimodal retrieval-augmented pipeline for generating accurate radiology reports using RadGraph. |
| Outcome: | The proposed multimodal retrieval-augmented pipeline outperforms state-of-the-art retrievers on language generation and radiology-specific metrics. |
Copied to clipboard
| Challenge: | Current methods for information extraction (IE) focus on integrating IE output with the database . a long-overlooked question is what counts as "relevant knowledge" |
| Approach: | They propose a task that emphasizes integration of IE output and the database . they introduce a benchmark and an LLM agent framework for this task . |
| Outcome: | The proposed task integrates IE output and the target database (or knowledge base) it meets common demands such as data infilling, row population, and column addition . |
Copied to clipboard
| Challenge: | Existing methods for dialogue state tracking still have a JGA of 60% on MultiWOZ 2.1 . break framework provides a simple yet effective way to generate dialogue state candidates . |
| Approach: | They propose a framework that generates k-best dialogue state candidates with beam search and re-ranks them to select the correct dialogue state. |
| Outcome: | The proposed framework pushes the joint goal accuracy to 80-90% on MultiWOZ 2.1-2.4. |
Copied to clipboard
| Challenge: | despite its potential to help users, NLP research on explicitation is limited because of the lack of adequate evaluation methods. |
| Approach: | They propose automatic methods to generate explicitations from a Wikipedia dataset . they use both intrinsic and extrinsic evaluation to evaluate the system's effectiveness . |
| Outcome: | The proposed system bridges the gap between the source speaker and the target audience . it is effective based on intrinsic and extrinsic evaluation, the authors show . |
Copied to clipboard
| Challenge: | Knowledge-to-text generators often struggle to faithfully generate descriptions for input facts . we propose a decoding-only method to reduce hallucinations . |
| Approach: | They propose a decoding-only method to generate accurate descriptions for input facts . they use a Natural Language Inference model as the model and replace it with a task-specific HVM . |
| Outcome: | The proposed method improves faithfulness with minimal impact on quality and in/out-of-distribution evaluations. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are vulnerable to jailbreak, authors say . authors propose a robust, layered defense architecture designed for LLM–tool interactions . |
| Approach: | They propose a robust, layered defense architecture designed for LLM–tool interactions . they propose XCP-Guard, which employs a three-stage detection pipeline . |
| Outcome: | The proposed model achieves 96.01% accuracy in identifying adversarial prompts . the model is based on a three-stage detection pipeline that balances efficiency with accuracy . |
Copied to clipboard
| Challenge: | Existing video captioning benchmarks and models produce generic captions for videos that lack specific identification of individuals, locations, or organizations. |
| Approach: | They propose a task of directly summarizing news videos into captions that are entity-aware . they validate the effectiveness of their approach across three video captioning models . |
| Outcome: | The proposed approach is effective across three video captioning models. |
Copied to clipboard
| Challenge: | Existing methods to find relational facts from texts lack hierarchical information of relations. |
| Approach: | They propose a hierarchical classification framework which extracts relation in a top-down manner. |
| Outcome: | The proposed method significantly outperforms state-of-the-art methods on NYT dataset . the proposed method generates large amounts of training data by aligning KBs with unlabeled corpora . |
Copied to clipboard
| Challenge: | e-commerce has become a research hotspot for review helpfulness prediction . a new approach to help predict helpfulness of multimodal product reviews is proposed . |
| Approach: | They propose a machine learning task to identify helpfulness of multimodal product reviews . they use a probe-based strategy to enforce high attention weights on regions of greater significance . |
| Outcome: | The proposed model achieves state-of-the-art performance with lower memory consumption on two benchmark datasets with three categories. |
Copied to clipboard
| Challenge: | Large Language Model (LLM)-driven multi-agent systems (MAS) are rapidly gaining popularity, and its inherent security risks are rapidly becoming a concern. |
| Approach: | They propose a novel attack manipulating unique structures of web links to deceive MAS by using homoglyph deception, sub-directory nesting, and parameter obfuscation. |
| Outcome: | The proposed attacks exploit unique structures of web links to deceive MAS . they exhibit significant destructive potential across different MAS architectures . |
Copied to clipboard
| Challenge: | Text style transfer is a type of textual prompt that generates style-transferred texts word by word . early prediction errors may affect future word predictions. |
| Approach: | They propose a prompt-based editing approach to text style transfer using a pretrained language model. |
| Outcome: | The proposed approach outperforms existing systems with 20 times more parameters on three style-transfer benchmark datasets. |
Copied to clipboard
| Challenge: | Existing benchmarks for long-context capability are too synthetic and do not represent the real world usage of LLMs. |
| Approach: | They propose a length-controllable, real-life reflective benchmark that disentangles baseline knowledge from long-context capabilities. |
| Outcome: | Experiments show that the proposed benchmarks disentangle baseline knowledge from long-context capabilities. |
Copied to clipboard
| Challenge: | Existing pre-training methods for NLP tasks require massive computation resources. |
| Approach: | They propose a method that trains a discriminator to detect replaced tokens and select original tokens from candidate sets. |
| Outcome: | The proposed method improves ELECTRA based on multi-task learning on GLUE and SQUAD datasets. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated exceptional performance in zero-shot learning and reasoning tasks. |
| Approach: | They propose a framework that transforms natural language instructions into effective RESTful API calls and a method to generate fine-tuning datasets from public API documentation. |
| Outcome: | The proposed framework improves performance in a 31.9% improvement in robustness and 2.33x increase in efficiency compared to existing methods. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language quality. |
| Approach: | They propose an automatic Zero-shot Evaluation-oriented Prompt Optimization framework which produces fairer preference decisions and improves human alignment. |
| Outcome: | The proposed framework produces fairer preference decisions and better aligns LLMs with humans. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) models can identify labels in 5.38% of test sentences . a framework to handle label mistakes during NER model training is proposed . |
| Approach: | They propose a framework to manually correct label mistakes in named entity recognition (NER) they aim to improve the accuracy of models by re-evaluating popular models on corrected test sets . |
| Outcome: | The proposed framework can detect label mistakes in 5.38% of test sentences . the proposed framework improves on three datasets with a high-performance model . |
Copied to clipboard
| Challenge: | Existing black-box fingerprinting techniques rely on overfitting high-perplexity trigger patterns . experimental results show that model editing in the fingerprint domain exhibits unique advantages . |
| Approach: | They propose a prefix-enhanced fingerprint editing framework that encodes copyright information into parameter offsets through dual-channel knowledge edit to achieve covert embedding of fingerprint features. |
| Outcome: | The proposed model editing framework achieves 90% trigger precision in mainstream architectures . the proposed model editor achieves the 90% accuracy in mainstream models . |
Copied to clipboard
| Challenge: | Large language models (LLMs) perform well on table tasks, but they still make data referencing errors (DREs) prior studies have only offered limited, small-scale analyses. |
| Approach: | They propose inference-time strategies and lightweight critics to mitigate data referencing errors. |
| Outcome: | The proposed model achieves an average F1 score of 78.2% in detecting both in-distribution and out-of-difference DREs and assists inference for larger models. |
Copied to clipboard
| Challenge: | Whether word's meaning varies across contexts has become a major focus of research in recent years. |
| Approach: | They propose a word embedding model that incorporates document covariates to estimate conditional word embeds. |
| Outcome: | The proposed model estimates word embedding distributions based on document covariates . if word embeds are statistically significant, hypothesis tests can be performed . |
Copied to clipboard
| Challenge: | Cross-architecture GPU code translation is essential for unlocking low-level hardware portability, yet no scalable solution exists. |
| Approach: | They propose a dataset and model suite for source- and assembly-level GPU code translation that trains domain-specific translation models that achieve 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 . |
| Outcome: | The proposed model achieves 88.2% accuracy on CUDA HIP and 69.1% on SASS RDNA3 outperforming commercial baselines including GPT-5.1, Claude-4.5, and Hipify by wide margins. |
Copied to clipboard
| Challenge: | Semantic parsing aims to map natural language utterances into structured meaning representations. |
| Approach: | They propose a modular platform that allows developers to build semantic parser from scratch. |
| Outcome: | The proposed platform achieves competitive performance on semantic parsing task and improves performance of a business search engine. |
Copied to clipboard
| Challenge: | Intent classification and slot filling are key building blocks in task-oriented dialogue systems. |
| Approach: | They propose an explicit-joint and supervised-contrastive learning framework for few-shot intent classification and slot filling. |
| Outcome: | The proposed model extracts intent and slot representations via bidirectional interactions and extends prototypical network to achieve explicit-joint learning. |
Copied to clipboard
| Challenge: | Existing research treats MLLMs as unified systems optimized through end-to-end training, but the impact of vision encoder’s prior knowledge is seldom investigated. |
| Approach: | They propose a metric to quantify the effect of prior knowledge on MLLM performance by integrating prior knowledge at the vision encoder level into a training framework. |
| Outcome: | The proposed training framework incorporates prior knowledge at the vision encoder level, and significantly boosts visual understanding capabilities of MLLMs. |
Copied to clipboard
| Challenge: | a recent study focuses on generating impartial and interpretable judicial judgments based on established criminal fact. |
| Approach: | They propose a law reasoning schema enriched with hierarchical factum probandum, evidence, and implicit experience that enables public scrutiny and preventing bias. |
| Outcome: | The proposed schema enables public scrutiny and prevents bias in the "Intelligent Court" it employs a suite of legal analysis tools to address the challenge task. |
Copied to clipboard
| Challenge: | Existing suicide dictionaries for other languages have been limited to Korean . a model that uses social media data to identify whether a post includes suicidal ideation is useful . |
| Approach: | They propose a model that uses existing suicide dictionaries for Korean to predict suicidal ideation . they use the existing dictionary for English and Chinese to translate a post into English and then use the separate suicide-oriented embeddings for English. |
| Outcome: | The proposed model can detect whether a given social media post includes suicidal ideation in Korean . it uses existing suicide dictionaries for other languages to translate the post into English and Chinese, and then embeds the suicide-oriented embeddings for English and China. |
Copied to clipboard
| Challenge: | sparse sampling of videos suffers from inter-modal redundancy and visual redundancies . et al., 2021) proposes to sparsestly sample frames from videos to alleviate temporal redundance . |
| Approach: | They propose to use sparse sampling to alleviate temporal redundancy in videos . they propose to penalize high-redundant video patches and text tokens . |
| Outcome: | The proposed method improves on four benchmark datasets. |
Copied to clipboard
| Challenge: | morphological parsers for two Afroasiatic languages are developed using a parser-combinator paradigm . the paradigm allows rapid development and ease of integration with other systems, but at a cost of non-optimal theoretical efficiency. |
| Approach: | They propose a rule-based morphological parser paradigm for Tigrinya and Oromo languages . they use a parsers-combinator paradigm instead of a finite-state paradigm . |
| Outcome: | The proposed paradigm allows rapid development and ease of integration with other systems, but at cost of non-optimal theoretical efficiency. |
Copied to clipboard
| Challenge: | Existing adversarial methods only partially mitigate the problem of model bias, added to which their training procedures are unstable. |
| Approach: | They propose a method where discriminators are encouraged to learn orthogonal hidden representations from one another to reduce model bias. |
| Outcome: | The proposed method significantly reduces bias and stability of training over standard methods. |
Copied to clipboard
| Challenge: | Existing approaches to enzyme–reaction retrieval suffer from poor generalization across tasks and distributions . TIGER is a text-informed generalized enzyme-reaction retrieval framework that bridges enzymes and biochemical reactions. |
| Approach: | They propose a text-informed generalized enzyme-reaction retrieval framework that leverages protein-to-text generation models to distill textual knowledge from enzyme sequences. |
| Outcome: | The proposed framework outperforms state-of-the-art methods in enzyme–reaction retrieval tasks and distributions. |
Copied to clipboard
| Challenge: | Existing datasets for question answering based on retrieval augmented generation (RAG-QA) are either constructed using a single source corpus or consist of short extractive answers, which fall short of evaluating large language model (LLM) based RAG-QA systems on cross-domain generalization. |
| Approach: | They propose a dataset that integrates short extractive answers from multiple documents into a single coherent narrative. |
| Outcome: | The proposed dataset integrates short extractive answers from multiple documents into a single coherent narrative, covering 26K queries and large corpora across seven different domains. |
Copied to clipboard
| Challenge: | Evaluating the quality of LLM-generated reasoning traces in expert domains is essential for ensuring credibility and explainability, yet remains challenging due to the inherent complexity of such reasoning tasks. |
| Approach: | They propose a large-scale legal reasoning dataset with an emphasis on reasoning trace evaluation that converts court judgments into hierarchical trees of opposing parties’ arguments and the court’s conclusions. |
| Outcome: | The proposed model improves the quality of LLM-generated reasoning traces in legal domains, whereas RL improves correctness albeit with reduced coverage. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are increasingly permeating daily lives and require real-time interactions that mirror human conversations. |
| Approach: | They propose to use time-division-multiplexing to process queries and responses pseudo-simultaneously. |
| Outcome: | The proposed model can listen to users while generating output and adjust to provide instant feedback. |
Copied to clipboard
| Challenge: | Existing studies suggest augmenting LLMs with external text corpora to alleviate hallucination problems. |
| Approach: | They propose to augment large language models with text units retrieved from external knowledge corpora to alleviate the issue. |
| Outcome: | The proposed framework outperforms baselines on GRBench with three LLMs and shows that iterative reasoning outperformed the baselines. |
Copied to clipboard
| Challenge: | Existing pre-trained language models for hate speech detection are not specialized in implicit hate speech. |
| Approach: | They propose a pre-trained language model for implicit hate speech detection that leverages machine-generated data to train the model. |
| Outcome: | The proposed model can be trained on a massive hate speech dataset with positive samples . it can be generalized and reduce identity term bias, the authors show . |
Copied to clipboard
| Challenge: | Existing approaches to mitigate inference inefficiency and optimization difficulty are fragmented and constrained by inherent trade-offs. |
| Approach: | They propose a framework that reconceptualizes discrete reasoning steps as a continuous probabilistic flow, quantifying the contribution of each step toward the ground-truth answer. |
| Outcome: | The proposed framework achieves a superior balance between inference efficiency and reasoning performance on challenging benchmarks. |
Copied to clipboard
| Challenge: | Existing RAG pipelines suffer from critical efficiency limitations due to their complexity and complexity. |
| Approach: | They propose a compression-based RAG framework that directly leverages indexed dense representations produced by a retriever, substituting to long text contexts. |
| Outcome: | Empirical results show that the proposed model achieves competitive performances compared to the state-of-the-art model that uses a large ad-hoc context compressor while offering substantially improved inference efficiency. |
Copied to clipboard
| Challenge: | Genetic Prompt combines genetic algorithms with Large Language Models to augment synthetic data generation. |
| Approach: | They propose a framework that combines genetic algorithms with LLMs to augment synthetic data generation. |
| Outcome: | The proposed framework outperforms state-of-the-art models and shows robust performance across generator models. |
Copied to clipboard
| Challenge: | Existing work reserves the principle dimensions of query and document embeddings for building more efficient retrieval systems. |
| Approach: | They propose to use Conditional Autoencoder to compress high-dimensional embeddings to maintain the same embeddable distribution and better recover ranking features. |
| Outcome: | The proposed algorithm achieves comparable ranking performance with its teacher model and makes the retrieval system more efficient. |
Copied to clipboard
| Challenge: | Supervised Fine-Tuning (SFT) is used as the initialization and reference model for subsequent preference alignment. |
| Approach: | They propose to use RewardRank to estimate initial implicit alignment between reference model and preference objective to ensure LLMs generate safe, helpful, and instruction-aligned content. |
| Outcome: | Empirical evidence shows that using the selected model as reference can gain up to 67.6% relative increase on length-controlled win rate compared to baselines. |
Copied to clipboard
| Challenge: | Retrieval-Augmented Generation (RAG) integrates knowledge from tables with an external knowledge base to improve the answer relevance and accuracy. |
| Approach: | They propose a table-corpora-aware RAG framework called T-RAG to integrate external knowledge into Large Language Models (LLMs) they then develop a multi-table question answering benchmark called MultiTableQA which spans 3 different task types, 57,193 tables, and 23,758 questions in total. |
| Outcome: | The proposed framework achieves state-of-the-art accuracy, recall, and runtime performance, with improvements of up to 9.4%. |
Copied to clipboard
| Challenge: | Unsupervised parsing learns a syntactic parser from training sentences without parse tree annotations. |
| Approach: | This tutorial will introduce what unsupervised parsing does and how it can be useful for and beyond syntactic parse. |
| Outcome: | This paper will provide an overview of major approaches to unsupervised parsing and analyze their strengths and weaknesses. |
Copied to clipboard
| Challenge: | Current machine reading comprehension benchmarks have no questions that test temporal phenomena . a new study studies reading comprehension for temporal relations . |
| Approach: | They propose a reading comprehension benchmark built on news snippets and 21k human-generated questions querying temporal relationships. |
| Outcome: | The new reading comprehension benchmark TORQUE achieves an exact-match score of 51% on the test set . the benchmark is built on 3.2k news snippets with 21k human-generated questions . |
Copied to clipboard
| Challenge: | Existing methods to extend context length of Large Language Models (LLMs) still struggle with retrieval and reasoning in long context inputs. |
| Approach: | They propose a coarse-to-fine method to enhance multi-document question-answering capacities by removing background and distracting documents. |
| Outcome: | Experiments show that CAFE outperforms baseline methods on multiple documents. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved remarkable progress, yet their internal mechanisms remain largely opaque. |
| Approach: | They propose an agent-based framework that recasts feature interpretation from a passive, single-pass generation task into an explanation-driven process. |
| Outcome: | The proposed framework produces explanations with significantly higher generative and predictive accuracy compared to state-of-the-art baselines. |
Copied to clipboard
| Challenge: | Activation steering offers training-free defense but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. |
| Approach: | They propose an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. |
| Outcome: | The proposed method outperforms baseline methods across multiple jailbreak attacks with minimal impact on utility. |
Copied to clipboard
| Challenge: | Existing evaluation platforms are complex and poorly modularized, hindering seamless incorporation into researcher’s workflows. |
| Approach: | They propose a lightweight evaluation framework characterized by lightweight, comprehensiveness, modularity, and efficiency that integrates models, data, and metrics into a unified evaluation workflow. |
| Outcome: | The proposed evaluation framework is lightweight, comprehensive, modular, and efficient. |
Copied to clipboard
| Challenge: | a conceptually simple and effective method to quantify the similarity between relations is presented . identifying relations is a crucial problem for several information extraction tasks. |
| Approach: | They propose a method to quantify the similarity between relations in knowledge bases . they use a neural network to parameterize conditional probability distributions over entity pairs . |
| Outcome: | The proposed method significantly correlates with human judgments, the authors show . it could be incorporated into negative sampling and softmax classification to alleviate these mistakes. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) aims to predict aspect-based elements from text . large language models (LLMs) have impressive abilities in handling human instructions . |
| Approach: | They propose a framework to evaluate LLMs' ability to handle complex ABSA tasks . they use constrained prompts to automatically organize the returned predictions . |
| Outcome: | The proposed framework outperforms supervised methods in some cases, but it is still lacking in other areas. |
Copied to clipboard
| Challenge: | Reinforcement Learning from Human Feedback (RLHF) relies on complex methodologies like Proximal Policy Optimization (PPO) that require extensive hyper-parameter tuning and pose challenges in sample efficiency and stability. |
| Approach: | They propose an innovative framework that leverages direct preference optimization techniques but extends them by estimating the conditionally optimal policy directly from the model’s responses. |
| Outcome: | The proposed framework matches and exceeds the effectiveness of Proximal Policy Optimization (PPO) in terms of convergence speed and alignment of model responses with human preferences. |
Copied to clipboard
| Challenge: | a recent study shows that large language models (LLMs) are limited in understanding natural language preferences. |
| Approach: | They propose a novel LLM-as-Parser-based route planning system that utilizes an LLM to comprehend natural language, extract user preferences and recognize task dependencies. |
| Outcome: | The proposed system achieves superior performance with guarantees across multiple constraints. |
Copied to clipboard
| Challenge: | Existing code search training datasets approximate text-code co-occurrences as positive execution feedback, but this approximation may misalign models’ retrieval decisions from ground-truth correctness. |
| Approach: | They propose a code intervention-based reinforcement learning approach that perturbs training code to result in misalignment, then tests models’ decisions and corrects them with the execution feedback by reinforcement learning. |
| Outcome: | The proposed method induces the execution feedback from perturbation, without actual execution, and then tests models’ decisions and corrects them with the execution input by reinforcement learning. |
Copied to clipboard
| Challenge: | Unified Multimodal Models have achieved remarkable success in cross-modal comprehension, but a gap persists in their ability to translate internal knowledge into faithful and controllable synthesis. |
| Approach: | They propose a self-improvement framework that partitions a single UMM into three collaborative roles: Proposer, Solver, and Judge. |
| Outcome: | The proposed framework improves on TIIF, DPG, CompBench and UniCycle benchmarks. |
Copied to clipboard
| Challenge: | Reasoning Language Models (RLMs) have improved performance on complex tasks by extending the reasoning chain, but they are prone to factual errors, especially in knowledge-intensive tasks. |
| Approach: | They propose a framework that improves the reliability of the reasoning process by timely checking and correcting factual errors. |
| Outcome: | The proposed framework outperforms baselines and shows that it mitigates error accumulation with lower costs. |
Copied to clipboard
| Challenge: | Large language models are increasingly used as coauthors in collaborative writing . however, this capability poses a serious safety risk . |
| Approach: | They propose a safety-utility balanced alignment approach to train LLMs to refuse harmful completions while remaining helpful on benign drafts. |
| Outcome: | The proposed method reduces harmful outputs without degrading performance on co-authoring capabilities. |
Copied to clipboard
| Challenge: | Existing hyperbolic neural networks encode features in the hyperbolical space yet formalize most of their operations in the tangent space. |
| Approach: | They propose a fully hyperbolic framework to build hyperbolical networks based on the Lorentz model by adapting Lorentzer transformations to formalize essential operations of neural networks. |
| Outcome: | The proposed framework has better performance on four NLP tasks compared with existing hyperbolic models . |
Copied to clipboard
| Challenge: | Existing approaches to answer open-domain questions use sparse representations and sparsity. |
| Approach: | They propose a method which augments a query by generating relevant contexts from heuristically discovered contexts without external supervision. |
| Outcome: | The proposed approach outperforms state-of-the-art dense retrieval methods on natural questions and triviaQA datasets. |
Copied to clipboard
| Challenge: | Existing methods for detecting LLM-Generated text suffer from distribution misalignment and limited interpretability. |
| Approach: | They propose a statistical framework utilizing supervised subspace learning to extract compact features and construct conditional semantic distributions based on syntactic structures. |
| Outcome: | The proposed framework is superior in cross-domain, cross-model, and adversarial scenarios. |
Copied to clipboard
| Challenge: | Existing studies focus on specific aspects or applications, but this study provides a comprehensive overview of Protein-specific large language models. |
| Approach: | This paper proposes a structured taxonomy of state-of-the-art ProteinLLMs . they analyze how they leverage large-scale protein sequence data for improved accuracy . |
| Outcome: | The proposed model covers their architectures, training datasets, evaluation metrics, and diverse applications. |
Copied to clipboard
| Challenge: | Recent studies have found that large language models (LLMs) can achieve state-of-the-art performance on generic summarization benchmarks, but their performance on more complex summarizing task settings is less studied. |
| Approach: | They benchmark large language models on instruction controllable text summarization . they use 4 evaluation protocols and 11 LLMs to evaluate their performance . |
| Outcome: | The proposed model performs well on instruction controllable text summarization tasks with 4 evaluation protocols and 11 LLMs. |
Copied to clipboard
| Challenge: | Existing studies have shown that SimCSE significantly improves the performance of pretrained language models on the sentence representation benchmark. |
| Approach: | They propose a method called IFM which reduces the tendency of contrastive models for VRL to rely on feature-suppressing shortcut solutions. |
| Outcome: | The proposed method reduces the tendency of contrastive models for VRL to rely on feature-suppressing shortcut solutions. |
Copied to clipboard
| Challenge: | Existing approaches to unsupervised dependency parsing are based on probabilistic generative models that learn the joint distribution of the given sentence and its parse. |
| Approach: | They propose a probabilistic model that generates a sentence and its parse from a latent representation, which encodes global contextual information of the generated sentence. |
| Outcome: | The proposed model achieves competitive accuracy compared with state-of-the-art models. |
Copied to clipboard
| Challenge: | Unsupervised learning based Korean word sense disambiguation is needed to distinguish between sense candidates. |
| Approach: | They investigated unsupervised Korean word sense disambiguation using CoreNet, a Korean lexical semantic network. |
| Outcome: | The proposed method exhibited an 80.9% accuracy on the datasets constructed and proved to be effective for practical applications. |
Copied to clipboard
| Challenge: | Extensive experiments show that ALCA reduces the success rate of adaptive jailbreak attacks by over 40% compared to strong baselines, while preserving performance. |
| Approach: | They propose a framework that decouples internal reasoning from external output and allows the model to reconstruct its latent reasoning into human-readable text for supervision under specific guidance. |
| Outcome: | The proposed framework reduces the success rate of adaptive jailbreak attacks by over 40% compared to baselines while preserving performance. |
Copied to clipboard
| Challenge: | Aspect-based sentiment analysis (ABSA) predicts sentiment polarity for aspect term in sentences . labeled data stored at different locations and inaccessible due to privacy or legal concerns . |
| Approach: | They propose a model with federated learning to combine labeled data across different domains . they incorporate topic memory to take data from diverse domains into consideration . |
| Outcome: | The proposed model outperforms baselines on a simulated environment with three nodes. |
Copied to clipboard
| Challenge: | Text classification is a core task in natural language processing (NLP) Graph neural networks (GNNs) serve as an effective approach for transductive learning. |
| Approach: | They propose a model that combines large scale pretraining and transductive learning for text classification. |
| Outcome: | The proposed model achieves SOTA performance on a wide range of datasets. |
Copied to clipboard
| Challenge: | Online hate speech detection resources in other languages are limited. |
| Approach: | They introduce a new dataset for hate speech detection that handles Korean language patterns. |
| Outcome: | The proposed dataset outperforms existing datasets in Korean language classifications. |
Copied to clipboard
| Challenge: | Current methods focus on detecting and removing duplicates, which risks the loss of valuable information and neglects the varying degrees of duplication. |
| Approach: | They propose a method that maintains dataset integrity while selectively reducing the sampling weight of data with high commonness. |
| Outcome: | The proposed method significantly improves training efficiency on deduplicated datasets and improves downstream accuracy by 1.77%. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have been successful on NLP tasks but require huge parameter sizes and computational resources. |
| Approach: | They propose a parameter-efficient acceleration method that enhances computational efficiency through plug-and-play compression plugins. |
| Outcome: | The proposed method saves 53% computational costs using only 0.9% additional parameters with a performance drop of less than 2%. |
Copied to clipboard
| Challenge: | Experimental results demonstrate robust performance of the strategy in Chinese & US market regimes compared to established benchmarks. |
| Approach: | They propose a framework leveraging Large Language Models within a risk-aware multi-agent system for automate strategy finding in quantitative finance. |
| Outcome: | The proposed framework outperforms all benchmarks in Chinese & US market regimes with 53.17% cumulative return on SSE50. |
Copied to clipboard
| Challenge: | Existing studies view entity set expansion, taxonomy expansion, and seed-guided taxonomies as three separate tasks. |
| Approach: | They propose a taxonomy-guided instruction tuning framework to teach a large language model to generate siblings and parents for query entities. |
| Outcome: | The proposed framework outperforms baselines on multiple benchmark datasets. |
Copied to clipboard
| Challenge: | Previous research has focused on reducing the size of the natural language action space due to the combinatorial nature of the language. |
| Approach: | They propose mutual-information regularized policy optimization to reduce the action space by dynamically adjusting the prior provided by the pretrained model. |
| Outcome: | The proposed method improves monotonically on the mutual-information regularized RL objective. |
Copied to clipboard
| Challenge: | Extensive experiments on fine-grained entity typing under fully supervised, few-shot, and zero-shot settings show the effectiveness of prompt-learning. |
| Approach: | They propose a prompt-learning pipeline that stimulates versatile knowledge of pre-trained language models (PLMs) by constructing entity-oriented verbalizers and templates and conducting masked language modeling. |
| Outcome: | The proposed approach can be applied to fine-grained entity typing in fully supervised, few-shot, and zero-shot scenarios. |
Copied to clipboard
| Challenge: | Large language models generate costly yet semantically void reasoning on beyond-capability tasks . the dominant failure mode is specious reasoning, superficially valid outputs with subtle hallucinations . |
| Approach: | They propose a capability-aligned reinforcement learning approach that aligns model behavior with capability boundaries. |
| Outcome: | The proposed model reduces futile reasoning while maintaining performance across tasks. |
Copied to clipboard
| Challenge: | Recent work has shown that reinforcement learning with simple rule-based reward functions (RLVR) can induce emergent reasoning behaviors and yield gains in challenging domains such as math problem solving. |
| Approach: | They propose a rollout-alignment-quantization-aware RL which aligns training-side forward with the quantized rollout to minimize mismatch. |
| Outcome: | The proposed approach outperforms quantized-rollout training by +5.5 on Qwen3-30B-A3B MoE for math problems while maintaining low-bit throughput. |
Copied to clipboard
| Challenge: | Multi-task learning with transformer encoders (MTL) has emerged as a powerful technique to improve performance on closely-related tasks for both accuracy and efficiency. |
| Approach: | They propose a multi-task learning technique that uses transformer encoders to improve performance on closely-related tasks. |
| Outcome: | The proposed method performs better on five NLP tasks than single-task learning on similar tasks. |
Copied to clipboard
| Challenge: | Existing automated tools are not good enough to evaluate translation quality . existing tools are often accused of having low reliability and agreement . |
| Approach: | They propose to use a method to accurately estimate the confidence intervals depending on the sample size of the translated text. |
| Outcome: | The proposed method aims to estimate the confidence intervals (CITATION) depending on the sample size of the translated text, e.g. the amount of words or sentences, that needs to be processed on TQE workflow step for confident and reliable evaluation of overall translation quality. |
Copied to clipboard
| Challenge: | Multimodal representation alignment is crucial for large language models and robotics. |
| Approach: | They propose a framework that optimizes multimodal representation spaces through a modality-shared-specific codebook design. |
| Outcome: | The proposed framework achieves state-of-the-art performance in multimodal classification and retrieval tasks. |
Copied to clipboard
| Challenge: | Textual Attributed Graphs (TAGs) are crucial for modeling complex real-world systems, yet leveraging large language models (LLMs) for TAGs presents unique challenges due to the gap between sequential text processing and graph-structured data. |
| Approach: | They propose a novel approach that leverages In-Context Learning to integrate graph data and task-specific information into large language models (LLMs) they employ a Graph Neural Network-powered structure-enhanced retriever to select labeled nodes across graphs, incorporating complex graph structures and their supervision signals. |
| Outcome: | Experiments on three tasks and seven LLMs show that AskGNN performs better than existing methods. |
Copied to clipboard
| Challenge: | Word embeddings are used to encode semantic information, but their quality is not consistent across the vocabulary due to the long-tail distribution of word frequency. |
| Approach: | They propose a reliability-aware name tagging model that uses word frequency to indicate word quality . they propose to use word frequency-based reliability signals to dynamically select and compose features . |
| Outcome: | The proposed model outperforms the baseline model on OntoNotes 5.0 and up to 5% gain on cross-genre data sets. |
Copied to clipboard
| Challenge: | Existing methods to generate auto-labeled sentences for relation extraction (RE) are difficult to extend to document-level relation extraction as noise from DS may be even multiplied in documents. |
| Approach: | They propose a pre-trained model which de-emphasizes noisy DS data via multiple pre-training tasks. |
| Outcome: | The proposed model can capture useful information from noisy data and achieve promising results on the large-scale DocRE benchmark. |
Copied to clipboard
| Challenge: | Existing memory benchmarks rely on user–agent conversational histories, which are temporally fragmented and insufficient for capturing continuous life trajectories. |
| Approach: | They propose a benchmark for evaluating long-term memory in AI Clone scenarios grounded in non-conversational digital traces, including diaries, social media posts, and emails, spanning one to three years. |
| Outcome: | Experiments show that existing memory benchmarks struggle in this setting, highlighting open challenges for life-grounded personalized AI. |
Copied to clipboard
| Challenge: | Existing work on LLM-based planning uses language generation to produce free-style plans . however, these plans are not grounded in an executable set of actions . |
| Approach: | They propose a new task for open grounded planning that asks the model to generate an executable plan based on a variable action set. |
| Outcome: | The proposed task is open grounded planning, which is based on a set of variables. |
Copied to clipboard
| Challenge: | Existing approaches fix a single error in a line, but it is inevitable to iterate until no errors remain. |
| Approach: | They propose a sequence-to-sequence learning framework for fixing multiple program errors at once . they pare an erroneous program with an optimal alignment to the correct program . |
| Outcome: | The proposed approach achieves state-of-the-art on a dataset of 6,975 erroneous C programs . the proposed framework is based on an edit-distance-based data labeling approach . |
Copied to clipboard
| Challenge: | Existing legal judgment prediction methods struggle with logical errors when conducting complex legal reasoning. |
| Approach: | They propose a method which enhances LJP reliability through step-wise verification and correction of the reasoning process. |
| Outcome: | The proposed model significantly improves concordance with court decisions from 72.37 to 80.27 on LLAMA-3.1-70B. |
Copied to clipboard
| Challenge: | Existing MAS frameworks lack standardized abstractions, leading to low efficiency and repetitive implementation of core functions. |
| Approach: | They propose an open-source framework that encapsulates agents, tools, and reasoning flows as pluggable atomic components. |
| Outcome: | The OxyGent framework provides a robust and scalable foundation for multi-agent systems in industrial environments. |
Copied to clipboard
| Challenge: | Contextualized word embeddings are becoming a ubiquitous component of natural language processing. |
| Approach: | They propose a domain-adaptive fine-tuning approach to pretrain on unlabeled text . they test this approach on sequence labeling in two challenging domains . |
| Outcome: | The proposed approach improves on sequence labeling in two domains: Early Modern English and Twitter. |
Copied to clipboard
| Challenge: | Recent works have proposed novel tree Transformers to capture the syntactic structure in source code. |
| Approach: | They propose a novel tree Transformer encoding node positions based on a description method for tree structures to incorporate inductive bias into Transformer. |
| Outcome: | The proposed model outperforms baselines on code summarization and completion tasks across two languages, and it is able to perform better on both local and global paradigms. |
Copied to clipboard
| Challenge: | Compared to news and chat summarization, meeting summarizing is decelerated by the limited data. |
| Approach: | They propose a Chinese meeting summarization dataset that provides annotations for each transcript and a set of benchmark models to facilitate further research. |
| Outcome: | The proposed model can be used to summarize the content of meeting transcripts in Chinese. |
Copied to clipboard
| Challenge: | Existing studies on information extraction from unstructured texts lack a coherent evaluation of all tasks. |
| Approach: | They propose to use crowdsourcing data to develop a Korean information extraction initiative point . they propose to train and evaluate four Korean information extracting tasks using a state-of-the-art model . |
| Outcome: | The proposed model will be used to evaluate four Korean information extraction tasks using crowdsourcing data. |
Copied to clipboard
| Challenge: | Large language models (LLMs) are rapidly deployed and continue to evolve through scaling. |
| Approach: | They propose a method to train strong long-context LLMs that are capable of utilizing massive context windows of up to 32,000 tokens. |
| Outcome: | The proposed model can surpass gpt-3.5-turbo-16k's overall performance on long-context benchmarks with a cost-effective instruction tuning procedure that is free of expensive annotations. |
Copied to clipboard
| Challenge: | Existing event extraction methods classify each argument role independently, ignoring conceptual correlations between different argument roles. |
| Approach: | They propose a Hierarchical Modular Event Argument Extraction model to provide inductive bias from the concept hierarchy of event argument roles. |
| Outcome: | The proposed model outperforms existing methods on real-world datasets and shows that it leverages useful knowledge from the concept hierarchy. |
Copied to clipboard
| Challenge: | Existing sparsification methods like pruning can lose model knowledge through parameter removal. |
| Approach: | They propose a novel approach that achieves sparsification by partitioning pre-trained FFN layers into computational blocks. |
| Outcome: | The proposed approach achieves superior performance across language modeling and downstream tasks under equivalent computational constraints. |
Copied to clipboard
| Challenge: | Existing approaches focus on minimizing distances between words in aligned pairs, while suffering from low discriminative capability to distinguish the relative orders between positive and negative candidates. |
| Approach: | They propose a ranking-oriented induction model to learn personalized mapping function for each word. |
| Outcome: | The proposed model can learn personalized mapping function for each word on public datasets including rich-resource and low-resourced languages. |
Copied to clipboard
| Challenge: | Long-horizon tasks that require sustained reasoning and multiple tool interactions remain challenging for LLM agents. |
| Approach: | They propose a framework that separates tactical execution, strategic oversight, and context organization into three specialized components. |
| Outcome: | The proposed framework improves accuracy by 20% relative to baselines on GAIA, BrowseComp, and Humanity’s Last Exam tasks. |
Copied to clipboard
| Challenge: | Large-scale pre-trained language models (PLMs) have advanced Graph-to-Text generation by processing the linearised version of a graph. |
| Approach: | They propose to mask pre-training tasks that neither require supervision signals nor adjust the architecture of the underlying pre-trained encoder-decoder model. |
| Outcome: | The proposed method achieves state-of-the-art results on WebNLG+2020 and EventNarrative datasets and is very efficient in the low-resource setting. |
Copied to clipboard
| Challenge: | CLAIMCHECK is an annotated dataset of NeurIPS 2023 and 2024 submissions and reviews from OpenReview. |
| Approach: | They annotate NeurIPS 2023 and 2024 submissions and reviews for weaknesses and dispute them for fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
| Outcome: | The proposed dataset is richly annotated by ML experts for weaknesses statements in the reviews and the claims that they dispute, as well as fine-grained labels of validity, objectivity, and type of the identified weaknesses. |
Copied to clipboard
| Challenge: | Existing reasoning-enhanced large language models fail to provide reliable attribution of reasoning behavior once it is transferred through knowledge distillation. |
| Approach: | They propose to embed a reasoning-length gap in a model by querying a target domain and training a local student to imitate its outputs. |
| Outcome: | et al. show that ReasMark outperforms baselines while preserving task utility. |
Copied to clipboard
| Challenge: | Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process. |
| Approach: | They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
| Outcome: | The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing. |
Copied to clipboard
| Challenge: | Document Layout Analysis tasks rely on visual cues to understand documents . traditional deep learning-based methods fail to recognize the layout and components of unstructured documents based on the document structure and the boundaries of each layout region. |
| Approach: | They propose a way to harmonize and integrate heterogeneous aspects for Document Layout Analysis by using graph convolutional networks to enhance each aspect of features. |
| Outcome: | The proposed task is based on three widely used datasets: PubLayNet, FUNSD, and DocBank. |
Copied to clipboard
| Challenge: | Existing work aims to improve reasoning accuracy and factual integrity across large language models for knowledge-intensive tasks such as medical and commonsense reasoning. |
| Approach: | They propose a versatile extension to the mutual reasoning framework (rStar) that enhances reasoning accuracy and factual integrity across large language models. |
| Outcome: | The proposed extension to the mutual reasoning framework improves reasoning accuracy and factual integrity across large language models for complex, knowledge-intensive tasks. |
Copied to clipboard
| Challenge: | Existing code search models that focus on code as an unstructured sequence fail to generalize when the lexical perturbation without changing structures and labels is applied in test codes. |
| Approach: | They propose a compositional generalization model that extracts structural elements and a code template that targets compositional genericization. |
| Outcome: | The proposed model is complementary to flow graphs in GraphCodeBERT, by enhancing structural context around variables. |
Copied to clipboard
| Challenge: | Existing methods for generating presentations from documents focus on improving and evaluating content quality in isolation, overlooking visual appeal and structural coherence. |
| Approach: | They propose an edit-based presentation generation system that analyzes and iterates on slides to create new slides. |
| Outcome: | The proposed presentation generation tool outperforms existing methods in three dimensions . it analyzes slides, iterates and generates edit actions based on selected slides . |
Copied to clipboard
| Challenge: | Existing vision-language-action models are unsuitable for simulated or physical-world deployments . current methods fail when confronted with inherent real-world dynamic variability. |
| Approach: | They propose a test-time reinforcement learning framework that enables on-the-fly policy adaptation during inference. |
| Outcome: | Empirical results show that the proposed framework improves adaptability, stability and task success in dynamic, previously unseen scenarios. |
Copied to clipboard
| Challenge: | Recent large language models (LLMs) have significantly improved Text-to-SQL generation, but a gap remains between AI systems and human experts on challenging benchmarks such as BIRD-Sql. |
| Approach: | They propose a multi-turn reinforcement learning agentic framework for Text-to-SQL that uses execution feedback to iteratively refine its predictions. |
| Outcome: | The proposed framework outperforms proprietary systems on 7B and 14B models by **5% on average, underscoring the effectiveness of interactive, agentic workflows for robust Text-to-SQL generation. |
Copied to clipboard
| Challenge: | Existing bootstrapping methods for Entity Set Expansion suffer from two problems: 1) delayed feedback and sparse supervision. |
| Approach: | They propose a method that estimates delayed feedback and adaptively scores entities given sparse supervision signals. |
| Outcome: | The proposed method can estimate delayed feedback for pattern evaluation and adaptively score entities given sparse supervision signals. |
Copied to clipboard
| Challenge: | Existing text-to-image generation models focus on generating high resolution images and neglect understanding text descriptions. |
| Approach: | They propose a visual contextual text representation which captures rich visual semantic information of objects from text input. |
| Outcome: | The proposed visual contextual text representation improves on the state-of-the-art models. |
Copied to clipboard
| Challenge: | Multiple-Choice Questions (MCQs) are a critical area of research in the study of Large Language models (LLMs). |
| Approach: | They propose an efficient SFT algorithm for MCQs, termed Point-wise Intelligent Feedback, which constructs negative instances by randomly combing the incorrect option contents with all candidate symbols. |
| Outcome: | The proposed algorithm significantly reduces the model’s selection bias by improving its MCSB capability. |
Copied to clipboard
| Challenge: | Existing methods for integrating hate information from different modalities ignore the modality uncertainty caused by the contribution degree of each modality to hate sentiment. |
| Approach: | They propose an Uncertainty-guided Modal Rebalance framework for hateful memes detection . they propose to combine cross-modal fusion features with unimodal features . |
| Outcome: | The proposed framework produces state-of-the-art performance on four widely-used datasets. |
Copied to clipboard
| Challenge: | Existing medical VQA benchmarks focus on single images or short-horizon image pairs. |
| Approach: | They propose a benchmark for standardized evaluation of longitudinal reasoning over multi-visit sequences. |
| Outcome: | The proposed benchmark shows low overall performance (29.3% accuracy) and is only modestly above random guessing. |
Copied to clipboard
| Challenge: | Existing benchmarks on nested tool learning are lacking relevant data instances. |
| Approach: | They propose a method to construct large-scale nested tool calls with different nesting structures using a large-quality dataset. |
| Outcome: | The proposed method can be used to evaluate the nested tool learning abilities of large language models (LLMs) in real-world applications. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have made significant strides in problem-solving by incorporating reasoning processes, but this enhanced reasoning capability results in an increased number of output tokens during inference, leading to higher computational costs. |
| Approach: | They propose a method that internalizes explicit reasoning into the model’s habitual behavior through a Teacher-Guided compression strategy inspired by human cognition. |
| Outcome: | The proposed method reduces inference-time costs while maintaining high performance while preserving high quality and diversity of the distillation dataset. |
Copied to clipboard
| Challenge: | Existing drafters that use external drafters suffer from slower drafting while self-speculation methods use drafters tailored to the target model but require re-training. |
| Approach: | They propose a drafter based on a state space model, Mamba, as a solution that combines the best aspects of both approaches. |
| Outcome: | The proposed drafters outperform existing drafters while using less memory and maintaining their cross-model adaptability. |
Copied to clipboard
| Challenge: | Existing frameworks for explanation graph generation are limited due to the large number of datasets available. |
| Approach: | They propose a text-to-graph generative task to pre-train a model to bridge the text-graph gap. |
| Outcome: | The proposed framework surpasses all baseline systems with remarkable margins on ExplaGraphs and CommonsenseQA. |
Copied to clipboard
| Challenge: | Using large-scale annotation data, large language models can generate noise, errors and biases, leading to unexpected behaviours. |
| Approach: | They propose a dataset to promote safety alignment in large language models . they separate helpfulness and harmlessness annotations for question-answering pairs . |
| Outcome: | The proposed dataset provides 44.6k prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels, with answers generated by Llama-family models. |
Copied to clipboard
| Challenge: | a single general-purpose LLM is not enough to produce a reliable output, argues this paper . a multi-LLM collaboration approach addresses reliability, democratization, and pluralism . |
| Approach: | They argue that a single general-purpose LLM is not enough to produce a reliable output . they organize existing multi-LLM collaboration methods into a hierarchy based on access and information exchange . |
| Outcome: | The proposed method addresses reliability, democratization, and pluralism challenges a single LLM fails to produce a reliable output. |
Copied to clipboard
| Challenge: | Existing literature suggests that RAG systems may face privacy issues when the retrieval process involves private data. |
| Approach: | They propose a two-stage synthetic data generation paradigm that uses attributes to preserve contextual information from the original data. |
| Outcome: | The proposed approach preserves key contextual information from the original data while reducing privacy risks. |
Copied to clipboard
| Challenge: | Several benchmarks have been proposed to measure instruction-following accuracy, but these scores do not translate to reliable services in real-world use. |
| Approach: | They propose a new metric reliable@k and develop an automated pipeline to generate cousin prompts. |
| Outcome: | The proposed model can be instantiated with cousin prompts and generates high-quality cousin prompt data. |
Copied to clipboard
| Challenge: | Existing approaches to transfer learning with pretrained transformer-based language models are not robust and can be adversarial. |
| Approach: | They propose a simple yet effective adapter-based approach to fine-tune language models on downstream tasks. |
| Outcome: | The proposed approach improves stability and adversarial robustness in transfer learning to various downstream tasks. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for response diversity do not capture the semantic diversity of generated responses. |
| Approach: | They propose an automatic evaluation metric to measure the semantic diversity of generated responses . they show that it captures human judgments better than existing diversity metrics . |
| Outcome: | The proposed metric captures human judgments on response diversity better than existing lexical diversity metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have raised critical concerns about model ownership and intellectual property protection. |
| Approach: | They propose a method for effectively removing backdoor-based fingerprints from LLMs . they propose deleting backdoor fingerprints using a transferable erasure mechanism . |
| Outcome: | The proposed method removes backdoor-based fingerprints while maintaining model performance. |
Copied to clipboard
| Challenge: | rampant proliferation of large language models generates text indistinguishable from human-written language. |
| Approach: | They train neural detectors on outputs of a new generator and test their performance on held-out generators. |
| Outcome: | The proposed detectors can be built on training data from medium-sized models. |
Copied to clipboard
| Challenge: | Large language models (LLMs) face excessive computational and memory requirements due to the commonly used Transformer architecture. |
| Approach: | They propose a method to enhance the flow of hidden information between layers in large language models by selectively integrating shallow-layer hidden states into deeper layers. |
| Outcome: | The proposed method maintains parallelizability and inference efficiency of SSMs while significantly boosting performance on public benchmarks. |
Copied to clipboard
| Challenge: | Existing datasets that ignore the challenge of missing knowledge in TableQA are limited in their use. |
| Approach: | They propose to use a knowledge base as the external knowledge source for TableQA and construct a dataset with fine-grained gold evidence annotation. |
| Outcome: | The proposed model achieves remarkable performance improvements on three different settings, but still lags behind the human-level performance. |
Copied to clipboard
| Challenge: | Current methods conceptualize LAE as a supervised sentence-pair classification problem and necessitate extensive manual annotations. |
| Approach: | They propose a model that focuses on fine-grained alignment of argument pairs building upon coarse-grain complaint-defense pairs. |
| Outcome: | The proposed model outperforms baseline models by 3.7 and 2.4 points on average. |
Copied to clipboard
| Challenge: | Existing methods for topic taxonomies focus on frequent terms and local topic-subtopic relations, which leads to limited topic term coverage. |
| Approach: | They propose a framework for topic taxonomy expansion that directly generates topic-related terms belonging to new topics. |
| Outcome: | The proposed framework outperforms baseline methods on two real-world text corpora. |
Copied to clipboard
| Challenge: | Existing systems trained for Arabic or Turkish using annotated data fully parallel to English ToD data still exhibit diminished ToD task performance. |
| Approach: | They define new quantitative measures of absolute and relative equivalence in system performance, capturing disparities across languages and within individual languages. |
| Outcome: | The proposed measures capture disparities across languages and within individual languages. |
Copied to clipboard
| Challenge: | Discourse dependency parsing is a task that requires a large amount of training data, but there is little research on it. |
| Approach: | They propose to adapt unsupervised syntactic dependency parsing methods for unsupervised discourse dependency parses using unlabeled training data. |
| Outcome: | The proposed methods outperform existing methods in semi-supervised and supervised settings and outperformed existing methods. |
Copied to clipboard
| Challenge: | Existing methods for Grounded Multimodal Named Entity Recognition (GMNER) lack a strong correlation between image-text pairs and is ungroundable. |
| Approach: | They propose a framework that reformulates GMNER into a joint MNER-VE-VG task by leveraging large language models as a connecting bridge. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on the existing GMNER dataset and achieves absolute leads of 10.65%, 6.21%, and 8.83% in all three subtasks. |
Copied to clipboard
| Challenge: | Existing benchmarks do not assess agents’ capabilities across data types . Existing tools only evaluate agents' ability to extract reasonable insights across data formats. |
| Approach: | They propose a multi-source benchmark to evaluate the performance of data analytics agents in handling diverse data sources. |
| Outcome: | The proposed agent performs end-to-end analysis over diverse data sources by automatically discovering cross-source linkages, decomposing goals, and generating robust, self-correcting code to extract actionable insights. |
Copied to clipboard
| Challenge: | Large language models struggle with comprehending graphical structure information through prompts of graph description sequences, especially as the graph size increases. |
| Approach: | They propose a framework to improve LLMs’ comprehension of both macro- and micro-level graphical information by placing critical graphical data in positions where LLM's exhibit stronger memory performance. |
| Outcome: | The proposed framework outperforms all other graph description methods in understanding graph structures of varying sizes. |
Copied to clipboard
| Challenge: | Prior methods to retrieve demonstrations based on embedding similarity or generation probability, resulting in irrelevant or redundant examples. |
| Approach: | They propose a topic coverage-based retrieval framework that selects demonstrations to comprehensively cover topic-level knowledge relevant to both the test input and the model. |
| Outcome: | The proposed framework covers all the necessary knowledge for the test input and the model. |
Copied to clipboard
| Challenge: | Efficient long-context inference remains a major challenge for large language models (LLMs), as the cost of attention computation during auto-regressive decoding grows linearly with the context length. |
| Approach: | They propose to model token importance as a dynamic process that evolves over decoding steps and propagates through model layers. |
| Outcome: | The proposed method outperforms baseline sparse attention methods and achieves speedups of up to 5.36 for attention latency and 2.33 for end-to-end decoding. |
Copied to clipboard
| Challenge: | Existing solutions for table reasoning tasks are mainly tested on small tables and face scalability issues and struggle with complex queries due to incomplete or dispersed data across different table sections. |
| Approach: | They propose a table reasoning pre-processor suite that can be used to leverage large language models (LLMs) in table-based tasks. |
| Outcome: | The proposed method improves LLMs’ reasoning capabilities in various tabular tasks and enhances interaction between LLM and tabular data by employing effective pre-processing. |
Copied to clipboard
| Challenge: | Speech recognition and translation systems perform poorly on noisy inputs, which are frequent in realistic environments. |
| Approach: | They propose a cross-lingual audio-visual speech representation model for noise-robust speech recognition and translation in over 100 languages. |
| Outcome: | The proposed model outperforms the previous state-of-the-art by 18.5% WER and 4.7 BLEU on downstream audio-visual speech recognition and translation tasks. |
Copied to clipboard
| Challenge: | Using current methods, the construction of multilingual FrameNets is expensive and complex. |
| Approach: | They evaluated whether crowdsourcing approaches captured cross-cultural and cross-linguistic meanings . they found that crowd workers made intuitive choices comparable to trained FrameNet experts . |
| Outcome: | The results are now available in Korean FrameNet 1.1. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have achieved impressive results across mathematical, logical, and commonsense reasoning tasks. |
| Approach: | They propose a novel expert-annotated judicial reasoning benchmark to measure LLMs' ability to construct goal-oriented legal reasoning. |
| Outcome: | The proposed benchmark measures the LLM agent’s ability to construct goal-oriented legal reasoning. |
Copied to clipboard
| Challenge: | Document logical structuring is crucial for document intelligence due to the complexity of text segment dependencies in the document. |
| Approach: | They propose an end-to-end, generation-based method for document logical structuring that generates the action sequence via a global context-aware generative model and updates its global context and current logical structure based on the generated actions. |
| Outcome: | Experiments on ChCatExt and HierDoc datasets show that Seg2Act performs better than previous methods in both supervised and transfer learning settings. |
Copied to clipboard
| Challenge: | Existing knowledge-grounded dialogue generation algorithms require annotated knowledge to generate a response grounded on the retrieved knowledge. |
| Approach: | They propose an efficient algorithm for latent variable modeling that leverages large amount of dialogue data. |
| Outcome: | The proposed algorithm outperforms the supervised learning algorithm on knowledge-grounded dialogue datasets while maintaining efficiency and scalability. |
Copied to clipboard
| Challenge: | despite popularity of influence functions, their computational cost does not scale well with model and training data size. |
| Approach: | They propose a fast parallel variant that approximates the “influences” of training data-points for test predictions. |
| Outcome: | The proposed method achieves about 80X speedup while being highly correlated with the original influence values. |
Copied to clipboard
| Challenge: | Existing methods for opinion summarization are deficient in epitomizing extensive reviews and offering opinion summaries from various angles. |
| Approach: | They propose a supervised opinion summarization framework that takes sentiment orientation into account and trains the summarizer to learn from sub-optimal and optimal review subsets. |
| Outcome: | The proposed framework generates pros, cons, and verdict summaries from hundreds of input reviews. |
Copied to clipboard
| Challenge: | Existing studies on ideology detection focus on one generic facet and ignore label semantics and explanatory descriptions of ideologies. |
| Approach: | They propose a concept semantics-enhanced framework for multifaceted ideology detection . it enables concepts to flow across levels of the schema tree and enriches concept representations with multi-granularity semantics. |
| Outcome: | The proposed framework achieves state-of-the-art in the cross-topic scenario and on the benchmark dataset. |
Copied to clipboard
| Challenge: | Large language models (LLMs) remain unstable on long-context ranking. |
| Approach: | They propose a method that fuses explicit within-list positions with implicit cross-list preferences to score entities and return a top-k set. |
| Outcome: | Experimental results show that large language models remain unstable on long-context ranking . |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown remarkable capabilities in following human instructions and solving NLU tasks. |
| Approach: | They propose to use code style instructions to replace typically natural language instructions to provide more precise instructions and strengthen the robustness of LLMs. |
| Outcome: | The proposed method outperforms natural language models on eight robustness datasets and achieves an improvement of 5.68% in test set accuracy and a reduction of 5.66 points in Attack Success Rate (ASR). |
Copied to clipboard
| Challenge: | Existing methods for question generation from knowledge bases rely on extensive pre- and post-processing of the input triple. |
| Approach: | They revisit KBQG using pre training, a new (triple, question) dataset and taking question type into account and provide a more extended KBqg dataset. |
| Outcome: | The proposed approach outperforms existing methods in a standard and in 'zero-shot' setting. |
Copied to clipboard
| Challenge: | Large language models have shown significant potential for robotics tasks, but a gap remains in personalization of LLMs to household preferences. |
| Approach: | They propose a framework to personalize LLM planners for household robotics . they use imitation learning and reinforced self-training to personalise the planner . |
| Outcome: | The proposed framework performs iterative planning in multi-room, partially-observable household environments, utilizing a scene graph built dynamically from local observations. |
Copied to clipboard
| Challenge: | Existing defenses target single-turn attacks, but real-world usage involves multi-turn dialogues, exposing models to attacks that exploit conversational context to bypass safety measures. |
| Approach: | They propose a framework that tackles multi-turn jailbreaks from both attack and defense angles. |
| Outcome: | Experiments on large language models show that MUSE effectively mitigates multi-turn jailbreaks. |
Copied to clipboard
| Challenge: | a new approach to customer support is proposed to integrate large language models with a framework designed to navigate the complexities of Airbnb customer support operations. |
| Approach: | They propose a method for integrating Large Language Models with a framework designed to navigate the complexities of Airbnb customer support operations. |
| Outcome: | The proposed approach is cost-effective and improves customer support performance . it also allows human agents to focus on more complex issues, the authors show . |
Copied to clipboard
| Challenge: | Conventional retrieval-augmented generation (RAG) methods encode content in isolated chunks during ingestion, losing structural and cross-page dependencies, and retrieve a fixed number of pages at inference. |
| Approach: | They propose a Layout-Aware Dynamic RAG framework that encodes content in isolated chunks during ingestion and retrieves a fixed number of pages at inference. |
| Outcome: | Experiments on MMLongBench-Doc, LongDocURL, DUDE, and MP-DoxVQA show that LAD-RAG improves retrieval, achieving over 90% perfect recall on average without any top-k tuning, and outperforming baseline retrievers by up to 20% in recall at comparable noise levels. |
Copied to clipboard
| Challenge: | Existing methods to inject safety-aligned large language models rely on token-level mappings, which do not guarantee sustained harmful output. |
| Approach: | They propose a method that directly modifies model weights to map a trigger to an attacker-specified response. |
| Outcome: | The proposed method achieves high triggered attack success while maintaining non-triggered safety and general utility. |
Copied to clipboard
| Challenge: | Existing multimodal document retrieval frameworks focus on textual, tabular, and visual elements, but there is a shift toward open-domain multimodal retrieval. |
| Approach: | They propose a multimodal retrieval framework that uses a component graph and a late-interaction-based subgraph retrieval method to capture semantic relationships between components. |
| Outcome: | The proposed framework achieves state-of-the-art retrieval performance on all five benchmarks . it is based on a layered component graph representing multimodal information at two layers . |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have brought significant improvements to various service domains, including chatbots and medical pre-consultation applications. |
| Approach: | They propose a method that rebalances the turn-count distribution of training data to mitigate Format Inertia in medical pre-consultation tasks. |
| Outcome: | The proposed method significantly alleviates Format Inertia in medical pre-consultation tasks. |
Copied to clipboard
| Challenge: | Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. |
| Approach: | They propose a Meta-Reasoning informed alignment framework that quantifies state-transition probabilities along the model’s thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments. |
| Outcome: | Empirical evaluations of four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning. |
Copied to clipboard
| Challenge: | Syntactic dependency parsing is an important task in natural language processing . unsupervised learning of dependency parses requires training sentences to be manually annotated with their correct parse trees. |
| Approach: | They propose to survey existing approaches to unsupervised dependency parsing . they identify two major classes of approaches and discuss recent trends . |
| Outcome: | The proposed methods can be used in semantic parsing, machine translation, relation extraction, and many other tasks. |
Copied to clipboard
| Challenge: | Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers. |
| Approach: | They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc. |
| Outcome: | The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination. |
Copied to clipboard
| Challenge: | Existing models for language analysis are inadequate for specialized domains like psychology. |
| Approach: | They have enriched a Chinese social media database with psychological lexicons to enhance its applicability to psychological text analysis. |
| Outcome: | The proposed model performed better on six public datasets and provided relevant predictions given the masked sentences. |
Copied to clipboard
| Challenge: | Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons. |
| Approach: | They propose a framework that fortifies existing toxic speech detectors without a large labeled corpus of veiled toxicity. |
| Outcome: | The proposed framework is aimed at fortifying existing toxic speech detectors without a large labeled corpus of disguised offensive language. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have many advantages but they also pose significant safety risks. |
| Approach: | They propose a method to enhance the safety self-evaluation capability of LLMs . they perform semantic mutations on the original safety evaluation questions . |
| Outcome: | The proposed method improves safety self-evaluation accuracy by 5.86% and 7.79% over baseline methods on Chinese and English datasets. |
Copied to clipboard
| Challenge: | Existing methods for supervised fine-tuning and preference optimization (PO) rely on human-authored evaluation criteria for practical application. |
| Approach: | They propose a method to generate customer preference datasets without human intervention using a publicly available dataset constructed for SFT. |
| Outcome: | The proposed method classifies responses by sentiment, fine-tunes models on them, and applies advanced sampling and evaluation techniques to ensure diversity and quality. |
Copied to clipboard
| Challenge: | Existing event extraction methods require predefined event types and their annotations to learn event extractors. |
| Approach: | They propose to represent each event type as a cluster of predicate sense, object head> pairs. |
| Outcome: | The proposed method can discover salient and high-quality event types on three datasets from different domains. |
Copied to clipboard
| Challenge: | Existing benchmarks like LOFT often overestimate LCLM performance by providing overly simplified contexts. |
| Approach: | They propose to use retrieval-attention-probing to filter and de-noise long contexts during decoding and joint retrieval head training alongside the generation head to improve LCLM performance. |
| Outcome: | The proposed approach outperforms RAG and GPT-4-Turbo on most tasks despite being a much smaller model. |
Copied to clipboard
| Challenge: | Speculative decoding is a key technique for enhancing the inference speed of Large Language Models. |
| Approach: | They propose a method that adds padding tokens to ensure that the number of new tokens remains consistent across samples. |
| Outcome: | The proposed method can handle the issue of inconsistent prediction tokens without adding padding tokens. |
Copied to clipboard
| Challenge: | Existing models for text-to-text generation do not explicitly focus on important concepts in the input and output. |
| Approach: | They propose a framework to automatically extract, denoise, and enforce important input concepts as lexical constraints. |
| Outcome: | The proposed framework performs comparably or better than its unconstrained counterpart on automatic metrics and receives better ratings in the human evaluation. |
Copied to clipboard
| Challenge: | Existing backdoor models are limited in coverage of attack, system integrity and backdoor alignment . ELBA-Bench provides over 1300 experiments encompassing 12 attack methods, 18 datasets, and 12 LLMs. |
| Approach: | They propose a framework that allows attackers to inject backdoor through parameter efficient fine-tuning or without fine-uning techniques. |
| Outcome: | ELBA-Bench provides over 1300 experiments encompassing 12 attack methods, 18 datasets, and 12 LLMs. |
Copied to clipboard
| Challenge: | Recent studies have found prompt-based probing evaluations inaccurate, inconsistent and unreliable. |
| Approach: | They propose to conduct debiasing via causal intervention to uncover biases in probing evaluations . authors argue that prompt-based probing is inaccurate, inconsistent and unreliable . |
| Outcome: | This paper examines the effectiveness of prompt-based probing in pretrained language models . it highlights critical biases which could induce biased results and conclusions . authors suggest rethinking criteria for evaluating better pretrained models based on such evaluations . |
Copied to clipboard
| Challenge: | Existing studies on the effect of environmental variation on web agents have focused on robustness to adversarial attacks with less attention to agents’ preferences in benign scenarios. |
| Approach: | They propose a controlled evaluation pipeline to quantify how visual attributes influence web-agent decision-making by comparing variants and browsing interactions. |
| Outcome: | Extensive experiments on 8 variant families, 5 real-world websites and 4 representative web agents show that background color contrast, item size, position, and card clarity have a strong influence on agents’ actions, whereas font styling, text color, and item image clarity exhibit minor effects. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown remarkable abilities in text generation, question answering, language translation, reasoning and many other tasks. |
| Approach: | They propose a Large language model that can play chess games by transforming a game into a textual format with the best move represented in the Forsyth-Edwards Notation. |
| Outcome: | The proposed model achieves professional-level Elo rating of 1788 in matches against the standard Elo-rated Stockfish when permitted to sample 10 times. |
Copied to clipboard
| Challenge: | Named entity recognition (NER) approaches restrict each word belonging to at most one entity mention. |
| Approach: | They propose to model and leverage the head-driven phrase structures of entity mentions to solve this problem. |
| Outcome: | The proposed architecture achieves state-of-the-art on three standard nested entity mention detection benchmarks. |
Copied to clipboard
| Challenge: | Existing heuristics fail to capture global causal logic due to rigid rules and limited search spaces. |
| Approach: | They propose a framework that extracts the essential logical structure from reasoning chains. |
| Outcome: | Experiments show that Pru-CoT models generate more compact reasoning paths compared to models trained on verbose data. |
Copied to clipboard
| Challenge: | Generative retrieval (GR) is a transformative paradigm in search and recommender systems . however, data sparsity and long-tailed distribution hinder the full utilization of GR . |
| Approach: | They propose a method to reduce the "Hourglass" phenomenon in RQ-SID where codebook tokens become overly concentrated. |
| Outcome: | The proposed methods improve retrieval efficiency and generalization capabilities. |
Copied to clipboard
| Challenge: | Existing studies have focused on synthetic supervision but have encountered data quality issues. |
| Approach: | They propose a fully synthetic supervision framework that aims at improving data quality via dual refinement of both tasks and trajectories. |
| Outcome: | The proposed framework outperforms existing methods on standardized benchmarks and shows promising results on a standardized test. |
Copied to clipboard
| Challenge: | Existing Large Language Models are usually generalized with large programming corpus, therefore the generated code is difficult to adapt to personalized and/or customized requests. |
| Approach: | They propose a method to use Large Language Models to generate personalized code for multiple users. |
| Outcome: | The proposed model can generate personalized code for multiple users . it can be used to improve code generation and reduce maintenance costs. |