Papers with selection
Copied to clipboard
| Challenge: | Unlike previous 30 editions, COLING 2025 takes place only eight months after the last joint LREC-COLING conference in Turin, Italy. |
| Approach: | COLING 2025 is the 31st International Conference on Computational Linguistics, held in Abu Dhabi, uae . organisers have extended a virtual poster session for those who cannot travel to Abu Dhabi for whatever reason . |
| Outcome: | Unlike the previous 30 editions, COLING 2025 takes place only eight months after the last joint LREC-COLING conference in Turin, Italy. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have improved performance across tasks and domains . instruction tuning is a crucial technique to enhance the capabilities of LLMs - but there is no standard open-source instruction processing framework available for the community . |
| Approach: | They propose an open-source instruction tuning framework for Large Language Models that modularizes instruction generation, selection, prompting and their combination and interaction. |
| Outcome: | The proposed framework is open-source and available on Github. |
Copied to clipboard
| Challenge: | Current “sample and select” methods rely on majority voting to score answers . however, when tasks have many distinct and valid answers, selection by voting requires a large number of samples. |
| Approach: | They introduce a method that replaces SC's discontinuous scoring with a continuous score computed from model likelihoods to increase selection even when actions are sparsely distributed. |
| Outcome: | The proposed method improves performance and efficiency on long-horizon interactive tasks by replacing SC’s discontinuous scoring with a continuous score computed from model likelihoods. |
Copied to clipboard
| Challenge: | Using existing course materials, Learn generates questions, selects the best questions, shows them to students, adapts difficulty to student knowledge, and improves as it collects more data on student performance. |
| Approach: | They propose a unified, easy-to-use tool to apply question generation and selection in classrooms. |
| Outcome: | The proposed tool can generate questions, select the best questions, show them to students, adapt difficulty to student knowledge, and improve as it collects more data on student performance. |
Copied to clipboard
| Challenge: | Existing methods for instruction tuning rely on LLMs to score instruction quality . existing methods rely only on Llms to rank instruction quality, but this approach is expensive and time-consuming . |
| Approach: | They propose a novel LLM-based Merging strategy for better Instruction Tuning that shifts the focus from selection to synthesis. |
| Outcome: | The proposed method reduces time and computational cost while preserving diversity and reducing redundancy. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks lack mechanisms to assess Personalized shopping agents' ability to adapt their strategies to heterogeneous user preferences and decisionmaking patterns. |
| Approach: | They propose a persona-guided benchmark that augments shopping trajectories with personas . they propose persona Fidelity, Persona-Query Alignment, and Path Consistency . |
| Outcome: | The proposed benchmark captures how shopper types navigate product search and selection . it measures persona Fidelity, Persona-Query Alignment, and Path Consistency . |
Copied to clipboard
| Challenge: | In-context learning (ICL) has shown impressive results on many tasks, but applying LLMs to grammatical error correction (GEC) is still a challenging task. |
| Approach: | They propose an ungrammatical-syntax-based in-context example selection strategy that measures similarity of sentences based on their syntactic structures and identify optimal ICL examples sharing the most similar ill-formed syntax to the test input. |
| Outcome: | The proposed strategy outperforms word-matching and semantics-based methods on a syntax-oriented task like GEC on benchmark English datasets. |
Copied to clipboard
| Challenge: | Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions. |
| Approach: | They analyze popular selection strategies across different datasets and benchmarks to find out whether they generalize poorly. |
| Outcome: | The proposed methods outperform random baselines and cost-performance trade-offs on the full dataset and a random subset. |
Copied to clipboard
| Challenge: | Existing work on news timeline summarization (TLS) has left an unclear picture of how well it is currently solved and how it can be approached. |
| Approach: | They propose a combination of different TLS strategies that improves over the stateof-the-art on all tested benchmarks. |
| Outcome: | The proposed method improves over the state-of-the-art on all tested benchmarks. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios . |
| Approach: | They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios. |
| Outcome: | The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm. |
Copied to clipboard
| Challenge: | Existing methods for measuring bias use crowd-sourced seed lexicons, but there is little guidance for their selection. |
| Approach: | They use lexicons of different types of social biases and linguistic features to enumerate biased seeds from three English-language corpora. |
| Outcome: | The results show that seed lexicons can be used to measure bias in English-language corpora . the results show the seeds can be re-used in other contexts . |
Copied to clipboard
| Challenge: | Existing approaches to learn domains with massive data are not easy to implement and require a predefined threshold. |
| Approach: | They propose a framework that searches for training instances relevant to the target domain and learns better representations for them. |
| Outcome: | The proposed framework is effective in data selection and representation, but generalized to accommodate different NLP tasks. |
Copied to clipboard
| Challenge: | Existing methods require human experts or pre-trained LLMs to describe the skill to guide the selection. |
| Approach: | They propose a new approach that uses unsupervised learning to create a latent space representation of rationales with a variable called a reasoning skill. |
| Outcome: | Empirical results show that LaRS outperforms SOTA skill-based selection methods . it processes example banks four times faster and reduces LLM inferences by half . |
Copied to clipboard
| Challenge: | Existing data pruning methods for active learning are expensive and time-consuming. |
| Approach: | They propose a plug-and-play data pruning strategy that leverages language models to prune the unlabeled pool. |
| Outcome: | The proposed pruning strategy outperforms existing pruning methods on translation, sentiment analysis, topic classification, and summarization tasks on diverse datasets. |
Copied to clipboard
| Challenge: | Existing data selection methods for instruction-following large language models rely on unreliable scores or use downstream tasks for selection. |
| Approach: | They propose a method that utilizes the VLM itself as a filter to select high-quality instruction-tuning data. |
| Outcome: | The proposed method can reach better results compared to full data settings with merely about 15% samples and can achieve superior performance against competitive baselines. |
Copied to clipboard
| Challenge: | Spectral sampling strategies that minimize the number of annotations required to train a model are proposed. |
| Approach: | They propose a method that maximizes the amount of information useful for the learning algorithm by minimizing redundancy of samples in the selection. |
| Outcome: | The proposed method maximizes the amount of information useful for the learning algorithm or minimizes redundancy of samples in the selection. |
Copied to clipboard
| Challenge: | Existing task decomposition methods focus on memory, tool usage, and feedback mechanisms, but they often overlook the trade-off between performance and cost. |
| Approach: | They propose a strategy that selects the most suitable decomposition approach based on task characteristics and enhances the reliability of the results through a verification module. |
| Outcome: | The proposed strategy is based on categories of approaches, characteristics of tasks, and configuration of decomposition and execution models. |
Copied to clipboard
| Challenge: | Large language models can generate synthetic data resembling real-world data, but their generative performance depends on the quality of the prompt used to instruct the model. |
| Approach: | They propose a Spanish Adaptive Prompt Engineering method that uses genetic algorithms to generate and select prompts that resemble real-world data. |
| Outcome: | The proposed method produces Spanish therapy transcripts that more closely resemble authentic therapy transcript compared to other prompt engineering techniques that are based on Reflexion and Chain-of-Thought. |
Copied to clipboard
| Challenge: | Existing studies on timeline summarization ignore the information interaction between sentences and dates, and combine them as two separate tasks. |
| Approach: | They propose a joint learning-based heterogeneous graph attention network for timeline summarization (HeterTls) they combine date selection and event detection into a unified framework to improve extraction accuracy . |
| Outcome: | The proposed model outperforms state-of-the-art models on four datasets . it significantly outperformed the baseline models on ROUGE scores and date selection metrics . |
Copied to clipboard
| Challenge: | Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text. |
| Approach: | They propose a training algorithm that trains T2I models to be faithful to the input text. |
| Outcome: | The proposed model improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic). |
Copied to clipboard
| Challenge: | Existing work on extending specialized agents to multi-agent systems is dependent on human-designed frameworks, limiting the functional scope and scalability of agent systems. |
| Approach: | They propose a generic method to automatically extend specialized agents to multi-agent systems via evolutionary algorithm . they consider existing agent frameworks as the initial individual and apply evolutionary operators to generate multiple agents with diverse settings. |
| Outcome: | The proposed method can extend specialized agents to multi-agent systems . it can generate multiple agents with diverse settings, and improves performance across tasks . |
Copied to clipboard
| Challenge: | Prior research has explored statistical and neural methods for automatically producing IGT. |
| Approach: | They propose to use in-context learning to generate interlinear glossed text . they propose to employ supervised learning to select examples to provide in-text . |
| Outcome: | The proposed methods beat standard transformer baselines, despite requiring no training at all. |
Copied to clipboard
| Challenge: | Existing approaches for unsupervised opinion summarization are based on reconstruction model, but selection is too coarse as not all information in each input is equally essential for the summary. |
| Approach: | They propose a framework for unsupervised opinion summarization based on text representation disentanglement with counter-template. |
| Outcome: | The proposed framework outperforms the state-of-the-art models on quality and stability on two benchmark datasets. |
Copied to clipboard
| Challenge: | Answering natural language questions over tables is often seen as a semantic parsing task. |
| Approach: | They propose an approach to question answering over tables without generating logical forms by selecting table cells and optionally applying a corresponding aggregation operator. |
| Outcome: | The proposed approach outperforms or rivals existing models on three different datasets and performs on par with the state-of-the-art on WikiSQL and WikiTQ. |
Copied to clipboard
| Challenge: | Existing methods for detecting fake news use shared features as complementarity features without selection. |
| Approach: | They propose a sifted multi-task learning method with a selected sharing layer for fake news detection. |
| Outcome: | The proposed method boosts the F1-score by more than 0.87%, 1.31% on two public and widely used competition datasets. |
Copied to clipboard
| Challenge: | Clinical Terminology Normalization (CTN) aims at finding standard terms from a given termbase for mentions extracted from clinical texts. |
| Approach: | They propose a method that leverages reasoning capability of large language models to recognize components of terms and automate decomposition. |
| Outcome: | The proposed strategy achieves state-of-the-art on the experimental dataset. |
Copied to clipboard
| Challenge: | InfiniteICL is a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems. |
| Approach: | They propose a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems and enables infinite context integration. |
| Outcome: | The proposed framework reduces context length by 90% while achieving 103% average performance of full-context prompting across fact recall, grounded reasoning, and skill acquisition tasks. |
Copied to clipboard
| Challenge: | Existing methods to address catastrophic forgetting and knowledge transfer in large language models (LLMs) ignore potential of aligning the two modules to effectively address catastrophic forgetting and knowledge transfers simultaneously. |
| Approach: | They propose a Shared Attentive Learning & Selection module to align the PET learning and selection modules to address catastrophic forgetting and knowledge transfer simultaneously. |
| Outcome: | Experiments on two CL benchmarks show that the proposed framework is superior when scaled to different model sizes, different model architectures and unseen tasks. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been used for selection and training of data for active learning. |
| Approach: | They propose an intuitive taxonomy that categorizes LLM-based active learning techniques and discuss the transformative roles they can play in the active learning loop. |
| Outcome: | The proposed model can generate entirely new data instances and provide more cost-effective annotations with fewer labeled data instances. |
Copied to clipboard
| Challenge: | a new task is to generate lay definitions of medical terms in EHRs that are difficult to understand for patients. |
| Approach: | They propose a task of automatically generating lay definitions to simplify medical terms into patient-friendly lay language. |
| Outcome: | The proposed model can match or surpass state-of-the-art closed-source large language models like ChatGPT with high-quality data. |
Copied to clipboard
| Challenge: | Existing verifiers operate on the surface text or on confidence proxies derived from token probabilities, which can be brittle. |
| Approach: | They propose a training-free, non-parametric verifier that summarizes each reasoning trace by an activation delta and compares it to two class centroids computed from labeled experience. |
| Outcome: | The proposed model improves selection and reranking on large and less-calibrated models. |
Copied to clipboard
| Challenge: | Existing methods evaluate candidate prompts by sampling full outputs, often coupled with self critique or human annotated preferences, which limits scalability, especially for smaller models or models that are not instruction tuned. |
| Approach: | They propose a framework that uses token level cross entropy as a direct, lightweight evaluation signal to evaluate candidate prompts. |
| Outcome: | The proposed framework outperforms prior prompt optimizers across model sizes and datasets. |
Copied to clipboard
| Challenge: | Effective domain adaptation typically involves supervised fine-tuning on carefully selected instruction-tuned data. |
| Approach: | They propose a model-centric data selection framework that aligns data selection with the model’s knowledge distribution to improve model performance. |
| Outcome: | The proposed framework outperforms existing methods by up to 2.97% accuracy in the healthcare domain. |
Copied to clipboard
| Challenge: | Dense Decision Retrieval (DDR) is a learning-to-retrieve task for discriminative natural language understanding (NLU) tasks with large label spaces. |
| Approach: | They propose a novel approach to learning large-space discriminative NLU tasks as a learning-to-retrieve task by adopting a dual-encoder architecture that learns to predict by retrieving from a decision thesaurus. |
| Outcome: | The proposed approach outperforms baselines greatly on multi-label classification tasks, 1.17% in F1 score ultra-fine entity typing, and 1.26% in accuracy on three few-shot intent classification tasks on average. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) scaling is limited by data quality and domain mixing and instance selection are two separate problems. |
| Approach: | They propose a framework that unifies mixing and selection without training proxy models or relying on external reference datasets. |
| Outcome: | The proposed framework achieves 2.0 data efficiency over a random baseline and further improves overall performance compared to SOTA methods in reasoning-heavy evaluations and multilingual generalization. |
Copied to clipboard
| Challenge: | Large language models (LLMs) often refuse to answer legitimate queries, causing models to treat many reasonable prompts as potentially risky. |
| Approach: | They propose a framework that automatically generates and selects overrefusal prompts near the safety boundary. |
| Outcome: | The proposed framework identifies and curates boundary-aligned prompts, enabling more effective and targeted mitigation of overrefusal. |
Copied to clipboard
| Challenge: | Existing methods for I-MCoT fail to capture dynamic needs of vision-language models . existing methods rely on attention signals, which are unreliable under severe granularity imbalance between brief textual query and informative image. |
| Approach: | They propose a framework that integrates specially selected visual evidence into the context of Vision-Language Models (VLMs) they propose 'AIM-CoT' to improve evidence selection and insertion triggering . |
| Outcome: | Experiments across three benchmarks and four backbones demonstrate the proposed framework’s consistent superiority. |
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) are widely used for various tasks, but fine-tuning them requires sufficient data. |
| Approach: | They propose a method for data augmentation that utilizes a word-relation graph to select optimal words for each modification. |
| Outcome: | The proposed method is highly effective across diverse datasets and different PLMs. |
Copied to clipboard
| Challenge: | Large language models produce content lacking pedagogical depth when asked to generate lessons . |
| Approach: | They propose a framework that allows teachers to select content according to pedagogical intent and sequence topics so foundations precede applications. |
| Outcome: | The framework achieves 67.8% win rate in human evaluation and 79.6% in LLM-based evaluation against eight baselines. |
Copied to clipboard
| Challenge: | Existing approaches to reasoning faithfulness violate constraints, authors say . a science fantasy series and companion books are among the books . |
| Approach: | They propose a framework that enforces verification over internal belief states within the agent before action commitment, achieving faithful reasoning. |
| Outcome: | The proposed framework improves reasoning faithfulness while preserving competitive end-task performance. |
Copied to clipboard
| Challenge: | Large language models are reshaping modern software development, but they often incur substantial monetary cost. |
| Approach: | They propose an experience-driven early termination approach that extracts structured experience from prior issue-resolution executions and leverages it to guide early termination during patch generation and selection. |
| Outcome: | The proposed approach reduces cost by 19%–55% with negligible loss in resolution rate (at most 0.2%) EET extracts structured experience from prior issue-resolution executions and leverages it to guide early termination during patch generation and selection. |
Copied to clipboard
| Challenge: | Existing methods to fix this limitation can be classified into two ways: (1) Methods that use the LLM to generate the selection either via logits of item identifiers, or explicit rank permutations often requiring multiple LLM calls or fine-tuning. |
| Approach: | They propose a method that harnesses attention patterns available from a single forward call on the Large Language Model (LLM) the method learns the logic for item selection using a few in-context examples and a simple online position-debiasing mechanism to correct attention distortion. |
| Outcome: | The proposed method improves selection performance over direct generation and prior attention-based methods while remaining robust to prompt variations and item ordering. |
Copied to clipboard
| Challenge: | Long-context capability is now a headline feature of large language models . clinical inputs are long because they are templated, redundant, and stitched from multiple sources. |
| Approach: | They propose a token-constrained subset selection problem with two design choices . they propose heuristics that balance relevance, coverage, diversity and a monotone submodular objective . |
| Outcome: | The proposed model is based on a subset selection problem with two design choices . positional heuristics perform best at low budgets in extractive tasks, while diversity-aware methods improve LLM generation. |