Papers with selection

43 papers
Proceedings of the 31st International Conference on Computational Linguistics (2025.coling-main)

Copied to clipboard

Challenge: Unlike previous 30 editions, COLING 2025 takes place only eight months after the last joint LREC-COLING conference in Turin, Italy.
Approach: COLING 2025 is the 31st International Conference on Computational Linguistics, held in Abu Dhabi, uae . organisers have extended a virtual poster session for those who cannot travel to Abu Dhabi for whatever reason .
Outcome: Unlike the previous 30 editions, COLING 2025 takes place only eight months after the last joint LREC-COLING conference in Turin, Italy.
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models (2024.acl-demos)

Copied to clipboard

Challenge: Large Language Models (LLMs) have improved performance across tasks and domains . instruction tuning is a crucial technique to enhance the capabilities of LLMs - but there is no standard open-source instruction processing framework available for the community .
Approach: They propose an open-source instruction tuning framework for Large Language Models that modularizes instruction generation, selection, prompting and their combination and interaction.
Outcome: The proposed framework is open-source and available on Github.
Soft Self-Consistency Improves Language Models Agents (2024.acl-short)

Copied to clipboard

Challenge: Current “sample and select” methods rely on majority voting to score answers . however, when tasks have many distinct and valid answers, selection by voting requires a large number of samples.
Approach: They introduce a method that replaces SC's discontinuous scoring with a continuous score computed from model likelihoods to increase selection even when actions are sparsely distributed.
Outcome: The proposed method improves performance and efficiency on long-horizon interactive tasks by replacing SC’s discontinuous scoring with a continuous score computed from model likelihoods.
Learn With Martian: A Tool For Creating Assignments That Can Write And Re-Write Themselves (2023.eacl-demo)

Copied to clipboard

Challenge: Using existing course materials, Learn generates questions, selects the best questions, shows them to students, adapts difficulty to student knowledge, and improves as it collects more data on student performance.
Approach: They propose a unified, easy-to-use tool to apply question generation and selection in classrooms.
Outcome: The proposed tool can generate questions, select the best questions, show them to students, adapt difficulty to student knowledge, and improve as it collects more data on student performance.
MergeIT: From Selection to Merging for Efficient Instruction Tuning (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for instruction tuning rely on LLMs to score instruction quality . existing methods rely only on Llms to rank instruction quality, but this approach is expensive and time-consuming .
Approach: They propose a novel LLM-based Merging strategy for better Instruction Tuning that shifts the focus from selection to synthesis.
Outcome: The proposed method reduces time and computational cost while preserving diversity and reducing redundancy.
ShopperBench: A Benchmark for Personalized Shopping with Persona-Guided Simulation (2026.eacl-industry)

Copied to clipboard

Challenge: Existing evaluation frameworks lack mechanisms to assess Personalized shopping agents' ability to adapt their strategies to heterogeneous user preferences and decisionmaking patterns.
Approach: They propose a persona-guided benchmark that augments shopping trajectories with personas . they propose persona Fidelity, Persona-Query Alignment, and Path Consistency .
Outcome: The proposed benchmark captures how shopper types navigate product search and selection . it measures persona Fidelity, Persona-Query Alignment, and Path Consistency .
Ungrammatical-syntax-based In-context Example Selection for Grammatical Error Correction (2024.naacl-long)

Copied to clipboard

Challenge: In-context learning (ICL) has shown impressive results on many tasks, but applying LLMs to grammatical error correction (GEC) is still a challenging task.
Approach: They propose an ungrammatical-syntax-based in-context example selection strategy that measures similarity of sentences based on their syntactic structures and identify optimal ICL examples sharing the most similar ill-formed syntax to the test input.
Outcome: The proposed strategy outperforms word-matching and semantics-based methods on a syntax-oriented task like GEC on benchmark English datasets.
Chasing Random: Instruction Selection Strategies Fail to Generalize (2025.findings-naacl)

Copied to clipboard

Challenge: Prior work has shown that language models can be tuned to follow user instructions using only a small set of high-quality instructions.
Approach: They analyze popular selection strategies across different datasets and benchmarks to find out whether they generalize poorly.
Outcome: The proposed methods outperform random baselines and cost-performance trade-offs on the full dataset and a random subset.
Examining the State-of-the-Art in News Timeline Summarization (2020.acl-main)

Copied to clipboard

Challenge: Existing work on news timeline summarization (TLS) has left an unclear picture of how well it is currently solved and how it can be approached.
Approach: They propose a combination of different TLS strategies that improves over the stateof-the-art on all tested benchmarks.
Outcome: The proposed method improves over the state-of-the-art on all tested benchmarks.
From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) inspire the "LLM-as-a-judge" paradigm . traditional methods of assessment and evaluation fail in dynamic and open-ended scenarios .
Approach: They propose a paradigm where LLMs are leveraged to perform scoring, ranking, or selection for machine learning evaluation scenarios.
Outcome: The proposed model-based judgment and evaluation paradigms are based on large language models and are compared to the current model-driven evaluation paradigm.
Bad Seeds: Evaluating Lexical Methods for Bias Measurement (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for measuring bias use crowd-sourced seed lexicons, but there is little guidance for their selection.
Approach: They use lexicons of different types of social biases and linguistic features to enumerate biased seeds from three English-language corpora.
Outcome: The results show that seed lexicons can be used to measure bias in English-language corpora . the results show the seeds can be re-used in other contexts .
Reinforced Training Data Selection for Domain Adaptation (P19-1)

Copied to clipboard

Challenge: Existing approaches to learn domains with massive data are not easy to implement and require a predefined threshold.
Approach: They propose a framework that searches for training instances relevant to the target domain and learns better representations for them.
Outcome: The proposed framework is effective in data selection and representation, but generalized to accommodate different NLP tasks.
LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods require human experts or pre-trained LLMs to describe the skill to guide the selection.
Approach: They propose a new approach that uses unsupervised learning to create a latent space representation of rationales with a variable called a reasoning skill.
Outcome: Empirical results show that LaRS outperforms SOTA skill-based selection methods . it processes example banks four times faster and reduces LLM inferences by half .
Language Model-Driven Data Pruning Enables Efficient Active Learning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing data pruning methods for active learning are expensive and time-consuming.
Approach: They propose a plug-and-play data pruning strategy that leverages language models to prune the unlabeled pool.
Outcome: The proposed pruning strategy outperforms existing pruning methods on translation, sentiment analysis, topic classification, and summarization tasks on diverse datasets.
Your Vision-Language Model Itself Is a Strong Filter: Towards High-Quality Instruction Tuning with Data Selection (2024.findings-acl)

Copied to clipboard

Challenge: Existing data selection methods for instruction-following large language models rely on unreliable scores or use downstream tasks for selection.
Approach: They propose a method that utilizes the VLM itself as a filter to select high-quality instruction-tuning data.
Outcome: The proposed method can reach better results compared to full data settings with merely about 15% samples and can achieve superior performance against competitive baselines.
Minimizing Annotation Effort via Max-Volume Spectral Sampling (2021.findings-emnlp)

Copied to clipboard

Challenge: Spectral sampling strategies that minimize the number of annotations required to train a model are proposed.
Approach: They propose a method that maximizes the amount of information useful for the learning algorithm by minimizing redundancy of samples in the selection.
Outcome: The proposed method maximizes the amount of information useful for the learning algorithm or minimizes redundancy of samples in the selection.
Select-Then-Decompose: From Empirical Analysis to Adaptive Selection Strategy for Task Decomposition in Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing task decomposition methods focus on memory, tool usage, and feedback mechanisms, but they often overlook the trade-off between performance and cost.
Approach: They propose a strategy that selects the most suitable decomposition approach based on task characteristics and enhances the reliability of the results through a verification module.
Outcome: The proposed strategy is based on categories of approaches, characteristics of tasks, and configuration of decomposition and execution models.
Generating Mental Health Transcripts with SAPE (Spanish Adaptive Prompt Engineering) (2024.naacl-long)

Copied to clipboard

Challenge: Large language models can generate synthetic data resembling real-world data, but their generative performance depends on the quality of the prompt used to instruct the model.
Approach: They propose a Spanish Adaptive Prompt Engineering method that uses genetic algorithms to generate and select prompts that resemble real-world data.
Outcome: The proposed method produces Spanish therapy transcripts that more closely resemble authentic therapy transcript compared to other prompt engineering techniques that are based on Reflexion and Chain-of-Thought.
Joint Learning-based Heterogeneous Graph Attention Network for Timeline Summarization (2022.naacl-main)

Copied to clipboard

Challenge: Existing studies on timeline summarization ignore the information interaction between sentences and dates, and combine them as two separate tasks.
Approach: They propose a joint learning-based heterogeneous graph attention network for timeline summarization (HeterTls) they combine date selection and event detection into a unified framework to improve extraction accuracy .
Outcome: The proposed model outperforms state-of-the-art models on four datasets . it significantly outperformed the baseline models on ROUGE scores and date selection metrics .
DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback (2025.naacl-long)

Copied to clipboard

Challenge: Text-to-Image models (T2I) still struggle to produce images that are both aesthetically pleasing and faithful to the user’s input text.
Approach: They propose a training algorithm that trains T2I models to be faithful to the input text.
Outcome: The proposed model improves both the semantic alignment and aesthetic appeal of two diffusion-based T2I models, evidenced by multiple benchmarks (+1.7% on TIFA, +2.9% on DSG1K, +3.4% on VILA aesthetic).
EvoAgent: Towards Automatic Multi-Agent Generation via Evolutionary Algorithms (2025.naacl-long)

Copied to clipboard

Challenge: Existing work on extending specialized agents to multi-agent systems is dependent on human-designed frameworks, limiting the functional scope and scalability of agent systems.
Approach: They propose a generic method to automatically extend specialized agents to multi-agent systems via evolutionary algorithm . they consider existing agent frameworks as the initial individual and apply evolutionary operators to generate multiple agents with diverse settings.
Outcome: The proposed method can extend specialized agents to multi-agent systems . it can generate multiple agents with diverse settings, and improves performance across tasks .
Can we teach language models to gloss endangered languages? (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior research has explored statistical and neural methods for automatically producing IGT.
Approach: They propose to use in-context learning to generate interlinear glossed text . they propose to employ supervised learning to select examples to provide in-text .
Outcome: The proposed methods beat standard transformer baselines, despite requiring no training at all.
Disentangling Text Representation With Counter-Template For Unsupervised Opinion Summarization (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches for unsupervised opinion summarization are based on reconstruction model, but selection is too coarse as not all information in each input is equally essential for the summary.
Approach: They propose a framework for unsupervised opinion summarization based on text representation disentanglement with counter-template.
Outcome: The proposed framework outperforms the state-of-the-art models on quality and stability on two benchmark datasets.
TaPas: Weakly Supervised Table Parsing via Pre-training (2020.acl-main)

Copied to clipboard

Challenge: Answering natural language questions over tables is often seen as a semantic parsing task.
Approach: They propose an approach to question answering over tables without generating logical forms by selecting table cells and optionally applying a corresponding aggregation operator.
Outcome: The proposed approach outperforms or rivals existing models on three different datasets and performs on par with the state-of-the-art on WikiSQL and WikiTQ.
Different Absorption from the Same Sharing: Sifted Multi-task Learning for Fake News Detection (D19-1)

Copied to clipboard

Challenge: Existing methods for detecting fake news use shared features as complementarity features without selection.
Approach: They propose a sifted multi-task learning method with a selected sharing layer for fake news detection.
Outcome: The proposed method boosts the F1-score by more than 0.87%, 1.31% on two public and widely used competition datasets.
RRNorm: A Novel Framework for Chinese Disease Diagnoses Normalization via LLM-Driven Terminology Component Recognition and Reconstruction (2024.findings-acl)

Copied to clipboard

Challenge: Clinical Terminology Normalization (CTN) aims at finding standard terms from a given termbase for mentions extracted from clinical texts.
Approach: They propose a method that leverages reasoning capability of large language models to recognize components of terms and automate decomposition.
Outcome: The proposed strategy achieves state-of-the-art on the experimental dataset.
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation (2025.findings-acl)

Copied to clipboard

Challenge: InfiniteICL is a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems.
Approach: They propose a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems and enables infinite context integration.
Outcome: The proposed framework reduces context length by 90% while achieving 103% average performance of full-context prompting across fact recall, grounded reasoning, and skill acquisition tasks.
SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to address catastrophic forgetting and knowledge transfer in large language models (LLMs) ignore potential of aligning the two modules to effectively address catastrophic forgetting and knowledge transfers simultaneously.
Approach: They propose a Shared Attentive Learning & Selection module to align the PET learning and selection modules to address catastrophic forgetting and knowledge transfer simultaneously.
Outcome: Experiments on two CL benchmarks show that the proposed framework is superior when scaled to different model sizes, different model architectures and unseen tasks.
From Selection to Generation: A Survey of LLM-based Active Learning (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been used for selection and training of data for active learning.
Approach: They propose an intuitive taxonomy that categorizes LLM-based active learning techniques and discuss the transformative roles they can play in the active learning loop.
Outcome: The proposed model can generate entirely new data instances and provide more cost-effective annotations with fewer labeled data instances.
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP (2024.findings-emnlp)

Copied to clipboard

Challenge: a new task is to generate lay definitions of medical terms in EHRs that are difficult to understand for patients.
Approach: They propose a task of automatically generating lay definitions to simplify medical terms into patient-friendly lay language.
Outcome: The proposed model can match or surpass state-of-the-art closed-source large language models like ChatGPT with high-quality data.
Your Reasoning Model is Secretly a Reward Model - Optimization-Free Verification from Experience (2026.acl-long)

Copied to clipboard

Challenge: Existing verifiers operate on the surface text or on confidence proxies derived from token probabilities, which can be brittle.
Approach: They propose a training-free, non-parametric verifier that summarizes each reasoning trace by an activation delta and compares it to two class centroids computed from labeled experience.
Outcome: The proposed model improves selection and reranking on large and less-calibrated models.
PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods evaluate candidate prompts by sampling full outputs, often coupled with self critique or human annotated preferences, which limits scalability, especially for smaller models or models that are not instruction tuned.
Approach: They propose a framework that uses token level cross entropy as a direct, lightweight evaluation signal to evaluate candidate prompts.
Outcome: The proposed framework outperforms prior prompt optimizers across model sizes and datasets.
3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection (2025.emnlp-main)

Copied to clipboard

Challenge: Effective domain adaptation typically involves supervised fine-tuning on carefully selected instruction-tuned data.
Approach: They propose a model-centric data selection framework that aligns data selection with the model’s knowledge distribution to improve model performance.
Outcome: The proposed framework outperforms existing methods by up to 2.97% accuracy in the healthcare domain.
Dense Retrieval as Indirect Supervision for Large-space Decision Making (2023.findings-emnlp)

Copied to clipboard

Challenge: Dense Decision Retrieval (DDR) is a learning-to-retrieve task for discriminative natural language understanding (NLU) tasks with large label spaces.
Approach: They propose a novel approach to learning large-space discriminative NLU tasks as a learning-to-retrieve task by adopting a dual-encoder architecture that learns to predict by retrieving from a decision thesaurus.
Outcome: The proposed approach outperforms baselines greatly on multi-label classification tasks, 1.17% in F1 score ultra-fine entity typing, and 1.26% in accuracy on three few-shot intent classification tasks on average.
UniGeM: Unifying Data Selection and Mixing via Geometric Exploration and Mining (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) scaling is limited by data quality and domain mixing and instance selection are two separate problems.
Approach: They propose a framework that unifies mixing and selection without training proxy models or relying on external reference datasets.
Outcome: The proposed framework achieves 2.0 data efficiency over a random baseline and further improves overall performance compared to SOTA methods in reasoning-heavy evaluations and multilingual generalization.
Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) often refuse to answer legitimate queries, causing models to treat many reasonable prompts as potentially risky.
Approach: They propose a framework that automatically generates and selects overrefusal prompts near the safety boundary.
Outcome: The proposed framework identifies and curates boundary-aligned prompts, enabling more effective and targeted mitigation of overrefusal.
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for I-MCoT fail to capture dynamic needs of vision-language models . existing methods rely on attention signals, which are unreliable under severe granularity imbalance between brief textual query and informative image.
Approach: They propose a framework that integrates specially selected visual evidence into the context of Vision-Language Models (VLMs) they propose 'AIM-CoT' to improve evidence selection and insertion triggering .
Outcome: Experiments across three benchmarks and four backbones demonstrate the proposed framework’s consistent superiority.
STAGE: Simple Text Data Augmentation by Graph Exploration (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained language models (PLMs) are widely used for various tasks, but fine-tuning them requires sufficient data.
Approach: They propose a method for data augmentation that utilizes a word-relation graph to select optimal words for each modification.
Outcome: The proposed method is highly effective across diverse datasets and different PLMs.
From Knowing to Teaching: Scaffolding Pedagogical Decisions for LLM Agent (2026.acl-long)

Copied to clipboard

Challenge: Large language models produce content lacking pedagogical depth when asked to generate lessons .
Approach: They propose a framework that allows teachers to select content according to pedagogical intent and sequence topics so foundations precede applications.
Outcome: The framework achieves 67.8% win rate in human evaluation and 79.6% in LLM-based evaluation against eight baselines.
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to reasoning faithfulness violate constraints, authors say . a science fantasy series and companion books are among the books .
Approach: They propose a framework that enforces verification over internal belief states within the agent before action commitment, achieving faithful reasoning.
Outcome: The proposed framework improves reasoning faithfulness while preserving competitive end-task performance.
EET: Experience-Driven Early Termination for Cost-Efficient Software Engineering Agents (2026.findings-acl)

Copied to clipboard

Challenge: Large language models are reshaping modern software development, but they often incur substantial monetary cost.
Approach: They propose an experience-driven early termination approach that extracts structured experience from prior issue-resolution executions and leverages it to guide early termination during patch generation and selection.
Outcome: The proposed approach reduces cost by 19%–55% with negligible loss in resolution rate (at most 0.2%) EET extracts structured experience from prior issue-resolution executions and leverages it to guide early termination during patch generation and selection.
Robust In-Context Selection via Online Learned Position-Corrected Attention (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to fix this limitation can be classified into two ways: (1) Methods that use the LLM to generate the selection either via logits of item identifiers, or explicit rank permutations often requiring multiple LLM calls or fine-tuning.
Approach: They propose a method that harnesses attention patterns available from a single forward call on the Large Language Model (LLM) the method learns the logic for item selection using a few in-context examples and a simple online position-debiasing mechanism to correct attention distortion.
Outcome: The proposed method improves selection performance over direct generation and prior attention-based methods while remaining robust to prompt variations and item ordering.
Budget-Aware Routing for Long Clinical Text (2026.findings-acl)

Copied to clipboard

Challenge: Long-context capability is now a headline feature of large language models . clinical inputs are long because they are templated, redundant, and stitched from multiple sources.
Approach: They propose a token-constrained subset selection problem with two design choices . they propose heuristics that balance relevance, coverage, diversity and a monotone submodular objective .
Outcome: The proposed model is based on a subset selection problem with two design choices . positional heuristics perform best at low budgets in extractive tasks, while diversity-aware methods improve LLM generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations