Papers with alignment

225 papers
Addressing Issues of Cross-Linguality in Open-Retrieval Question Answering Systems For Emergent Domains (2023.eacl-demo)

Copied to clipboard

Challenge: a lack of cross-lingual training data in emergent domains makes it difficult to train on emerging domains.
Approach: They propose a cross-lingual open-retrieval question answering system for COVID-19 . their system adopts a corpus of scientific articles to ensure that retrieved documents are reliable.
Outcome: The proposed system outperforms BM25 baselines in cross-lingual settings.
Alignment, Acceptance, and Rejection of Group Identities in Online Political Discourse (N18-4)

Copied to clipboard

Challenge: linguistic alignment is a robust and robust form of communication accommodation, and has been detected in a variety of linguistic interactions, ranging from speed dates to the Supreme Court.
Approach: They propose a model to examine alignment in Twitter conversations across antagonistic groups.
Outcome: The proposed model adapts the WHAM alignment model to examine alignment in Twitter conversations across antagonistic groups.
Inverse Reinforcement Learning Meets Large Language Model Alignment (2025.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide a comprehensive review of recent advances in LLM alignment . it will highlight the necessity of constructing neural reward models from human data .
Approach: This tutorial will provide a comprehensive review of recent advances in LLM alignment through the lens of inverse reinforcement learning.
Outcome: This tutorial will provide a comprehensive review of recent advances in LLM alignment through the lens of inverse reinforcement learning (IRL).
Being Kind Isn’t Always Being Safe: Diagnosing Affective Hallucination in LLMs (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly engaged in emotionally vulnerable conversations that extend beyond information seeking to moments of personal distress.
Approach: They propose AHaBench, a benchmark of 500 mental-health-related prompts with expert-informed reference responses, evaluated along three dimensions: Emotional Enmeshment, Illusion of Presence, and Fostering Overdependence.
Outcome: The proposed model is based on 500 mental-health-related prompts with expert-informed reference responses and a 5K-instance preference dataset enabling direct preference optimization (DPO) for alignment with emotionally responsible behavior.
Data and Model Centric Approaches for Expansion of Large Language Models to New languages (2025.emnlp-tutorials)

Copied to clipboard

Challenge: Existing LLMs mainly support English alongside a handful of high resource languages . this leaves a major gap for most low-resource languages despite increasing pace of research .
Approach: This tutorial examines approaches to expand the language coverage of LLMs . they look at tokenizer training, pre-training, instruction tuning, alignment, evaluation, etc.
Outcome: This tutorial examines approaches to expand the language coverage of LLMs . it provides an efficient and viable path to bring LLM technologies to low-resource languages .
Continual Learning of Large Language Models (2025.emnlp-tutorials)

Copied to clipboard

Challenge: This tutorial explores the challenges of continual learning in large language models . participants will learn strategies to mitigate forgetting and manage data and evaluation pipelines .
Approach: This tutorial offers a comprehensive exploration of continual learning in the context of large language models.
Outcome: This tutorial explores the challenges of continual learning in large language models . participants will learn how to manage data and evaluation pipelines and adapt responsibly .
Call for Rigor in Reporting Quality of Instruction Tuning Data (2025.acl-short)

Copied to clipboard

Challenge: Instruction tuning is crucial for adapting large language models (LLMs) to user intentions.
Approach: They propose to use hyperparameters for training models that are often selected arbitrarily without adequate justification to make arbitrary conclusions.
Outcome: The results show that arbitrary hyperparameter decisions can make any arbitrary conclusion.
Automatic Pair Construction for Contrastive Post-training (2024.findings-naacl)

Copied to clipboard

Challenge: Large language models (LLMs) have unprecedented proficiency in a wide array of tasks.
Approach: They propose a way to construct contrastive data using preference pairs from multiple models of varying strengths using SLiC and DPO.
Outcome: The proposed method outperforms existing models like Orca in the comparison of SLiC and DPO with SFT baselines.
Utilizing Language-Image Pretraining for Efficient and Robust Bilingual Word Alignment (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that unsupervised word translation is more accurate and robust without parallel corpora.
Approach: They propose a method for unsupervised word translation that leverages visual observations and pretrained language-image models to align words.
Outcome: The proposed method improves on the state-of-the-art language-image pretraining method for bilingual word alignment.
Versatile Framework for Song Generation with Prompt-based Control (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for song generation fail to generate vocals with prompt-based control and proper alignment.
Approach: VersBand is a multi-task song generation framework for synthesizing high-quality songs with prompt-based control.
Outcome: Experimental results show that VersBand performs better than baseline models across multiple song generation tasks.
Justice in Judgment: Unveiling (Hidden) Bias in LLM-assisted Peer Reviews (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies show that large language models carry implicit biases across race, gender, and religion . prior studies documented such biase based on text generation and classification tasks .
Approach: They investigate bias in large language models by controlling metadata on author metadata . authors found affiliation bias favoring authors from highly ranked institutions .
Outcome: The proposed model favors authors from highly ranked institutions, the authors show . the model also favors author affiliations from highly-ranked institutions .
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) require high quality instruction data for effective alignment, especially in code generation tasks where expert curated datasets are expensive to produce.
Approach: They propose a scalable algorithm for synthesizing large-scale, high quality coding instructions using evolutionary principles.
Outcome: The proposed approach generates 7.5 million coding instructions with a small seed population and is highly parallelizable and effective even with weaker generator models.
Aligning Generative Language Models with Human Values (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods for learning human values do not consider contextual and abstract nature of human values.
Approach: They propose a reinforcement learning based method that embeds human values judgements into each step of language generation.
Outcome: The proposed method improves on human values judgements and shows higher alignment performance.
Data-Efficiently Learn Large Language Model for Universal 3D Scene Perception (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods for 3D scene understanding are limited to specific downstream tasks, hindering their practicality in real-world applications.
Approach: They propose a 3D visual perceptual ability and advanced reasoning capabilities for 3D scenes by aligning 3D representations into the feature space of advanced LLMs.
Outcome: The proposed system achieves a 82.2% relative score compared with state-of-the-art methods with limited data.
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety (2024.acl-demos)

Copied to clipboard

Challenge: a rapid development of Chinese large language models poses big challenges for efficient LLM evaluation.
Approach: They propose an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety.
Outcome: The evaluation platform OpenEval benchmarks Chinese LLMs across capability, alignment and safety.
sudoLLM: On Multi-role Alignment of Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework that allows users to control access rights has not been extensively studied in the large language model realm.
Approach: They propose a framework that allows users to control access rights in a multi-role manner.
Outcome: The proposed framework improves alignment, generalization and resistance to prefix-based jailbreaking attacks.
A Bidirectional Transformer Based Alignment Model for Unsupervised Word Alignment (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for learning word alignment include statistical word aligners (e.g. GIZA++) Existing word alignment models employ a target-to-source attention mechanism which can provide rough word alignments but with a low accuracy.
Approach: They propose a bidirectional Transformer based alignment model for unsupervised learning of the word alignment task.
Outcome: The proposed model outperforms both previous neural word alignment approaches and the popular statistical word aligner GIZA++ on three word alignment tasks.
Effectively Aligning and Filtering Parallel Corpora under Sparse Data Conditions (2020.acl-srw)

Copied to clipboard

Challenge: Parallel corpora are key to developing good machine translation systems, but abundant parallel data is hard to come by for languages with a low number of speakers.
Approach: They propose an unsupervised alignment method that can handle rich morphology by removing incorrect translations and segments containing extraneous data.
Outcome: The proposed method maximizes the number of correctly translated segments in a corpus and minimises noise by removing incorrect translations and segments containing extraneous data.
Mitigating the Alignment Tax of RLHF (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, which is also known as the alignment tax.
Approach: They propose to use a model averaging technique to find the most powerful alignment-forging Pareto front among RLHF algorithms.
Outcome: The proposed method achieves the strongest alignment-forging Pareto front among competing methods.
Adopting the Word-Pair-Dependency-Triplets with Individual Comparison for Natural Language Inference (C18-1)

Copied to clipboard

Challenge: Existing approaches to perform natural language inference ignore syntactic dependency among words or use tree-LSTM to generate sentence representation with irrelevant information.
Approach: They propose to perform natural language inference with Word-Pair-Dependency-Triplets . they propose to compare the triplets of a given passage-pair to make judgement more interpretable .
Outcome: The proposed approach is better than most of the approaches that use tree structures and comparable to other state-of-the-art approaches.
Reinforced Cross-modal Alignment for Radiology Report Generation (2022.findings-acl)

Copied to clipboard

Challenge: Medical images are widely used in clinical decision-making, where writing radiology reports can be enhanced by automatic solutions to alleviate physicians’ workload.
Approach: They propose an approach with reinforcement learning over a cross-modal memory to better align visual and textual features for radiology report generation.
Outcome: The proposed approach improves cross-modal alignment on two English radiology report datasets and human evaluation confirms the results.
Monolingual Phrase Alignment as Parse Forest Mapping (2023.starsem-1)

Copied to clipboard

Challenge: Existing methods for phrase alignment are based on unordered tree mapping . syntactic ambiguities can affect alignment quality, so we expand it to parse forests instead of 1-best trees.
Approach: They propose to expand existing method to align parse forests rather than 1-best trees, where syntactic structures and phrase alignment are simultaneously identified.
Outcome: The proposed method improves the state-of-the-art method by aligning forests rather than 1-best trees.
Just One is Enough: An Existence-based Alignment Check for Robust Japanese Pronunciation Estimation (2025.emnlp-industry)

Copied to clipboard

Challenge: Existence-based alignment has been used to detect pronunciation errors in Japanese NLP, but finding reliable attention heads remains challenging.
Approach: They propose a method that detects and corrects pronunciation errors in Japanese by using beam search.
Outcome: The proposed method reduces hallucinations and omissions and improves pronunciation estimation by over 2.5%.
Tethering Broken Themes: Aligning Neural Topic Models with Labels and Authors (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies suggest that topic models do not align well with human intentions.
Approach: They propose a method to align neural topic models with both labels and authorship information.
Outcome: The proposed method improves existing models in terms of topic quality and alignment.
LEEETs-Dial: Linguistic Entrainment in End-to-End Task-oriented Dialogue systems (2024.findings-naacl)

Copied to clipboard

Challenge: Linguistic entrainment is a phenomenon where linguistic patterns employed by conversational participants converge to one another.
Approach: They propose methods for achieving dialogue entrainment in a task-oriented dialogue system using shared vocabulary.
Outcome: The proposed model produces significantly better entrainment than the base model.
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are often prompted to produce natural language explanations, but the features driving the answer are often different from those emphasized in their explanations.
Approach: They propose a large-scale benchmark linking model decisions with diverse explanations and attribution vectors across datasets, methods, and model families to address this gap.
Outcome: The proposed model generates an answer where the word NLP in the prompt has high feature importance.
Negating Negatives: Alignment with Human Negative Samples via Distributional Dispreference Optimization (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to steer LLMs towards human preference suffer from noisy positive-negative training pairs.
Approach: They propose a distributional preference optimization method which maximizes discrepancy between dispreferred responses and generated non-negative ones.
Outcome: The proposed method achieves comparable generation quality and surpasses the latest strong baselines in producing less harmful and more informative responses with better training stability and faster convergence.
Not that much power: Linguistic alignment is influenced more by low-level linguistic features rather than social power (P18-1)

Copied to clipboard

Challenge: linguistic alignment between interlocutors of higher power is attributed to their relative social power, but studies on low-level linguistic features do not account for these factors.
Approach: They characterize the effect of power on alignment with logistic regression models in two datasets and find it vanishes after controlling for low-level features such as utterance length.
Outcome: The proposed model shows that the effect vanishes or is reversed after controlling for low-level features such as utterance length.
Massively Multilingual Document Alignment with Cross-lingual Sentence-Mover’s Distance (2020.aacl-main)

Copied to clipboard

Challenge: Document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.
Approach: They propose an unsupervised scoring function that leverages cross-lingual sentence embeddings to compute the semantic distance between documents in different languages.
Outcome: The proposed scoring function outperforms baseline methods on high-resource language pairs, 15% on mid-resourced language pairs and 22% on low-resourcing language pairs.
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning? (2026.eacl-industry)

Copied to clipboard

Challenge: Clinical Question-Answering (CQA) industry systems rely on Large Language Models (LLMs).
Approach: They propose a framework that applies alignment at inference time rather than through SFT to help CQA users achieve consistent reasoning.
Outcome: MEDASSESS-X improves Accuracy, Factual Consistency and Safety by up to 50%.
Wait, that’s not an option: LLMs Robustness with Incorrect Multiple-Choice Options (2025.acl-long)

Copied to clipboard

Challenge: Using a framework that combines instruction-following with critical reasoning, we show that the ability of LLMs to override defaults when faced with invalid options is impaired by alignment techniques.
Approach: They propose a framework for evaluating LLMs’ capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers.
Outcome: The proposed framework improves models' ability to override defaults when faced with invalid options while minimizing the impact of model size and training techniques on the model.
Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Multilingual human preference data are difficult to obtain at scale, making it challenging to extend this framework to diverse languages.
Approach: They propose a method where a reward model is trained on preference data in one source language and applied to other target languages.
Outcome: The proposed approach is effective under comprehensive evaluation settings, including human evaluation.
Inducing and Using Alignments for Transition-based AMR Parsing (2022.naacl-main)

Copied to clipboard

Challenge: Abstract Meaning Representation parsers rely on node-to-word alignments, but lack the complexity of the pipeline.
Approach: They propose a neural aligner for abstract meaning representation that learns node-to-word alignments without relying on pipelines.
Outcome: The proposed approach improves accuracy and generalization from AMR2.0 to AMR3.0 corpora.
OpenKI: Integrating Open Information Extraction and Knowledge Bases with Relation Inference (N19-1)

Copied to clipboard

Challenge: Existing methods for knowledge extraction and alignment are limited in quality and performance.
Approach: They propose to integrate OpenIE extractions in the form of (subject, predicate, object) triples with Knowledge Bases (KB)
Outcome: The proposed method improves state-of-the-art for OpenIE extractions and boosts performance on OpenIE from semi-structured data.
Unraveling LLM Jailbreaks Through Safety Knowledge Neurons (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved significant progress in alignment, ensuring safer and more reliable outputs.
Approach: They propose a neuron-level interpretability method that focuses on the role of safety-related knowledge neurons to improve model robustness against jailbreak attacks.
Outcome: The proposed method reduces attack success rates across multiple LLMs and outperforms all baseline defenses.
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Vision-Language Models have demonstrated remarkable capabilities in processing both visual and textual information.
Approach: They examine the challenge of alignment and misalignment in LVLMs through an explainability lens.
Outcome: The findings highlight the need for standardized evaluation protocols and in-depth explainability studies.
LLM-Personalize: Aligning LLM Planners with Human Preferences via Reinforced Self-Training for Housekeeping Robots (2025.coling-main)

Copied to clipboard

Challenge: Large language models have shown significant potential for robotics tasks, but a gap remains in personalization of LLMs to household preferences.
Approach: They propose a framework to personalize LLM planners for household robotics . they use imitation learning and reinforced self-training to personalise the planner .
Outcome: The proposed framework performs iterative planning in multi-room, partially-observable household environments, utilizing a scene graph built dynamically from local observations.
Dissecting Human and LLM Preferences (2024.acl-long)

Copied to clipboard

Challenge: a recent study shows that human and Large Language Model preferences are important for model fine-tuning and evaluation.
Approach: They dissect the preferences of human and 32 different Large Language Models to understand their quantitative composition.
Outcome: The proposed model is compared with 32 different large language models using real-world user-model conversations.
MatchTime: Towards Automatic Soccer Game Commentary Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing data on soccer commentary are often unsatisfactory, and the quality of existing data is often poor.
Approach: They propose to manually annotate timestamps for 49 soccer matches and then use them to create a model to correct and filter existing data.
Outcome: The proposed model improves the viewing experience of soccer and can be trained on the curated dataset.
Retrieved Sequence Augmentation for Protein Representation Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Using multiple sequence alignments (MSA) to extract evolutionary knowledge is limited.
Approach: They propose to use multiple sequence alignments to augment protein representations . they propose to employ Retrieved Sequence Augmentation to enhance protein representation learning .
Outcome: The proposed method surpasses MSA Transformer by 5% in structural and property prediction tasks while being 373 times faster.
Language Model Decoding as Likelihood–Utility Alignment (2023.findings-eacl)

Copied to clipboard

Challenge: Existing studies only compare decoding algorithms in narrow scenarios, and their findings do not generalize across tasks.
Approach: They propose a taxonomy of misalignment mitigation strategies to provide a unifying view of decoding as a tool for alignment.
Outcome: The proposed taxonomy combines likelihood and utility assumptions to provide general statements about decoding as a tool for alignment across tasks.
RS-DPO: A Hybrid Rejection Sampling and Direct Preference Optimization Method for Alignment of Large Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Reinforcement learning with human feedback (RLHF) is widely employed to align large language models with user intent.
Approach: They propose to combine rejection sampling and direct preference optimization to improve alignment with user intent by identifying pairs of contrastive samples from human annotator and alternative LLMs.
Outcome: The proposed method outperforms existing methods including RS, PPO, and DPO in a limited resource environment.
Cross-lingual AMR Aligner: Paying Attention to Cross-Attention (2023.findings-acl)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) graphs embed the semantics of a sentence in a directed acyclic graph, where concepts are represented by nodes, semantic relations between concepts by edges, and the co-references by reentrant nodes.
Approach: They propose a novel aligner for Abstract Meaning Representation graphs that scales cross-lingually and can align units and spans in sentences of different languages.
Outcome: The proposed aligner achieves state-of-the-art in the benchmarks and can scale cross-lingually.
CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling (2025.findings-acl)

Copied to clipboard

Challenge: Reward modeling in large language models is susceptible to reward hacking . flawed reward signals often lead to outputs that optimize for spurious correlates .
Approach: They propose a new approach that generates dynamic, context-relevant criteria to ground the reward model prior to producing reward scores.
Outcome: The proposed approach generates dynamic, context-relevant criteria to ground the model prior to producing reward scores.
Mind the Gap: Multilingual Divide in LLM Bias Detection and Reasoning (2026.acl-srw)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly deployed in multilingual settings . but most bias evaluation remains English-centric and ignores how bias manifests within reasoning .
Approach: They evaluate large language models with supervised fine-tuning and preference optimization . they find that bias varies substantially across languages, with consistent degradation in non-English settings .
Outcome: The proposed model improves in English, Dutch, Spanish, and Turkish using the MBBQ benchmark.
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit impressive performance on a variety of tasks from text summarization to zero-shot common-sense reasoning.
Approach: They propose to manipulate the embedding space of mLLMs by manipulating its activations to steer generation into the desired direction.
Outcome: The proposed model interventions improves alignment of cross-lingual representations in multilingual large language models with up to 2x improvements in top-1 accuracy on cross-linguistic retrieval tasks.
SMATCH++: Standardized and Extended Evaluation of Semantic Graphs (2023.findings-eacl)

Copied to clipboard

Challenge: Existing graph-alignment metrics that measure graph distances are not reliable, we show . metric is spread out and does not provide upper bounds for extended tasks.
Approach: They propose a metric to measure a distance between graphs by aligning nodes and counting matching graph triples.
Outcome: The proposed method reduces search space and improves scoring by reducing the number of errors.
Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback (2026.eacl-long)

Copied to clipboard

Challenge: Novelty assessment is a central yet understudied aspect of peer review . manuscript submissions double roughly every 15 years, and individual reviewers now complete an average of 14 reviews per year.
Approach: They propose a structured approach for automated novelty evaluation that models expert reviewer behavior through three stages: content extraction, retrieval and synthesis of related work, and structured comparison for evidence-based assessment.
Outcome: The proposed approach outperforms existing LLM-based baselines on 182 ICLR 2025 submissions with human-annotated reviewer novelty assessments.
A Language-First Approach for Procedure Planning (2023.findings-acl)

Copied to clipboard

Challenge: Developing intelligent agents requires the ability to produce plans on the fly based on visual observations.
Approach: They propose a language-first procedure planning framework with a modularized design . they first align current and goal observations with corresponding steps and then use a pre-trained LM to predict intermediate steps.
Outcome: The proposed framework matches state-of-the-art procedures on COIN and CrossTask benchmarks.
Large Language Models Do Multi-Label Classification Differently (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied.
Approach: They propose to use initial probability distributions to analyze output distributions of LLMs at each label generation step to find out how LLM models perform multi-label classification.
Outcome: The proposed methods improve alignment and predictive performance over existing methods.
Learning to Paraphrase for Alignment with LLM Preference (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit the issue of paraphrase divergence, which means that when a question is phrased in a slightly different but semantically similar way, LLM may output a wrong response . retraining faces challenges in meeting the computational costs and privacy security demands of LLMs.
Approach: They propose a black-box method that enhances model performance by paraphrasing questions in expressions preferred by the model.
Outcome: The proposed method improves performance by paraphrasing questions in expressions preferred by the model.
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) rely on safety alignment to avoid malicious user inputs.
Approach: They employ weak classifiers to explain LLM safety through the intermediate hidden states.
Outcome: The proposed model can identify malicious and normal inputs and detect malicious ones without jailbreak.
A Systematic Investigation of KB-Text Embedding Alignment at Scale (2021.acl-long)

Copied to clipboard

Challenge: Knowledge bases (KBs) and text often contain complementary knowledge.
Approach: They propose a framework for aligning KB and text embeddings for joint reasoning . they also evaluate alignment methods to infuse textual information into KB embeddables .
Outcome: The proposed framework can be used to predict link prediction on emerging entities and events using textual information.
From Detection to Explanation: Effective Learning Strategies for LLMs in Online Abusive Language Research (2025.coling-main)

Copied to clipboard

Challenge: Abusive language detection requires commonsense reasoning, world knowledge and linguistic nuances that evolve over time.
Approach: They propose a knowledge-guided version of Llama-2 instruction fine-tuned for multi-class abusive language detection and explanation generation that mitigates bias and generates explanations that are relevant to the text and coherent with human reasoning.
Outcome: The proposed model mitigates bias and generates explanations that are relevant to the text and coherent with human reasoning, with an average 48.76% better alignment with human judgment.
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for directing language model outputs are limited in their accuracy due to a distributional gap . existing methods train static value functions on trajectories sampled exclusively from the base policy .
Approach: They propose a framework to bridge a distributional gap in the accuracy of value functions . they propose RLHF to align language models with human values and task requirements .
Outcome: The proposed framework reduces computational costs and improves value function accuracy by leveraging principled value function optimization.
Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have explored the working mechanisms of In-Context Learning (ICL) however, they mainly focus on classification and simple generation tasks, limiting their broader application to more complex generation tasks in practice.
Approach: They propose an efficient Progressive In-Context Alignment method that embeds the task function learned from demonstrations into the separator token representation.
Outcome: The proposed method surpasses vanilla ICL and achieves comparable performance to other alignment tuning methods.
How Transliterations Improve Crosslingual Alignment (2025.coling-main)

Copied to clipboard

Challenge: Recent studies show that post-aligning multilingual pretrained language models improve crosslingual alignment, but it is unclear how and why this is achieved.
Approach: They propose to explicitly evaluate crosslingual alignment by adding transliterations to models using original and transliterated data.
Outcome: The proposed approach improves crosslingual alignment even for random sentences.
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs (2025.emnlp-main)

Copied to clipboard

Challenge: Existing reward models have a high performance on benchmarks, but performance degradation is often due to overfitting.
Approach: They propose to explicitly train reward models to assign similar scores to paraphrases to improve their robustness.
Outcome: The proposed model reduces degradation by half for the Chat Hard subset in RewardBench.
Extracting and Understanding the Superficial Knowledge in Alignment (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have shown that alignment of large language models with human values and preferences requires substantial data and computation resources.
Approach: They propose a method to extract and isolate superficial knowledge from aligned models by focusing on the shallow modifications to the final token selection process.
Outcome: The proposed method extracts and isolates superficial knowledge from aligned models, focusing on the shallow modifications to the final token selection process.
Multilingual Factor Analysis (P19-1)

Copied to clipboard

Challenge: Existing methods for multilingual word embeddings are based on the observation that word embeds exhibit similar structures across languages.
Approach: They propose a latent variable-based model that fits a multilingual dictionary to learn multilingual word representations offline.
Outcome: The proposed model is robust to noise in the embedding space making it suitable for distributed representations learned from noisy corpora.
LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing cross-lingual topic models depend on sparse bilingual resources and often yield incoherent or weakly aligned topics.
Approach: They propose a framework that integrates LLM-guided topic refinement with self-consistency uncertainty quantification to enable black-box, stable, and scalable enhancement of cross-lingual topic models.
Outcome: Experiments on multilingual corpora show that the proposed framework achieves superior topic coherence and alignment while reducing reliance on bilingual dictionaries and expensive LLM calls.
Assessing Non-autoregressive Alignment in Neural Machine Translation via Word Reordering (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing non-autoregressive neural machine translation models that implicitly model dependencies are sub-optimal in handling word order errors.
Approach: They propose to learn a non-autoregressive language model that can be combined with Viterbi decoding to achieve better reordering performance.
Outcome: The proposed model outperforms state-of-the-art reordering mechanisms under different word permutation settings with a 2-27 BLEU improvement, suggesting high potential for word alignment in NAT.
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity (2025.findings-emnlp)

Copied to clipboard

Challenge: Using real-world datasets, we conduct the most comprehensive study to date, auditing various state-of-the-art reward models across nine sensitive attributes, including age, gender, ethnicity, etc.
Approach: They propose a method to mitigate group disparities in reward modeling by using real-world data.
Outcome: The proposed method is based on a population-based dataset with nine demographic attributes, including gender, ethnicity, age, gender, and ethnicity.
Conformal Feedback Alignment: Quantifying Answer-Level Reliability for Robust LLM Alignment (2026.findings-eacl)

Copied to clipboard

Challenge: Existing uncertainty-aware approaches weight preferences, but ignore reliability of the answers being compared.
Approach: They propose a framework that grounds preference weighting in Conformal Predictions to address this problem.
Outcome: The proposed framework improves alignment robustness and data efficiency across different datasets.
ALIGNMEET: A Comprehensive Tool for Meeting Annotation, Alignment, and Evaluation (2022.lrec-1)

Copied to clipboard

Challenge: Summarization is a challenging problem, and it is difficult to create, correct, and evaluate the summaries manually.
Approach: They propose an open-source tool for meeting annotation, alignment, and evaluation . the tool aims to provide an efficient and clear interface for fast annotation .
Outcome: The proposed tool is open-source and installable from PyPI.
Exploring the Relationship between Alignment and Cross-lingual Transfer in Multilingual Transformers (2023.findings-acl)

Copied to clipboard

Challenge: despite lack of explicit cross-lingual training data, multilingual models can achieve cross-linguistic transfer.
Approach: They find alignment is significantly correlated with cross-lingual transfer . they advocate for further research on realignment methods for smaller models .
Outcome: The proposed method outperforms XLM-R Large in POS-tagging between English and Arabic by +15.8 accuracy.
Icon2: Aligning Large Language Models Using Self-Synthetic Preference Data via Inherent Regulation (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) require high quality preference datasets to align with human preferences.
Approach: They propose a framework that leverages inherent regulation of LLMs’ representation space for efficient and tailored preference dataset construction, named Icon2.
Outcome: The proposed framework improves performance on benchmarks like AlpacaEval 2.0 and Arena-Hard while reducing computational costs by up to 48.1%.
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluations of large language models (LLMs) focus on a single output per example, which limits our understanding of LLM performance variability in real-world applications.
Approach: They explore the performance differences between greedy decoding and sampling and identify benchmarks’ consistency regarding non-determinism and examine unique model behaviors.
Outcome: The proposed model outperforms sampling methods and greedy decoding outperformed other models.
PROMPTEVALS: A Dataset of Assertions and Guardrails for Custom Production Large Language Model Pipelines (2025.naacl-long)

Copied to clipboard

Challenge: Large language models fail to follow instructions or meet developer expectations when running in production . a dataset of 2087 LLM pipeline prompts with 12623 assertion criteria is larger than previous collections .
Approach: They propose a dataset of 2087 LLM pipeline prompts with 12623 assertion criteria . they fine-tuned Mistral and Llama 3 models outperform GPT-4o by 20.93% on average .
Outcome: The proposed dataset outperforms GPT-4o and mistral models in generating assertions and offers reduced latency and improved performance.
Interpretable Safety Alignment via SAE-Constructed Low-Rank Subspace Adaptation (2026.acl-long)

Copied to clipboard

Challenge: Prior work has shown that safety behaviors are governed by low-rank structures . Low-Rank Adaptation (LoRA) consistently underperforms full fine-tuning and reinforcement learning on safety benchmarks .
Approach: They propose a safety alignment system that disentangles safety-relevant directions into monosemantic features and constructs an interpretable safety subspace from SAE directions.
Outcome: Empirically, the proposed model achieves 99.6% safety rates across multiple model families and scales . low-rank Adaptation consistently underperforms full fine-tuning and reinforcement learning on safety benchmarks compared with previous methods .
Deep Subjecthood: Higher-Order Grammatical Features in Multilingual BERT (2021.eacl-main)

Copied to clipboard

Challenge: a recent study has shown that multilingual BERT encodes sentences in structurally meaningful ways.
Approach: They analyze how morphosyntactic alignment manifests across embedding spaces of languages . they train classifiers to recover subjecthood of mBERT embedds in transitive sentences .
Outcome: The proposed model encodes a high-order grammatical feature of morphosyntactic alignment across languages . the results show that the classifier distributions reflect the morphological alignment of their training languages based on the results .
Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for training large language models require additional annotations to adjust to shifted distributions.
Approach: They propose an algorithm that allows LLMs and reward models to update alternatively via a min-max game to improve their alignment.
Outcome: The proposed framework improves existing alignment baselines in terms of LLM helpfulness and harmlessness.
LUNA: Learning Slot-Turn Alignment for Dialogue State Tracking (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods exploit the utterances of all dialogue turns to assign value to slots . this can lead to suboptimal results due to information introduced from irrelevant utterrances .
Approach: They propose a SLot-TUrN Alignment enhanced approach to assign slot value . they explicitly align each slot with its most relevant utterance and then predict the corresponding value based on this aligned utteration.
Outcome: The proposed approach achieves state-of-the-art on three multi-domain task-oriented dialogue datasets.
Deciphering Rumors: A Multi-Task Learning Approach with Intent-aware Hierarchical Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Social networks are rife with noise and misleading information, presenting multifaceted challenges for rumor detection.
Approach: They propose a new multi-task learning framework that mines latent intentions and rumor semantic features . they propose to use event-level and intent-level strategies to establish cognitive anchors .
Outcome: The proposed framework improves the effectiveness of rumor detection and addresses the challenges present in the field.
MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing Benchmark (2021.eacl-main)

Copied to clipboard

Challenge: Existing datasets for task-oriented dialog systems are limited and expensive . current models are based on the simple intent and slot detection paradigm for non-compositional queries.
Approach: They propose to use a multilingual dataset to scale semantic parsing models to new languages . they demonstrate an average improvement of +6.3 points on Slot F1 for existing datasets .
Outcome: The proposed model achieves an average improvement of +6.3 points on Slot F1 over existing models.
Pseudo-Likelihood Training for Reasoning Diffusion Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are the backbone of modern natural language processing and are powering applications ranging from code generation to autonomous agents.
Approach: They propose a framework that uses pseudo-likelihood based objective for alignment of diffusion based language models (dLLMs).
Outcome: The proposed method matches or surpasses dLLM training baselines on various coding and mathematical reasoning benchmarks.
Benchmarking Direct Preference Optimization for Medical Large Vision–Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Large vision-language models (LVLMs) are gaining traction in clinical tasks such as diagnostic support, report generation, and medical question answering.
Approach: They present a systematic evaluation of nine DPO variants applied to two leading medical LVLMs.
Outcome: The proposed model improves alignment and reduces severe hallucinations, but yields inconsistent gains over supervised fine-tuning.
ChatGPT Rates Natural Language Explanation Quality like Humans: But on Which Scales? (2024.lrec-main)

Copied to clipboard

Challenge: Traditionally, evaluating NLEs through gathering human judgments is a tedious task due to the subjective nature of human evaluations.
Approach: They examine the alignment between ChatGPT and human assessments across multiple scales and compare them using paired comparisons and dynamic prompting.
Outcome: The proposed model aligns better with humans in coarser scales and provides semantically similar examples in the prompt.
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for large language models (LLMs) are limited to small sample and fail to demonstrate LLM susceptibility to context with potential human bias.
Approach: They propose a benchmark for evaluating LLM investment decision-making when faced with uncertainty and possible human-biased opinions.
Outcome: The proposed model can herd the explicit bias in context and even exceed human performance in predicting future stock return.
What Media Frames Reveal About Stance: A Dataset and Study about Memes in Climate Change Discourse (2025.findings-emnlp)

Copied to clipboard

Challenge: Media framing is a method of shaping public perceptions of issues, but the interaction between stance and media frame remains unexplored.
Approach: They propose to use a dataset of climate-change memes annotated with stance and media frames to conceptualize and computationally explore this interaction.
Outcome: The proposed dataset includes 1,184 climate-change memes sourced from 47 subreddits and enables analysis of frame prominence over time and communities.
Mutual Gaze and Linguistic Repetition in a Multimodal Corpus (2022.lrec-1)

Copied to clipboard

Challenge: a study of linguistic repetitions and mutual understanding is conducted . we find no compelling correlation between mutual gaze and duration of the event .
Approach: They investigate the correlation between mutual gaze and linguistic repetition, a form of alignment, which they take as evidence of mutual understanding.
Outcome: The proposed method is based on the Multisimo corpus, a multimodal corpus which provides authentic task-based interactions among three participants.
Transformer-based Causal Language Models Perform Clustering (2025.findings-naacl)

Copied to clipboard

Challenge: Recent studies have shown great improvements in instruction-following capability through additional training for instruction- following tasks.
Approach: They propose to use a Transformer-based causal language model to study instruction-following capabilities.
Outcome: The proposed model learns task-specific information by clustering data within its hidden space, with this clustering process evolving dynamically during learning.
The Pluralistic Moral Gap: Understanding Moral Judgment and Value Differences between Humans and Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Existing studies have shown that Large Language Models (LLMs) are not fully aligned with human moral judgments.
Approach: They propose a dataset of 1,618 real-world moral dilemmas paired with a distribution of human moral judgments consisting of a binary evaluation and a free-text rationale to examine how closely LLMs align with human moral judgements.
Outcome: The proposed model reproduces human judgments only under high consensus; alignment deteriorates sharply when human disagreement increases.
E-Verify: A Paradigm Shift to Scalable Embedding-based Factuality Verification (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing factuality verification methods follow a Decompose-Then-Verify paradigm, which improves granularity but suffers from poor scalability and efficiency.
Approach: They propose a Decompose-Embed-Interact paradigm that shifts factuality verification from costly text-level reasoning to efficient alignment in embedding space.
Outcome: The proposed paradigm shifts factuality verification from costly text-level reasoning to efficient alignment in embedding space .
Aligning Cross-lingual Sentence Representations with Dual Momentum Contrast (2021.emnlp-main)

Copied to clipboard

Challenge: Existing work uses sentences within the same batch as negatives, which suffers from easy negatives.
Approach: They propose to align sentence representations from different languages into a unified embedding space . they adapt MoCo to further improve the quality of alignment .
Outcome: The proposed model achieves state-of-the-art on several tasks.
Concept Space Alignment in Multilingual LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Multilingual large language models generalize somewhat across languages, but it is unclear whether this is a result of improved, implicit alignment, or of something else, e.g., linguistic overlap or semi-parallel subsets of training data.
Approach: They hypothesize that implicit alignment is the reason for generalization in multilingual large language models.
Outcome: The proposed model generalizes well across languages, but lacks linearity.
Constructing Your Model’s Value Distinction: Towards LLM Alignment with Anchor Words Tuning (2025.findings-emnlp)

Copied to clipboard

Challenge: a study of large language models (LLMs) shows that they can generate outputs that are honest, positive, harmless, etc.
Approach: They propose a method that amplifies logits difference between positive and negative tokens . they propose to use the logits gap to generate positive and positive tokens after alignment .
Outcome: The proposed method achieves effective alignment, but requires fewer computational resources compared to training-time alignment methods.
Synergizing In-context Learning with Hints for End-to-end Task-oriented Dialog Systems (2024.emnlp-main)

Copied to clipboard

Challenge: Existing end-to-end task-oriented dialogue systems require extensive training datasets to perform well.
Approach: They propose a system that synergizes LLMs with task-specific hints to improve alignment in low-data settings.
Outcome: The proposed model improves alignment in low-data settings while retaining competitive performance in full-data environments.
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs (2026.eacl-long)

Copied to clipboard

Challenge: Existing approaches to align pre-trained LLMs with instructions for one property are difficult to fine-tune.
Approach: They propose a mixture-of-experts-based fusion mechanism that models alignment as a controllable drift within the subspace, guided by a drift-regularization loss to balance competing alignment dimensions.
Outcome: Extensive evaluations of three benchmark datasets show that H3Fusion outperforms each individually aligned model by 11.37% and provides stronger robustness compared to the state-of-the-art LLM ensemble approaches by 13.77% and model-merging approaches by 6.18 %.
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent studies suggest different methods to improve multilingual word representations in contextualized settings including techniques that align between source and target embedding spaces.
Approach: They propose to use Optimal Transport as an alignment objective during fine-tuning to improve multilingual contextualized representations for downstream cross-lingual transfer.
Outcome: The proposed method achieves better performance on two tasks (XNLI and XQuAD) and is competitive with existing methods.
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation (2025.findings-acl)

Copied to clipboard

Challenge: Existing alignment methods focus on reactive feedback, where immediate human perception is leveraged to judge sampled model responses as preference data for post-training.
Approach: They propose a proof-of-concept framework that projects how model-generated advice could propagate through societal systems on a macroscopic scale over time, enabling more robust alignment.
Outcome: The proposed framework achieves 20% improvement on existing safety benchmarks and an average win rate exceeding 70% against strong baselines.
Table-based Fact Verification With Salience-aware Learning (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for fact verification use tabular data with tokens, but training requires labeled training data.
Approach: They propose a system that identifies token-level salience in the statement with probing-based saliency estimation.
Outcome: The proposed system improves on TabFact benchmark by replacing non-salient terms with tokens.
Joint Multilingual Knowledge Graph Completion and Alignment (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing work on multilingual KG completion has focused on entity and relation alignments, but understanding of how it can aid multilingual alignments is limited.
Approach: They propose to combine two components that jointly accomplish KG completion and alignment.
Outcome: The proposed model outperforms existing competitive baselines on a public multilingual benchmark and achieves state-of-the-art results.
When to Trust LLMs: Aligning Confidence with Response Quality (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods express reliability by confidence level, but lack objective guidance . Existing approaches express reliability but lack guidance on when to trust LLMs .
Approach: They propose a reward-based approach to align confidence with quality to ensure reliability . they propose 'conqORD' to help model to verbalize greater confidence for higher quality responses .
Outcome: Experiments show that CONQORD significantly improves confidence and response accuracy . the proposed approach can be used to determine reliability of large language models .
Zero-shot Cross-lingual Alignment for Embedding Initialization (2024.findings-acl)

Copied to clipboard

Challenge: CrossInit initializes embeddings into similar geometrical structures across languages in unsupervised manner.
Approach: They propose a method that initializes embeddings into similar geometrical structures across languages in an unsupervised manner.
Outcome: The proposed method demostrates similar patterns in low-resource and dissimilar languages.
Teaching Language Models to Self-Improve by Learning from Language Feedback (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) generate content that can be untruthful or harmful.
Approach: They propose a method that leverages model feedback for alignment . they use a base language model to generate initial responses, critiqued and refined .
Outcome: The proposed method outperforms strong baselines across diverse tasks and model sizes.
Reasoning Like a Doctor: Improving Medical Dialogue Systems via Diagnostic Reasoning Process Alignment (2024.findings-acl)

Copied to clipboard

Challenge: Medical dialogue systems have attracted significant attention for their potential to act as medical assistants.
Approach: They propose a framework that emulates clinicians' diagnostic reasoning processes and aligns with clinician preferences through thought process modeling.
Outcome: The proposed framework generates appropriate responses that relies on abductive and deductive diagnostic reasoning analyses and aligns with clinician preferences through thought process modeling.
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent efforts in LLM alignment focus on instance-wise supervision, costing substantial . ZEBRA binarizes response pairs by evaluating the quality and similarity of their origin models .
Approach: They propose a model behavior-wise zero-annotation framework that binarizes preference data . ZEBRA binarized response pairs by evaluating the quality and similarity of their origin models .
Outcome: The proposed framework achieves comparable alignment performance to instance-supervised methods .
AbsVis – Benchmarking How Humans and Vision-Language Models “See” Abstract Concepts in Images (2025.emnlp-main)

Copied to clipboard

Challenge: Abstract concepts like mercy and peace lack clear visual grounding, and therefore challenge humans and models to provide suitable image representations.
Approach: They propose a dataset of 675 images annotated with 14,175 concept–explanation attributions from humans and two Vision-Language Models where each concept is accompanied by a textual explanation.
Outcome: The proposed dataset compares human and VLM attributions in terms of diversity, abstractness, and alignment, and shows that overlapping concepts are most preferred.
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective (2022.findings-emnlp)

Copied to clipboard

Challenge: In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks.
Approach: They propose a new probing method that is based on image captioning to first empirically study the cross-modal semantics alignment of VLP models.
Outcome: The proposed method analyzes captions generated by five popular VLP models to reveal how well they align with visual words and how well these align with images.
Improving Alignment in LVLMs with Debiased Self-Judgment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for aligning LVLMs rely on external datasets, human annotations or complex post-processing.
Approach: They propose a method that generates a debiased self-judgment score for LVLMs . this self-evaluation metric is created internally by the model without external resources .
Outcome: The proposed approach outperforms existing methods in reducing hallucinations and safety concerns.
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words? (2024.emnlp-main)

Copied to clipboard

Challenge: Despite their unprecedented capabilities, large language models (LLMs) often output erroneous information, which may lead users to overly rely on their false output.
Approach: They formalize faithful response uncertainty based on the gap between the model’s intrinsic confidence in the assertions it makes and the decisiveness by which they are conveyed.
Outcome: The proposed model is poor at faithfully conveying uncertainty on knowledge-intensive questions.
DavIR: Data Selection via Implicit Reward for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: 6% of Alpaca dataset selected with DavIR can steer both LLaMA and Gemma models to produce superior performance compared to the same models trained on the full 52K dataset.
Approach: They propose a model-based data selection method for post-training Large Language Models . they generalize Reducible Holdout Loss to core-set selection problem of causal language modeling .
Outcome: The proposed method can steer both LLaMA and Gemma models to superior performance compared to the same models trained on the full 52K dataset.
A Layer-wise Analysis of Supervised Fine-Tuning (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning ignore depth-dependent heterogeneity of instruction-following . a critical gap remains in understanding where these changes occur across the model's depth and which layers are essential for instruction- following.
Approach: They propose a method which selectively updates critical intermediate layers . they show that effective alignment is architecturally localized rather than distributed .
Outcome: The proposed method outperforms standard LoRA up to 10.2% on GSM8K with reduced parameter overhead.
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods rely on inference-time interventions, which are limited in attention adaptation or require additional supervision.
Approach: They propose a framework for automatic attention alignment tuning that leverages weak labels from SAM and selectively modifies visually-critical attention heads to improve alignment while minimizing interference.
Outcome: The proposed framework outperforms state-of-the-art models on medical VQA and report generation benchmarks.
Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization (2025.emnlp-main)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains.
Approach: They propose an approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs.
Outcome: The proposed approach outperforms state-of-the-art approaches in performance across multiple model scales and widely adopted benchmark datasets.
LAReQA: Language-Agnostic Answer Retrieval from a Multilingual Pool (2020.emnlp-main)

Copied to clipboard

Challenge: LAReQA tests for “strong” cross-lingual alignment, requiring semantically related cross-language pairs to be closer in representation space than unrelated same-language pair.
Approach: They propose a new benchmark for language-agnostic answer retrieval from a multilingual candidate pool that tests for "strong" cross-lingual alignment . they augment training data via machine translation and find that model performance is improved by augmenting training data through machine translation .
Outcome: The proposed task is based on multilingual BERT (mBERT) and XLM-R.
Token-Aware Editing of Internal Activations for Large Language Model Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to optimize the behavior of large language models neglect misalignment discrepancies among tokens, resulting in deviant alignment direction and inflexible editing strength.
Approach: They propose a token-aware editing approach to exploit the misalignment discrepancy among tokens to enhance activation probing and facilitate intervention.
Outcome: Extensive experiments on three alignment capabilities demonstrate the efficacy of the proposed approach surpassing baseline by 25.8% on the primary metric of truthfulness with minimal cost.
Value FULCRA: Mapping Large Language Models to the Multidimensional Spectrum of Basic Human Value (2024.naacl-long)

Copied to clipboard

Challenge: Existing work specifies values as risk criteria formulated in the AI community, e.g., fairness and privacy protection, suffering from poor clarity, adaptability and transparency.
Approach: They propose a value alignment paradigm based on Schwartz's Theory of Basic Values as an instantiation and propose 'BaseAlign' to support this paradigm.
Outcome: The proposed model covers existing risks and anticipates unidentified ones with a low-data set.
A Thorough Examination of Decoding Methods in the Era of LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Decoding methods are essential for converting language models from next-token predictors into practical task solvers.
Approach: They propose to evaluate decoding methods in general-purpose large language models . they find that decoding method performance is notably task-dependent .
Outcome: The proposed methods perform task-dependently and are influenced by alignment, model size, and quantization.
How Far Can In-Context Alignment Go? Exploring the State of In-Context Alignment (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have demonstrated that In-Context Learning (ICA) can align Large Language Models (LLMs) with human preferences without requiring parameter adjustments.
Approach: They investigate the effectiveness of each part in enabling ICA to function effectively and examine how variants in these parts impact alignment performance.
Outcome: The proposed model can comprehend human instructions without parameter adjustments.
Annotation alignment: Comparing LLM and human annotations of conversational safety (2024.emnlp-main)

Copied to clipboard

Challenge: We examine whether LLMs and humans agree when annotating the safety of user-chatbot conversations.
Approach: They leverage a recent DICES dataset in which 350 conversations are each rated for safety by 112 annotators spanning 10 race-gender groups.
Outcome: The LLMs annotators are compared to human annotator demographic groups and can predict when one group finds a conversation unsafe .
From Real to Synthetic: Synthesizing Millions of Diversified and Complicated User Instructions with Attributed Grounding (2025.acl-long)

Copied to clipboard

Challenge: a pursuit of diverse, complex, and large-scale instruction data is crucial for automatically aligning large language models . authors: methods that generate synthetic instructions at scale suffer from limited grounding sources . attributed grounding is a technique that can be used to align language models with human .
Approach: They synthesize 1 million instructions using attributed grounding and a bottom-up synthesis process that leverages web documents to generate a situation, then a meaningful instruction.
Outcome: The proposed framework achieves leading performance on benchmarks and scales with more web corpora.
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for complex instruction-following with elaborate constraints rely on a weaker model, especially GPT-4, limiting their application.
Approach: They propose a Multi-granularity Self-Contrastive Training framework to improve instruction alignment without relying on a stronger model.
Outcome: The proposed framework improves instruction-following with elaborate constraints without external supervision on coarse and fine granularity.
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models (2026.acl-long)

Copied to clipboard

Challenge: Masked diffusion language models have achieved significant progress in language modeling . however, the systematic analysis and empirical validation of their alignment on general tasks remains underexplored.
Approach: They propose a framework that analyzes the bias and variance of preference optimization loss and gradient based on Direct Preference Optimization.
Outcome: The proposed model outperforms its SFT-only predecessor on general benchmarks . it consistently outperformed other strong language models and ARMs on general tasks .
HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in Reinforcement Learning with Verifiable Rewards (RLVR) have gained significant attention due to their objective and verifiably verifier reward signals.
Approach: They propose to exploit RLVR for alignment reversibility by using GRPO to reverse alignment with merely 64 harmful prompts without responses.
Outcome: The proposed method outperforms fine-tuning and RLHF in reasoning and code generation tasks while maintaining general capabilities.
Aligning Images and Text with Semantic Role Labels for Fine-Grained Cross-Modal Understanding (2022.lrec-1)

Copied to clipboard

Challenge: Currently, image retrieval systems can retrieve relevant results for diverse inputs, but they do not provide a way to intentionally inject variety into the search results.
Approach: They propose a multimodal dataset that combines semantic annotations with image bounding boxes.
Outcome: The proposed system improves image retrieval performance and flexibility.
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for inference-time steering are limited by their limitations . Angular Steering violates norm preservation, causing distribution shift and generation collapse .
Approach: They propose a method that uses a norm-preserving rotation formulation to maintain activation distribution integrity and discriminative layer selection to apply steering only where features exhibit opposite-signed class alignment.
Outcome: Experiments show that Selective Steering achieves higher attack success rates than prior methods while maintaining zero perplexity violations and approximately 100% capability retention on standard benchmarks.
On Diversified Preferences of Large Language Model Alignment (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) can be fine tuned with human feedback, but human preferences can be diversified due to annotators’ different tastes, which hinders the effectiveness of LLM alignment methods.
Approach: They propose a calibration error metric to evaluate large language models (LLMs) and a multi-objective reward learning method to enhance the calibration performance of RMs on shared preferences.
Outcome: The proposed model can be adopted as a key calibration error and MORE can achieve superior alignment performance.
Dynamic Parallel Tree Search for Efficient LLM Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Recent methods focus on search accuracy while overlooking computational efficiency.
Approach: They propose a parallelism framework that dynamically optimizes reasoning path in inference.
Outcome: The proposed framework improves efficiency by 2-4 on average while maintaining or even surpassing existing reasoning algorithms in accuracy.
A Systematic Examination of Preference Learning through the Lens of Instruction-Following (2025.naacl-long)

Copied to clipboard

Challenge: a recent study has found that preference learning is a key tool for enhancing LLM training and alignment.
Approach: They use a synthetic data generation pipeline to generate 48,000 unique instruction-following prompts with 23 verifiable constraints to obtain preference pairs.
Outcome: The proposed pipeline generates 48,000 unique instruction-following prompts with 23 verifiable constraints that enable fine-grained and automated quality assessments of model responses.
Propaganda Signals in LLMs: Perspectival Divergence and Narrative Framing in the Russia-Ukraine War (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models are increasingly used to explain, summarize, and translate real-world events . a recent study examined whether LLMs reproduce conflict-specific propaganda .
Approach: They evaluate LLMs under several prompting contexts to determine which side they are closer to . they find model-specific leanings and technique profiles that persist across prompts .
Outcome: The proposed model outputs align with competing narratives from different information ecosystems.
Sample Efficient Alignment Learning With Episodic Control (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing parametric methods for aligning large language models with task objectives are limited.
Approach: They propose a non-parametric framework that aligns large language models with task objectives . they use a key-value memory to store associations between generated text and its corresponding values .
Outcome: The proposed framework outperforms state-of-the-art baselines on harmless, helpful, and summarization tasks.
PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models exhibit reasonable multilingual abilities, despite predominantly English-centric pretraining.
Approach: They propose a framework that establishes multilingual alignment prior to language model pretraining and preserves this alignment using a code-switching strategy during pretraining.
Outcome: Experiments in a synthetic English to English-Clone setting show that PreAlign outperforms standard multilingual joint training in language modeling, zero-shot cross-lingual transfer, and cross-linguistic knowledge application.
How Reliable is Multilingual LLM-as-a-Judge? (2025.findings-emnlp)

Copied to clipboard

Challenge: LLMs are a popular evaluation strategy, but their reliability in multilingual evaluation remains uncertain.
Approach: They evaluate five models from different model families across five diverse tasks involving 25 languages.
Outcome: The models perform poorly across languages and average Fleiss’ Kappa is 0.3 .
Optimal Transport-based Alignment of Learned Character Representations for String Similarity (P19-1)

Copied to clipboard

Challenge: String similarity models are crucial for record linkage, data integration, search and entity resolution systems.
Approach: They propose a model that encodes the characters of each string, aligns the encodings using Sinkhorn Iteration and scores the alignment with a convolutional neural network.
Outcome: The proposed model outperforms state-of-the-art and classical similarity models on four of the five datasets and improves performance by applying it to cross-document coreference.
CommunityLM: Probing Partisan Worldviews from Language Models (2022.coling-1)

Copied to clipboard

Challenge: Political polarization is accelerating as political discourse diverges linguistically . et al. ( 2017) show that partisanship makes reliable predictions about an individual's word understanding .
Approach: They propose a framework that probes community-specific responses to a survey using community language models CommunityLM.
Outcome: The proposed framework can query the worldview of any group of people given a sufficiently large sample of their social media discussions or media diet.
The MWN.PT WordNet for Portuguese: Projection, Validation, Cross-lingual Alignment and Distribution (2020.lrec-1)

Copied to clipboard

Challenge: Lexical semantic networks are pervasive in natural language processing . Lexical ontologies play a key role in virtually all major applications .
Approach: The present paper presents the MWN.PT WordNet for Portuguese . it is the largest high quality, manually validated and cross-lingually integrated wordnet of Portuguese based on the Princeton WordNet of English .
Outcome: The MWN.PT WordNet for Portuguese includes 41,000 concepts expressed by 38,000 lexical units.
Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization (2025.acl-long)

Copied to clipboard

Challenge: Large language models generate unintended outputs due to their unsupervised nature.
Approach: They propose a method to construct preference pairs of selected and rejected LLMs by repeated random sampling to improve alignment performance.
Outcome: The proposed method improves performance as the sample size increases.
Split and Merge: Aligning Position Biases in LLM-based Evaluators (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have shown promise as automated evaluators for assessing the quality of answers generated by AI systems.
Approach: They propose an alignment-based system that calibrates position bias in a lightweight yet effective manner by taking into account both length and semantics and combining them into a single prompt.
Outcome: Extensive experiments with six LLMs on 11,520 answer pairs show that PORTIA significantly improves consistency and consistency rates with humans.
AlignBench: Benchmarking Chinese Alignment of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment.
Approach: They propose a multi-dimensional benchmark for evaluating LLMs’ alignment in Chinese with 8 main categories, 683 real-scenario rooted queries and corresponding human verified references.
Outcome: The benchmark uses a human-in-the-loop data curation pipeline, 683 real-scenario rooted queries and human verified references.
ReAlign: Structured Revision for Small Language Model Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: weak policies struggle to generate informative on-policy samples and suffer from unstable gradients when trained on off-police signals from stronger models.
Approach: They propose a training framework that combines stability of on-policy learning with reviser-assisted supervision.
Outcome: The proposed training framework outperforms strong preference optimization baselines on AlpacaEval-2 and Arena-Hard.
WhitenedCSE: Whitening-based Contrastive Learning of Sentence Embeddings (2023.acl-long)

Copied to clipboard

Challenge: Extensive experiments on seven semantic textual similarity tasks show our method achieves consistent improvement over the contrastive learning baseline and sets new states of the art.
Approach: They propose a whitening-based contrastive learning method for sentence embedding learning which combines contrastive and shuffled group whitening.
Outcome: The proposed method achieves better alignment and uniformity on seven semantic textual similarity tasks.
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for medical vision-language models overlook modality misalignment . HSCR generates high-quality preference data with higher sampling probability .
Approach: They propose a hierarchical self-contrastive reward approach that addresses two challenges in alignment . they leverage the inherent capability of Med-VLMs to generate dispreferred responses .
Outcome: The proposed approach improves accuracy and trustworthiness of medical vision-label models with 2,000 training entries.
LimaCost: Data Valuation for Instruction Tuning of Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Instruction tuning is an effective approach for aligning large language models with human intentions.
Approach: They propose a data quality measure that exhibits a strong correlation with model performance.
Outcome: The proposed measure exhibits a strong correlation with model performance.
RIFT: Repurposing Negative Samples via Reward-Informed Fine-Tuning (2026.findings-acl)

Copied to clipboard

Challenge: Reward Informed Fine-Tuning (RIFT) is an effective and robust alternative to expensive expert data for LLM alignment.
Approach: They propose a reward-informed fine-tuning framework that utilizes all self-generated samples to learn from both positive and negative trajectories.
Outcome: The proposed framework outperforms both RFT and Supervised Fine-Tuning (SFT) on mathematical benchmarks.
Dynamic Steering With Episodic Memory For Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing activation steering methods apply a single sentence-level steering vector uniformly across all tokens, ignoring LLMs’ token-wise, auto-regressive nature.
Approach: They propose a framework that aligns LLMs to given demonstrations by steering at the token level conditioned on the input query.
Outcome: The proposed framework surpasses baselines across safety, style transfer, and role-playing tasks, demonstrating improved alignment as demonstration scales.
Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation (2026.acl-long)

Copied to clipboard

Challenge: Recent systems exhibit conversational competence but lack structured mechanisms to evaluate adherence to core therapeutic principles.
Approach: They propose a framework to evaluate therapist-like responses for clinically grounded appropriateness and effectiveness using an ordinal scale.
Outcome: The proposed framework achieves an F-1 score of 63.34 versus the baseline Qwen3 score of 38.56 .
Learning Personalized Alignment for Evaluating Open-ended Text Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Traditional evaluation metrics rely heavily on lexical similarity with human-written references, showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences.
Approach: They propose an interpretable evaluation framework that evaluates alignment with specific human preferences by providing detailed comments and fine-grained scoring.
Outcome: The proposed framework outperforms GPT-4 in Kendall correlation and accuracy with zero-shot reviewers.
Enhancing Alignment using Curriculum Learning & Ranked Preferences (2024.findings-emnlp)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is an effective technique that leverages pairwise preference data to align LLMs to human preferences.
Approach: They propose to use pairwise preference data to create multiple preference pairs for a given prompt.
Outcome: The proposed method outperforms standard DPO on MTbench, Vicuna bench, and WizardLM with a score of 7.43 on the test sets.
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for merging large language models often overlook safety alignment during merging, leading to misaligned models.
Approach: They propose to combine safety and domain-specific data to optimize model merging techniques . they propose to use this data to maximize model alignment .
Outcome: The proposed method allows for models that excel in both domain expertise and alignment.
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to QA provide inaccurate answers but lack mechanisms to iteratively refine poor queries.
Approach: They propose a biomedical question answering agent that performs self-critic query refinement . they propose re-reflection methods that kick in only after full retrieval is completed .
Outcome: a biomedical question answering agent achieves 78.32% accuracy on PubMedQA . the proposed approach provides practical assistance to clinicians and biomedically researchers .
Curriculum Consistency Learning for Conditional Sentence Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Consistency learning (CL) has proven to be a valuable technique for improving the robustness of conditional sentence generation models.
Approach: They propose a strategy that guides models to learn consistency in alignment with their current capacity to differentiate between features.
Outcome: The proposed strategy delivers +2.0 accuracy point improvement compared with vanilla IT and +0.7 COMET scores over traditional CL methods in MT tasks.
Better Alignment with Instruction Back-and-Forth Translation (2024.findings-emnlp)

Copied to clipboard

Challenge: et al., 2023) proposes a method to improve instruction-tuning data . e.g., we generate synthetic instructions using the backtranslation approach .
Approach: They propose a method to improve instruction-tuning data using web-based inputs . they generate synthetic instructions using the backtranslation approach and filter the generated data .
Outcome: The proposed method improves the quality of instruction-tuning data based on preprocessed texts . it yields better AlpacaEval win rates than direct distillation .
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation (2025.findings-acl)

Copied to clipboard

Challenge: Existing automated singing annotation (ASA) methods tackle isolated aspects of the annotation pipeline.
Approach: They propose a framework that addresses transcription, alignment, and refined style annotations.
Outcome: The proposed framework delivers comprehensive multi-level annotations encompassing: (1) precise phoneme-audio alignment, (2) robust note transcription and temporal localization, (3) expressive vocal technique identification, and (4) global stylistic characterization including emotion and pace.
Pedagogical Alignment of Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often used without pedagogical fine-tuning and provide immediate answers rather than guiding students through the problem-solving process.
Approach: They propose a method for constructing large-scale preference datasets using synthetic data generation techniques that eliminates the need for manual annotation.
Outcome: The proposed methods outperform standard supervised fine-tuning (SFT) and improve alignment accuracy by 13.1% and 8.7% respectively.
Aligners: Decoupling LLMs and Alignment (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) need to be aligned with human expectations to ensure their safety and utility in most applications.
Approach: They propose to decouple LLMs and alignment by training *aligner* models that can be used to align any LLM on an as-needed basis.
Outcome: The proposed model can be used to align any LLM for a given criteria on an as-needed basis.
Joint Completion and Alignment of Multilingual Knowledge Graphs (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for knowledge graph completion are incomplete, as curators struggle to keep up with the real world.
Approach: They propose a multitask approach to solve missing facts in incomplete Knowledge Graphs . they add a relation representation to the existing KG embedding scheme .
Outcome: The proposed system outperforms existing models in seven languages compared to existing models . it also outperformed existing models, underscoring the value of joint alignment and completion.
Beyond Demographics: Aligning Role-playing LLM-based Agents Using Human Belief Networks (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing large language models can be prompted to role-play as individuals with particular demographic traits, but results are often human-like.
Approach: They found that seeding LLM-based agents with a single belief improved alignment . they say that role-playing based on demographic information does not improve alignment a .
Outcome: The proposed approach improves LLM alignment with human behavior . seeding agents with a single belief improves alignment for topics related to the belief network .
FRACTAL: Fine-Grained Scoring from Aggregate Text Labels (2025.acl-long)

Copied to clipboard

Challenge: Recent studies indicate that finer sentence or span-level labels provide more accurate and interpretable feedback for LLM optimization.
Approach: They propose a suite of models to disaggregate response-level labels into sentence-level (pseudo-)labels through Multiple Instance Learning and Learning from Label Proportions (LLP) formulations.
Outcome: The proposed model can reach 93% of the performance of the fully supervised baseline while requiring only around 10% of the gold labels.
The Price of Format: Diversity Collapse in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Instruction-tuned large language models employ structured templates to enforce format consistency during inference.
Approach: They fine-tune instruction-tuning large language models with structured templates and evaluate their results across three axes: downstream task performance, alignment behavior, and output diversity.
Outcome: The proposed model generates semantically similar outputs even under high temperature sampling and structural tokens in templates significantly constrain the model’s output space.
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Recent research has shown that carefully crafted jailbreak inputs can induce large language models to produce harmful outputs, despite safety measures such as alignment.
Approach: They propose a method for generating highly effective Jailbreak attacks that selectively strengthen or weaken attention among different parts of the prompt.
Outcome: The proposed attacks amplify the success rate of existing Jailbreak algorithms while lowering generation cost.
Representation Potentials of Foundation Models for Multimodal Alignment: A Survey (2025.emnlp-main)

Copied to clipboard

Challenge: foundation models learn highly transferable representations through large-scale pretraining on diverse data.
Approach: They examine the representation potentials of foundation models by examining their latent capacity to capture task-specific information within a single modality while providing a transferable basis for alignment and unification across modalities.
Outcome: The foundation models exhibit remarkable similarities across architectures and modalities, the authors show . the models can capture task-specific information within a single modality while providing a transferable basis for alignment and unification across modality.
Linguistic Alignment Predicts Learning in Small Group Tutoring Sessions (2025.findings-emnlp)

Copied to clipboard

Challenge: Cognitive science offers rich theories of learning and communication, yet these are often difficult to operationalize at scale.
Approach: They investigate linguistic alignment in a longitudinal dataset of real-world tutoring interactions and associated student test scores.
Outcome: The proposed method can be applied to real-world tutoring interactions and student test scores.
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs (2026.acl-long)

Copied to clipboard

Challenge: Recent self-training approaches have reduced reliance on human-labeled data, which limits their scalability.
Approach: They propose a team-based self-play algorithm that iteratively refines alignment without additional human supervision.
Outcome: The proposed algorithm outperforms baselines and LLM benchmarks in the self-supervised setting.
Improving Discriminative Capability of Reward Models in RLHF Using Contrastive Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Current methods rely on ranking losses to teach reward model to assess preferences, but they are susceptible to noise and ambiguous data, often failing to deeply understand human intentions.
Approach: They propose a method that incorporates contrastive learning into the reward modeling process to enhance generalization and stabilize the reinforcement learning training process.
Outcome: The proposed method enhances generalization of the reward model, stabilizes the reinforcement learning training process, and improves the final alignment with human preferences.
Profiling LLM’s Copyright Infringement Risks under Adversarial Persuasive Prompting (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models have demonstrated impressive capabilities in text generation but raise concerns regarding potential copyright infringement.
Approach: They propose a structured persuasion workflow to analyze the influence of persuasive prompts on LLM outputs.
Outcome: The proposed method analyzes the influence of persuasive prompts on LLM outputs.
BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on the use of Large Language Models (LLMs) focus primarily on English, leaving the cross-lingual generalization of aligned behavior underexplored.
Approach: They propose a structured generator-extractor pipeline and a multi-dimensional distributional analysis framework to examine how narrative attributes vary across languages, models, and social conditions.
Outcome: The proposed model reveals substantial cross-lingual variability in narrative generation patterns, indicating that distributions observed in English do not always exhibit similar characteristics in other languages, particularly in lower-resource settings.
CycleAlign: Iterative Distillation from Black-box LLM to White-box Models for Better Human Alignment (2024.findings-acl)

Copied to clipboard

Challenge: Existing language models that generate harmful responses are constrained by their inherent capability.
Approach: They propose to align large language models with human preferences from AI feedback.
Outcome: The proposed framework improves the alignment of large language models with human preferences from AI feedback.
ExPerT: Effective and Explainable Evaluation of Personalized Long-Form Text Generation (2025.findings-acl)

Copied to clipboard

Challenge: Evaluating personalized text generated by large language models is challenging, as only the LLM user, i.e. prompt author, can reliably assess the output.
Approach: They propose an explainable reference-based evaluation framework that leverages an LLM to extract atomic aspects and their evidences from the generated and reference texts, match the aspects, and evaluate their alignment based on content and writing style.
Outcome: The proposed framework achieves a 7.2% improvement in alignment with human judgments compared to the state-of-the-art evaluation methods.
HAF-RM: A Hybrid Alignment Framework for Reward Model Training (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have focused on enhancing reward models through data improvements, following the conventional training framework for reward models that directly optimizes the predicted rewards.
Approach: They propose a hybrid alignment framework **HAF-RM** that incorporates additional constraint on token-level policy probabilities in addition to the reward score.
Outcome: The proposed framework can supervise the internal preference model at the token level and optimize the mapping layer of the reward model at sequence level.
DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Inference-time alignment approaches still face limitations due to policy-specific value functions and latency during the inference phase.
Approach: They propose an efficient and policy-agnostic preference optimization method that avoids time latency associated with token generation.
Outcome: The proposed method achieves a favorable trade-off between alignment quality and inference-time latency.
LlmFixer: Fix the Helpfulness of Defensive Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Several defense strategies have been introduced to defend against jailbreak attacks, but these strategies weakened the usefulness of large language models.
Approach: They propose a framework that acts on large language models equipped with any defense strategy to recover their usefulness.
Outcome: The proposed framework can be used on large language models to recover their usefulness without updating the parameters of a defensive large language model.
Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection (2026.findings-acl)

Copied to clipboard

Challenge: Existing quantization methods focus on general metrics like perplexity or accuracy on standard benchmarks.
Approach: They propose a method that preserves fairness- and safety-critical weights during quantization.
Outcome: The proposed method reduces bias and safety degradation without costly retraining or alignment while maintaining trustworthiness while retaining efficiency.
EpiCaR: Knowing What You Don’t Know Matters for Better Reasoning in LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to improving reasoning abilities of large language models incur a significant calibration cost.
Approach: They propose an epistemic learning problem that integrates reasoning and calibration into an iterative supervised training framework.
Outcome: The proposed method achieves Pareto-superiority over standard baselines in accuracy and calibration.
Debiasing Reward Models via Causally Motivated Inference-Time Intervention (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches for mitigating spurious features in RMs focus on response length . Existing methods focus on RM activation, resulting in performance trade-offs .
Approach: They propose a method that uses neurons to suppress spurious features in RMs at inference time.
Outcome: The proposed method reduces sensitivity to spurious features without inducing performance trade-offs on RM benchmarks.
Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a transparent brain with accessible parameters that encode extensive knowledge, which can be analyzed, located and transferred.
Approach: They propose a new paradigm that aligns parametric spaces of LLMs using several training steps without following training.
Outcome: The proposed model aligns parametric spaces across scales using only training steps without following training.
Reinforcement Learning for Large Language Models via Group Preference Reward Shaping (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning Large Language Models (LLMs) are expensive and sensitive to reward model quality.
Approach: They propose a method that leverages preference-based comparisons rather than precise numerical rewards.
Outcome: Experiments show that GPRS outperforms critic-model-free RL algorithms on RLHF and reasoning tasks.
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) struggle with factual inaccuracies, a critical issue in clinical NLP applications where errors could lead to serious consequences.
Approach: They propose a pipeline that leverages >100B parameter GPT variants to act as synthetic experts to generate edit feedback without additional human annotations.
Outcome: The proposed pipeline aims to improve the quality of clinical note summarizations by generating edit feedback without human annotations.
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on table-based reasoning focus on a single gold table, not multiple tables . a persistent demand for robust table understanding systems is resulting from the complexity of table data .
Approach: They propose a MT-RAIG Bench to evaluate systems on Retrieval-Augmented Insight Generation over Mulit-Tables.
Outcome: The proposed framework improves human quality judgments on the generated insights.
Towards Aligning Language Models with Textual Feedback (2024.emnlp-main)

Copied to clipboard

Challenge: Using textual feedback, language models can be trained to learn from textual inputs.
Approach: They propose an approach that aligns language models with user preferences expressed in text.
Outcome: The proposed approach outperforms PPO on toxicity reduction, summarization, and dialog response tasks while achieving the same performance with only 20% of the samples.
Language Models Resist Alignment: Evidence From Data Compression (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) may exhibit undesirable behaviors due to the inevitable biases and harmful content present in training.
Approach: They propose to investigate the elasticity of large language models by examining their performance.
Outcome: The proposed model performance declines rapidly before reverting to the pre-training distribution, the authors show . the proposed model weight and code are available at pku-lm-res ist-alignment.github.io.
Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models focus on improving performance . however, language prior conflict leads to suboptimal vision-language alignment .
Approach: They propose a method to decouple the alignment process from language prior interference . they use a proxy LLM to detach from language interference during pretraining .
Outcome: The proposed method improves training performance and generalizes training data.
How do Language Models Reshape Entity Alignment? A Survey of LM-Driven EA Methods: Advances, Benchmarks, and Future (2025.emnlp-main)

Copied to clipboard

Challenge: Entity alignment (EA) is critical for knowledge graph (KG) integration.
Approach: They propose a taxonomy that categorizes methods in three stages: data preparation, feature embedding, and alignment.
Outcome: The proposed taxonomy categorizes methods in three key stages: data preparation, feature embedding, and alignment.
Enhancing Language Model Alignment: A Confidence-Based Approach to Label Smoothing (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have remarkable capabilities across various domains . Reinforcement Learning with Human Feedback (RLHF) phase is crucial for training . label smoothing is a technique that replaces hard labels with soft labels .
Approach: They propose a method that iteratively updates the label smoothing parameter based on preference labels and model forecasts.
Outcome: The proposed method improves the performance of large language models on state-of-the-art alignment tasks.
YinYang-Align: A new Benchmark for Competing Objectives and Introducing Multi-Objective Preference based Text-to-Image Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Recent controversies highlight the need for robust alignment mechanisms in text-to-image systems.
Approach: They propose a framework to evaluate T2I systems across six contradictory alignment objectives . objectives highlight key trade-offs such as artistic freedom and cultural sensitivity .
Outcome: The proposed framework achieves superior alignment across all objectives.
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Empirical evaluations on eight recent LLMs reveal that DRPO significantly enhances alignment performance, enabling base models to outperform their SFT/RLHF-tuned counterparts.
Approach: They propose a tuning-free approach to self-alignment called Dynamic Rewarding with Prompt Optimization (DRPO) it leverages a dynamic rewarding mechanism to identify and rectify alignment weaknesses .
Outcome: The proposed approach outperforms existing methods and is highly adaptable to various alignment challenges.
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
PrimeX: A Dataset of Worldview, Opinion, and Explanation (2025.emnlp-main)

Copied to clipboard

Challenge: Recent work shows that an individual's worldview -or beliefs about the overall character of the world -can explain persistent behavioral patterns and correlates with personality, well-being, political, religious, and demographic variables.
Approach: They develop a dataset of public opinion survey data from 858 US residents with written explanations from the respondents for why they hold specific opinions and the Primal World Belief survey for assessing respondent worldview.
Outcome: The proposed model can be used to better represent an individual's belief system and improve opinion prediction.
The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: persona prompting is increasingly used in large language models to simulate views of various sociodemographic groups.
Approach: They use open-source LLMs to study how persona prompts influence LLM simulations . they use role adoption formats and demographic priming strategies to study marginalized groups .
Outcome: The results show that the choice of demographic priming and role adoption strategy significantly impacts their portrayal.
Subjective Behaviors and Preferences in LLM: Language of Browsing (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) fuel expectations that a single trained model can effectively align with preferences of myriad users for a given task within a domain.
Approach: They introduce clusterwise LM training, HeTLM, appropriate for subjective behaviors . authors say small LM outperforms large pretrained LMs; heterogeneous cluster specific set of parameters outperformed single LM .
Outcome: The proposed model outperforms large pretrained or finetuned models in the domain of subjective behavior and preferences.
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have a global audience, so alignment must extend to cultural resonance.
Approach: They propose a framework that frames alignment as a conditional capacity separation problem.
Outcome: The proposed framework outperforms both dense baselines and semantic-only MoEs on three large language models.
HomoGraphAdapter: A Homogeneous Graph Neural Network as an Effective Adapter for Vision-Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing adaptation methods overlook structural knowledge between text and image modalities or create overly complex graphs containing redundant information for alignment.
Approach: They propose a method to adapt visual models to downstream tasks using text and image modalities.
Outcome: The proposed method improves classification accuracy by 1.51% for 1-shot and 0.74% for 16-shot on 11 datasets.
DeAL: Decoding-time Alignment for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are expected to generate content aligned with human preferences.
Approach: They propose a framework that allows the user to customize reward functions and enables Decoding-time Alignment of LLMs (DeAL).
Outcome: The proposed framework allows the user to customize reward functions and enables Decoding-time Alignment of LLMs.
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts (2025.emnlp-main)

Copied to clipboard

Challenge: Current large language models struggle to answer questions that span tens of thousands of tokens.
Approach: They evaluate 1–4 hop QA over 64k–128k-token excerpts from 83 novels . they find consistent accuracy drops with increased hops and context length .
Outcome: The novelhopqa benchmark evaluates 1–4 hop QA over 64k–128k-token excerpts from 83 public-domain novels.
MAPLE: Multi-Aspect Panels of LLM Evaluators for Open-Ended Questions (2026.findings-acl)

Copied to clipboard

Challenge: LLM-as-a-Judge uses LLMs to evaluate open-ended questions . however, the discrepancy between LLM generated evaluations and human evaluations remains a critical problem in this field .
Approach: They propose a framework that orchestrates evaluations across multiple criteria using multiple LLMs.
Outcome: The proposed framework achieves superior alignment with human evaluations compared to baselines.
T-REG: Preference Optimization with Token-Level Reward Regularization (2025.acl-long)

Copied to clipboard

Challenge: Reinforcement learning from human feedback (RLHF) is a dominant approach for large language models to follow instructions and produce meaningful alignment.
Approach: They propose a method that leverages human feedback to optimize large language models . they propose to use sequence-level and token-level rewards to optimize preference .
Outcome: The proposed method outperforms baseline methods on Alpaca Eval 2 and Arena-Hard benchmarks.
Rationalize and Align: Enhancing Writing Assistance with Rationale via Self-Training for Improved Alignment (2025.findings-acl)

Copied to clipboard

Challenge: Existing writing assistants rely on supervised fine-tuning to optimize models for multiple revisions.
Approach: They propose a framework that enhances WA performance with rationale and alignment.
Outcome: The proposed framework outperforms state-of-the-art WAs and the closed-source GPT-4o by 3.9 and 7.1 points on average across eight well-established writing-related test sets.
Segment, Embed, and Align: A Universal Recipe for Aligning Subtitles to Signing (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches for aligning spoken language text to sign language videos rely on end-to-end training tied to a specific language or dataset.
Approach: They propose a universal approach for aligning spoken language text with corresponding timestamps to sign language videos using a lightweight dynamic programming procedure.
Outcome: The proposed method can be used on four sign language datasets and is highly efficient on CPU.
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment (2025.acl-long)

Copied to clipboard

Challenge: Typical approaches to training large language models rely on limited contrasting patterns . contrasting data is limited and models are susceptible to harmful response tendencies .
Approach: They propose a framework that integrates contrasting patterns across the prompt, model, and pipeline levels.
Outcome: The proposed framework outperforms existing methods in the comparison of RQ1 and RQ2 . the proposed framework significantly outperformed existing methods, leading to more comprehensive alignment.
Towards Better Value Principles for Large Language Model Alignment: A Systematic Evaluation and Enhancement (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) show remarkable performance across tasks . alignment with human values is critical for their responsible development.
Approach: They propose a framework that evaluates value principles along three desirable properties . they propose supervised fine-tuning, reinforcement learning-based approaches .
Outcome: The proposed framework improves value principles along the three desirable properties of LLMs.
Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models (2026.acl-long)

Copied to clipboard

Challenge: Large language models are increasingly used as conversational agents that adopt personas and role-play characters at user request.
Approach: They propose to examine how persona agreeableness influences sycophancy across 13 small, open-weight language models ranging from 0.6B to 20B parameters.
Outcome: The proposed model consists of 275 personas and exposes them to 4,950 sycophancy-eliciting prompts spanning 33 topic categories.
Symmetrical Visual Contrastive Optimization: Aligning Vision-Language Models with Minimal Contrastive Images (2025.acl-long)

Copied to clipboard

Challenge: Recent studies have shown that Large Vision-Language Models (VLMs) tend to neglect image content and over-rely on language-model priors, resulting in errors in visually grounded tasks and hallucinations.
Approach: They propose a novel finetuning objective that steers the model toward capturing important visual details and aligning them with corresponding text tokens.
Outcome: The proposed method achieves up to 22% reduction in hallucinations and significant gains in vision-centric and general tasks while maintaining or improving the model's general abilities.
Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy (2026.findings-acl)

Copied to clipboard

Challenge: a limited amount of annotated data has slowed progress in machine learning for low-resource languages . a sentiment label records an annotator's final decision, but it is not a valid record of the annotation's interpretation.
Approach: They propose a large-scale Telugu sentiment classification dataset annotated with sentiment labels and human-selected rationales from multiple native speakers.
Outcome: The proposed model improves classification performance, explanation quality, and social bias by incorporating human rationales.
ReviewGrounder: Improving Review Substantiveness with Rubric-Guided, Tool-Integrated Agents (2026.acl-long)

Copied to clipboard

Challenge: Rapid rise in AI conference submissions has driven increasing exploration of large language models (LLMs) for peer review support.
Approach: They propose a peer review benchmarking tool based on paper-specific rubrics and a rubric-guided framework that decomposes reviewing into drafting and grounding stages.
Outcome: The proposed framework outperforms baselines with stronger/larger backbones in both alignment with human judgments and rubric-based review quality across 8 dimensions.
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly used as automated evaluators . et al., 2024: strong labels can foster trust but also undermine it .
Approach: They show that LLMs' source labels bias trust judgments by humans . they use eye-tracking data to analyze LLM internal states during judgment .
Outcome: The proposed model is biased by disclosed source labels, the authors show . eye-tracking data show humans rely heavily on source labels for judgments .
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work focused on improving alignment by refining the diffusion process, ignoring the role of the text encoder, which guides the diffusion.
Approach: They investigate how semantic information is distributed across token representations in text-to-image prompts by patching techniques to uncover encoding patterns.
Outcome: The proposed model can improve alignment and generation quality by modifying the diffusion stage and the cross-attention mechanism.
SpiralThinker: Latent Reasoning through an Iterative Process with Text–Latent Interleaving (2026.findings-acl)

Copied to clipboard

Challenge: Existing latent reasoning methods lack mechanisms to ensure stable reasoning dynamics in latent space and a systematic way to interleave implicit and explicit reasoning.
Approach: They propose a framework that performs iterative updates over latent representations while enabling interleaved reasoning across latent and textual steps.
Outcome: SpiralThinker achieves state-of-the-art among latent reasoning baselines.
ScholaWrite: A Dataset of End-to-End Scholarly Writing (2026.acl-long)

Copied to clipboard

Challenge: SCHOLAWRITE traces the multi-month journey from initial drafts to final manuscripts . authors demonstrate the value of capturing scientists’ cognitive writing process .
Approach: They present a dataset of end-to-end scholarly writing tracing the multi-month journey from initial drafts to final manuscripts.
Outcome: The first dataset of end-to-end scholarly writing traces the multi-month journey from initial drafts to final manuscripts over four months.
CityVG: Contrastive Fine-Tuning and Reward-Based Chain-of-Thought Reasoning for Zero-Shot City-Scale 3D Visual Grounding (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for 3D visual grounding are limited to small-scale indoor data or require heavy supervision.
Approach: They propose a contrastive fine-tuning strategy to align textual queries with urban scene graphs.
Outcome: The proposed framework achieves strong zero-shot localization performance and generalizes effectively to unseen urban environments.
Sounding vs. Being an Expert: Disentangling Authority, Register and Cultural Impact in Sycophantic LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models exhibit sycophancy, a tendency to align with user assertions even when they conflict with factual correctness.
Approach: They propose an adversarial evaluation framework that isolates two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
Outcome: The proposed framework disentangles two drivers of credibility: explicit authority (credentials) and implicit authority (linguistic register).
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization (2026.acl-long)

Copied to clipboard

Challenge: prevailing RLHF methods such as PPO and DPO depend on large-scale binary preference annotations.
Approach: They propose a method which converts natural feedback into continuous preference trajectories and optimizes them using the novel TraceBias algorithm.
Outcome: The proposed approach outperforms PPO and DPO in a variety of domains and improves alignment by up to 7.6% across diverse LLMs and preference domains.
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment (2025.emnlp-main)

Copied to clipboard

Challenge: Video-language models excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity.
Approach: They propose a framework that trains video-LLMs to distinguish accurate representations from carefully crafted adversarial examples.
Outcome: Experiments show that VideoPASTA improves performance without human annotation or captioning . the framework can be used on various state-of-the-art video-LLMs with no human annotation .
Massively Multilingual Joint Segmentation and Glossing (2026.acl-long)

Copied to clipboard

Challenge: Existing models generate morpheme-level glosses but assign them to whole words without predicting the actual morphological boundaries, making them less interpretable and therefore untrustworthy to human annotators.
Approach: They propose to use neural networks to predict interlinear glosses and morphological segmentation from raw text.
Outcome: The proposed model outperforms GlossLM on glossing and beats open-source models on segmentation, glossing, and alignment.
CARE: Multilingual Human Preference Learning for Cultural Awareness (2025.emnlp-main)

Copied to clipboard

Challenge: Language Models are tuned with human preferences to produce helpful responses, but the impact of preference tuning on the ability to handle culturally diverse queries remains understudied.
Approach: They introduce a multilingual resource that contains culturally specific questions and 31.7k responses with human judgments.
Outcome: The proposed model outperforms models with stronger initial cultural performance . the proposed model has gaps in the literature on culturally relevant data .
Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to align large language models with human preferences are limited in generalizability due to distribution shift, preference label noise, and mismatch of challenging samples with model capacity.
Approach: They propose a framework that constructs preference pairs with varying difficulty levels and then produces a specific curriculum for reward model training.
Outcome: The proposed framework improves generalizability of reward models by a significant margin without incurring additional inference costs compared to existing non-curriculum baselines.
SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition Inference (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for generating specifications are limited and often fail to infer semantic specifications such as pre-/postconditions.
Approach: They propose a framework that treats LLMs as exploratory reasoners rather than one-shot generators.
Outcome: The proposed framework outperforms state-of-the-art methods in accuracy and completeness of generated postconditions.
WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback (2026.acl-long)

Copied to clipboard

Challenge: Traditional alignment methods rely on human annotations and are subjective and misalignment with real-world user preferences.
Approach: They propose a framework that leverages in-situ user feedback during conversations with LLMs to create preference datasets automatically.
Outcome: The proposed framework identifies and classifies user feedback to LLM responses between conversation turns and creates examples of preferred and dispreferred responses according to user preferences.
Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another Begins (2026.acl-long)

Copied to clipboard

Challenge: Large language model agents exhibit action boundary blindness, granularity confusion, scope creep and boundary ambiguity . Explicit boundary prompting improves ABS by 0.08–0.13 across all models .
Approach: They propose four automatic metrics that require no human annotation to detect boundary blindness . they propose to use a multi-label attribution framework to validate the models .
Outcome: Experiments with seven large language model agents show that the best model achieves only 0.424 ABS . Explicit Boundary Prompting improves ABS by 0.08–0.13 across all models .
Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have found that speech language models fail to capture brain-relevant semantics beyond low-level features.
Approach: They analyze multimodal models to assess their alignment with MEG brain recordings . they find text embeddings from multimodal and unimodal models significantly outperform unilateral models .
Outcome: a new study shows that text-based models outperform unimodal models in alignment with brain recordings during naturalistic story listening.
When Cultures Meet: Multicultural Text-to-Image Generation (2026.findings-acl)

Copied to clipboard

Challenge: a new task to evaluate text-to-image generation models for multicultural scenes is unexplored.
Approach: They propose a benchmark task to evaluate text-to-image models in multicultural settings . they use a dataset of 9,000 images spanning five countries, three age groups, two genders, 25 historical landmarks, and five languages to analyze behavior .
Outcome: The proposed benchmark analyzes the behavior of state-of-the-art models across multiple dimensions including alignment, image quality, aesthetics, knowledge, and fairness.
Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing preference learning-based approaches rely on proprietary models to construct preference datasets, causing a distributional mismatch between the proprietary and target models.
Approach: They propose a framework that aligns LVLMs using in-distribution data derived from the model's intrinsic knowledge.
Outcome: The proposed framework surpasses baselines in hallucination mitigation while requiring only 5.2k samples.
Do LLM Agents Really Mimic Humans? Diagnosing and Aligning Microeconomic Behaviors in Macro-ABMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on replicating macro-level stylized facts while neglecting verification of micro-level decision-making.
Approach: They propose a framework that replicates macro-level stylized facts while ignoring micro-level decision-making.
Outcome: The proposed framework improves alignment with human trends and captures behavioral heterogeneity.
Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that deep vision-only and language-only models project inputs into a partially aligned representational space.
Approach: They investigate whether a model's representational code is semantically shared . they find that alignment peaks in mid-to-late layers of both model types .
Outcome: a forced-choice "Pick-a-Pic" task shows human preferences for image-caption matches are mirrored in embedding spaces across vision-language model pairs.
Alignment Data Map for Efficient Preference Data Selection and Diagnosis (2026.findings-acl)

Copied to clipboard

Challenge: constructing high-quality preference datasets faces scalability challenges due to prohibitive cost and complexity of human annotation.
Approach: They propose a tool to identify and select effective preference data by LLM-as-a-judge, explicit reward model, and reference-based approaches.
Outcome: The proposed tool reduces annotation costs while preserving alignment effectiveness.
Survey Response Generation: Generating Closed-Ended Survey Responses In-Silico with Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies focus on generating closed-ended survey responses with large language models, whereas LLMs are typically trained to generate open-ended text.
Approach: They evaluate the impact of various Survey Response Generation Methods on simulated responses by generating closed-ended responses from large language models.
Outcome: The proposed methods perform best in individual-level and subpopulation-level alignment.
Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment (2026.acl-long)

Copied to clipboard

Challenge: Existing methods assess suitability primarily through student likelihood, favoring trajectories that align closely with the student model’s current behavior but overlooking more informative ones.
Approach: They propose a Rank–Surprisal Ratio metric that captures both alignment and informativeness to assess the suitability of a reasoning trajectory.
Outcome: The proposed metric captures both alignment and informativeness to assess the suitability of a reasoning trajectory.
HypoEval: Hypothesis-Guided Evaluation for Natural Language Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing frameworks for LLM-as-a-judge use zero-shot setting without consulting any human input, which leads to low alignment, or fine-tune LLMs on labeled data, which requires a non-trivial number of samples.
Approach: They propose a hypothesis-guided evaluation framework that uses a small corpus of human evaluations to generate more detailed rubrics for human judgments and incorporates a checklist-like approach to combine LLM’s assigned scores on each decomposed dimension to acquire overall scores.
Outcome: The proposed framework outperforms existing frameworks in both human rankings and human scores with 30 human evaluations and fine-tunes LLMs on labeled data with 3 times more human evaluation by 11.95%.
AdaFuse: Adaptive Ensemble Decoding for Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing ensemble approaches to large language models lack flexibility for mid-generation adaptation.
Approach: They propose an adaptive ensemble decoding framework that dynamically selects semantically appropriate fusion units during generation.
Outcome: The proposed framework outperforms existing ensemble frameworks on open-domain QA, arithmetic reasoning, and machine translation tasks.
Thinking Alignment of Scenario-Oriented User Simulation (2026.acl-long)

Copied to clipboard

Challenge: Existing user simulators based on prompting to role-play or SFT focus on imitating textual utterances without considering multi-faceted cognitive processes that underlie human decision-making during interactions.
Approach: They construct a user-simulator dataset that augments 51k human–LLM conversations by reconstructing the user’s inner reasoning during and at the end of each dialogue.
Outcome: The proposed user simulators augment 51k human–LLM conversations by reconstructing the user’s inner reasoning both during and at the end of each dialogue.
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging (2026.findings-acl)

Copied to clipboard

Challenge: We show that the "alignment tax" of post-training is framed as a drop in task accuracy.
Approach: They propose a more holistic view of the alignment tax by framing it as a drop in accuracy and a degradation of model calibration.
Outcome: The proposed method improves accuracy beyond both parents while recovering calibration lost during alignment.
Selective Contrastive Learning For Gloss Free Sign Language Translation (2026.acl-long)

Copied to clipboard

Challenge: Recent SLT systems adopt CLIP-like Vision-Language pretraining, but the random in-batch contrast provides few, batch-dependent negatives.
Approach: They propose a method to train sign video-text similarity over a time period of 3 months . they use a random in-batch contrast strategy to track negative video- text similarity .
Outcome: The proposed system improves sign language translation by focusing on challenging negatives . the results show that the random in-batch contrast provides few negatives and noisy supervision .
CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations (2026.findings-acl)

Copied to clipboard

Challenge: Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment.
Approach: They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation.
Outcome: The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations.
Overcoming Copyright Barriers in Corpus Distribution Through Non-Reversible Hashing (2026.acl-long)

Copied to clipboard

Challenge: Annotated corpora are crucial in the field of natural language processing, but are difficult to exchange among researchers.
Approach: They propose a method to lawfully share the annotations of any sequential copyrighted corpus.
Outcome: The proposed method is robust to reasonable divergences in the version of the copyrighted data owned by the user.
Influence-based Online Experience Selection for Effective RLHF (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for RL fail to establish an interpretable connection between data and optimization objectives.
Approach: They propose a data selection method that dynamically estimates the influence of individual training samples on policy optimization.
Outcome: The proposed method significantly improves training effectiveness with fewer optimization steps.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations