Papers with understanding

113 papers
From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning, Efficiency and beyond (2024.lrec-tutorials)

Copied to clipboard

Challenge: This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs.
Approach: This tutorial will review cutting-edge research in MLLMs and examine the impact of ML in learning and reasoning.
Outcome: This course will review cutting-edge research in MLLMs and examine the impact of ML models on learning, learning, and multimodal reasoning.
Towards Generation and Recognition of Humorous Texts in Portuguese (2023.eacl-srw)

Copied to clipboard

Challenge: This PhD thesis focuses on the automatic generation and recognition of verbal punning humor in Portuguese.
Approach: They propose to combine natural language generation and cognitive processing to generate and recognize verbal humor in Portuguese.
Outcome: The proposed methods aim to generate and recognize humor in Portuguese, an underdeveloped language compared to English.
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research (2025.acl-demo)

Copied to clipboard

Challenge: Language agents powered by large language models (LLMs) have demonstrated remarkable capabilities in understanding, reasoning, and executing complex tasks.
Approach: They propose a flexible framework that addresses engineering overhead and insufficient evaluation frameworks for fair comparison.
Outcome: The proposed framework simplifies language agent development and establishes a foundation for reproducible agent research.
Multi-task Learning of Spoken Language Understanding by Integrating N-Best Hypotheses with Hierarchical Attention (2020.coling-industry)

Copied to clipboard

Challenge: Existing methods to integrate hypotheses into speech recognition systems are noisy and can cause information loss.
Approach: They propose to integrate hypotheses into multi-task learning and transfer learning to improve performance.
Outcome: The proposed model improves domain and intent classification by 19% and 37% compared to current methods . the proposed model could recover transcription and rewrite the query for a better understanding .
FlowDelta: Modeling Flow Information Gain in Reasoning for Conversational Machine Comprehension (D19-58)

Copied to clipboard

Challenge: Existing machine comprehension models focus on a single-turn setting and do not account for previous reasoning processes.
Approach: They propose to explicitly model the information gain through the dialogue reasoning . they propose to apply the proposed mechanism to other machine comprehension models .
Outcome: The proposed model achieves state-of-the-art performance in a conversational QA dataset QuAC and a sequential instruction understanding dataset SCONE.
MFinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset (2025.findings-acl)

Copied to clipboard

Challenge: Existing financial benchmarks rely on news articles, earnings reports, or announcements, making it challenging to capture the real-world dynamics of financial meetings.
Approach: They propose a multilingual, multi-sector, and multi-task dataset called MFinMeeting that supports English, Chinese, and Japanese .
Outcome: The proposed benchmark supports English, Chinese, and Japanese, enhancing comprehension of financial discussions in diverse linguistic contexts.
Inconsistent dialogue responses and how to recover from them (2024.findings-eacl)

Copied to clipboard

Challenge: Existing methods to assess and bolster utterance consistency of chat systems have been shown difficult to detect.
Approach: They propose to use annotators to write dialogue responses and recovery utterances to assess and bolster utteration consistency of chat systems.
Outcome: The proposed dataset significantly improves the detection and resolution of inconsistencies in chat conversations.
Language Models Understand Us, Poorly (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent large language models have achieved impressive results on benchmark tasks.
Approach: They examine three views of human language understanding: as-mapping, as-reliability and as-representation.
Outcome: The authors argue that language models are inadequate and that they can't understand us . they also argue that as-representation advances a science of understanding .
Measuring Alignment Bias in Neural Seq2seq Semantic Parsers (2022.starsem-1)

Copied to clipboard

Challenge: Sequence-to-sequence semantic parsers with attention mechanisms have changed the research landscape . emergence of seq2seq models have led to questions about alignments .
Approach: They investigate whether seq2seq models can handle both simple and complex alignments.
Outcome: The proposed model performs better on monotonic and complex alignments compared to monotonic models .
Walia-LLM: Enhancing Amharic-LLaMA by Integrating Task-Specific and Generative Datasets (2024.findings-emnlp)

Copied to clipboard

Challenge: Low-resource languages are left behind due to the unavailability of resources.
Approach: They propose to integrate task-specific and generative datasets to improve language model performance for Amharic by fine-tuning an Amharican instruction fine-to-tuned model.
Outcome: The proposed model shows promising results in different NLP tasks and compares translated instruction datasets with the original model.
A Trip Towards Fairness: Bias and De-Biasing in Large Language Models (2024.starsem-1)

Copied to clipboard

Challenge: a little or a large bias in CtB-LLMs may cause huge harm . LLaMA and OPT families have an important bias in gender, race, religion, and profession.
Approach: They propose to debiase three families of Very Large-Language Models with LORA to reduce bias by 4.12 points in the normalized stereotype score.
Outcome: The proposed model reduces bias up to 4.12 points in the normalized stereotype score.
Automating Horizon Scanning in Future Studies (2022.lrec-1)

Copied to clipboard

Challenge: Existing studies collect enough information to predict drastic social changes in the mid- or long-term future.
Approach: They propose document retrieval and comment generation tasks for automating horizon scanning by analyzing a dataset that contains 2,266 manually collected news articles with comments written by experts.
Outcome: The proposed tasks are more efficient than previous methods and the proposed models are more accurate.
Generation-driven Contrastive Self-training for Zero-shot Text Classification with Instruction-following LLM (2024.eacl-long)

Copied to clipboard

Challenge: a novel method to train a smaller model with LLMs for zero-shot text classification requires immense computational resources due to their substantial model size.
Approach: They propose a method which leverages the generative power of large language models to train a smaller model.
Outcome: The proposed method outperforms state-of-the-art methods when limited data is available.
Do Video Language Models really understand the video contexts? (2025.naacl-srw)

Copied to clipboard

Challenge: Recent advances in VideoQA performance have shown that visual language models are effective but the processes of understanding and reasoning in VLMs remain under-explored.
Approach: They propose a framework that incorporates a fine-grained question generation and answering process to measure how well VLMs understand video question answering tasks.
Outcome: The proposed framework incorporates a fine-grained question generation and answering process to measure how well the responses generated by VLMs align with what the model understands.
QUILL: Query Intent with Large Language Models using Retrieval Augmentation and Multi-stage Distillation (2022.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive results on a variety of text understanding tasks.
Approach: They propose a two-stage distillation approach that allows retrieval augmentation to be carried over without the increased compute associated with it.
Outcome: The proposed approach can carry over the gains of retrieval augmentation without suffering the increased compute typically associated with it.
Sensitivity as a Complexity Measure for Sequence Classification Tasks (2021.tacl-1)

Copied to clipboard

Challenge: Existing complexity metrics provide limited practical insight into complexity differences between tasks.
Approach: They propose a theoretical framework for understanding and predicting the complexity of sequence classification tasks using a new extension of the theory of Boolean function sensitivity.
Outcome: The proposed framework predicts the complexity of sequence classification tasks using a new method . it shows that low-sensitivity functions are easier to learn for LSTMs than lexical classifiers .
PromptPrism: A Linguistically-Inspired Taxonomy for Prompts (2026.findings-eacl)

Copied to clipboard

Challenge: PromptPrism is a linguistically-inspired taxonomy that enables prompt analysis across three hierarchical levels.
Approach: They propose a linguistically-inspired taxonomy that enables prompt analysis across three hierarchical levels: functional structure, semantic component, and syntactic pattern.
Outcome: The proposed taxonomy bridges traditional language understanding with modern LLM research . it improves prompt quality and improves model performance across tasks .
Testing the Effect of Code Documentation on Large Language Model Code Understanding (2024.findings-naacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive abilities in recent years with regards to code generation and understanding.
Approach: They propose to provide an LLM with "incorrect" documentation that can greatly hinder code understanding, while incomplete or missing documentation does not seem to significantly affect an LRM's ability to understand code.
Outcome: The proposed model can generate and understand code in a language with high documentation quality while lacking documentation does not significantly affect the ability to understand code.
Self-Debiasing Large Language Models: Zero-Shot Recognition and Reduction of Stereotypes (2025.naacl-short)

Copied to clipboard

Challenge: Large language models exhibit harmful social biases, but they are often difficult to train and modify.
Approach: They leverage the zero-shot capabilities of large language models to reduce stereotyping . they introduce a technique called zero- shot self-debiasing to reduce bias .
Outcome: The proposed technique reduces stereotyping across nine different social groups while relying on the LLM itself and a simple prompt.
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
Reconfidencing LLMs from the Grouping Loss Perspective (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to calibrate confidence scores for large language models often overlook biases towards certain groups, such as specific nationalities.
Approach: They propose a method to calibrate confidence scores of Large Language Models by considering different groups, a process they call reconfidencing.
Outcome: The proposed method mitigates biases against minority groups, the authors show . they show that the proposed method is more reliable than existing methods .
RewardBench: Evaluating Reward Models for Language Modeling (2025.findings-naacl)

Copied to clipboard

Challenge: Evaluating reward models presents an opportunity to understand the opaque technologies used for alignment of language models.
Approach: They present a benchmark dataset and code-base for evaluation of reward models . they use prompt-chosen-rejected trios to benchmark how they perform on queries .
Outcome: The proposed dataset compares RMs with other models on a set of questions.
Quantifying the Contextualization of Word Representations with Semantic Class Probing (2020.findings-emnlp)

Copied to clipboard

Challenge: Pretrained language models are effective in solving NLP tasks, but there are still questions about how and why they work so well.
Approach: They use BERT to quantify contextualization by studying the extent of inference . they show that top layer representations support highly accurate inference of semantic classes .
Outcome: The proposed model is highly accurate, but weak in the lower layers . it is more task-specific after finetuning while lower layers are more transferable .
Learning to Refuse: Towards Mitigating Privacy Risks in LLMs (2025.coling-main)

Copied to clipboard

Challenge: Large language models (LLMs) exhibit remarkable capabilities in understanding and generating natural languages, but can inadvertently memorize private information, posing significant privacy risks.
Approach: They propose to use a dataset to evaluate machine unlearning methods for protecting personal data in a realistic scenario.
Outcome: The proposed model outperforms baseline methods by 5.65 points and protects target individuals’ personal data while maintaining general capabilities.
PuzzLing Machines: A Challenge on Learning From Small Data (2020.acl-main)

Copied to clipboard

Challenge: a benchmark dataset of 81 languages is released to test deep neural models' human-like reasoning and generalization skills.
Approach: They propose a challenge on learning from small data using Rosetta Stone puzzles from Linguistic Olympiads for high school students.
Outcome: The proposed benchmark consists of Rosetta Stone puzzles from Linguistic Olympiads for high school students.
Thinking beyond the anthropomorphic paradigm benefits LLM research (2026.acl-long)

Copied to clipboard

Challenge: anthropomorphism is an automatic and unconscious response that occurs even in advanced technical expertise.
Approach: They argue that anthropomorphism is an automatic and unconscious response . they identify and examine five assumptions that shape research across the LLM development lifecycle .
Outcome: The proposed framework challenges assumptions that shape research across the LLM development lifecycle and offers promising directions for LLMs.
G-SPEED: General SParse Efficient Editing MoDel (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated incredible capabilities in understanding, generating, and manipulating languages.
Approach: They propose a general SParse Efficient Editing MoDel which can fulfill diverse editing requirements through a single model while maintaining low computational costs.
Outcome: The proposed model can fulfill diverse editing requirements through a single model while maintaining low computational costs.
Measuring Context-Word Biases in Lexical Semantic Datasets (2022.emnlp-main)

Copied to clipboard

Challenge: Existing pretrained contextualized models have been used to evaluate word-in-context representations in many lexical semantic tasks.
Approach: They propose to quantify the degree of context or word biases in existing datasets by probing masked input.
Outcome: The proposed model performs better when both word and context are available than with masked input.
The Two Paradigms of LLM Detection: Authorship Attribution vs Authorship Verification (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting texts generated by large language models are disputed . authors argue that there are limitations in the current technology .
Approach: They propose to make LLM detectors robust against domain shifts and build benchmarks . they argue that the limitations lie elsewhere, and open the realm of authorship analysis technology .
Outcome: The proposed method systematically analyzes the benchmarks and validates it using state-of-the-art detectors.
GraPPI: A Retrieve-Divide-Solve GraphRAG Framework for Large-scale Protein-protein Interaction Exploration (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models and Retrieval-Augmented Generation frameworks have accelerated drug discovery, but integrating models into workflows remains challenging.
Approach: They propose a large-scale knowledge graph-based retrieve-divide-solve agent pipeline RAG framework to support large-level PPI signaling pathway exploration.
Outcome: The proposed framework is based on large-scale knowledge graphs and can be used to analyze protein-protein interactions.
MentalManip: A Dataset For Fine-grained Analysis of Mental Manipulation in Conversations (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on mental manipulation focus on context-free content and face challenges in identifying implicit toxicity.
Approach: They propose a dataset that analyzes mental manipulation and its components . they propose to use 4,000 fictional dialogues to identify the techniques utilized for manipulation .
Outcome: The proposed dataset enables a comprehensive analysis of mental manipulation . it shows that leading-edge models inadequately identify and categorize manipulative content .
Faithfulness-Aware Decoding Strategies for Abstractive Summarization (2023.eacl-main)

Copied to clipboard

Challenge: Existing studies on faithfulness of abstractive summarization have focused on decoding strategies.
Approach: They propose two faithfulness-aware generation methods to further improve faithfulness . they propose to use a distillation approach to generate faithful summaries with greedy decoding .
Outcome: The proposed methods improve faithfulness across two datasets as evaluated by automatic faithfulness metrics and human evaluation.
Sparse Latents Steer Retrieval-Augmented Generation (2025.acl-long)

Copied to clipboard

Challenge: In this study, we uncover interpretable latents that govern RAG behavior in large language models . Sparse Autoencoders are used to control large language model (LLM) behavior .
Approach: They leverage Sparse Autoencoders within the LLaMA Scope to uncover latents that govern RAG behaviors.
Outcome: The proposed model can be used to control large language models without architectural modifications.
MemoReader: Large-Scale Reading Comprehension through Neural Memory Controller (D18-1)

Copied to clipboard

Challenge: Existing approaches to machine reading comprehension are limited in understanding, up to a few paragraphs, failing to comprehend lengthy documents.
Approach: They propose a deep neural network architecture to handle a long-range dependency in RC tasks.
Outcome: The proposed method outperforms existing methods especially for lengthy documents.
FFE-Hallu: Hallucinations in Fixed Figurative Expressions: A Benchmark of Idioms and Proverbs in the Persian Language (2026.eacl-long)

Copied to clipboard

Challenge: Figurative language, especially fixed figurative expressions, poses unique challenges for large language models . Unlike literal phrases, FFEs are culturally grounded and often non-compositional, making them vulnerable to figurativ hallucination .
Approach: They propose a benchmark to evaluate LLMs' ability to generate, detect, and translate fixed figurative expressions in Persian.
Outcome: The proposed benchmarks show that LLMs still struggle with figurative language expressions . the benchmarks are based on 600 carefully curated examples spanning three tasks .
MultiMET: A Multimodal Dataset for Metaphor Understanding (2021.acl-long)

Copied to clipboard

Challenge: Metaphor is a linguistic phenomenon and a cognitive phenomenon structuring human thought, authors say . previous studies focused on texts, partly due to the unavailability of ground truth labels of multimodal metaphor .
Approach: They propose a multimodal metaphor dataset that integrates multimodal text and image . it contains 10,437 text-image pairs with multimodal annotations of occurrences .
Outcome: The proposed dataset examines multimodal cues and their interplay.
Knowledge Graph-Driven Memory Editing with Directional Interventions (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are hampered by inaccuracies and outdated information.
Approach: They propose a framework that constructs knowledge graphs using available information to guide the direction of knowledge editing.
Outcome: The proposed framework allows consistent, aligned, and stable information during large-scale editing scenarios.
Constructing Emotional Consensus and Utilizing Unpaired Data for Empathetic Dialogue Generation (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for dialogue empathy focus on the emotion flow in one direction, from context to response.
Approach: They propose a dual-generative model to construct emotional consensus and use unpaired data to produce pseudo paired empathetic samples.
Outcome: The proposed model outperforms baseline models in producing coherent and empathetic responses.
Automated Fact Checking: Task Formulations, Methods and Future Directions (C18-1)

Copied to clipboard

Challenge: Recent research on fact checking has focused on misinformation . however, relevant papers and articles have been published in research communities that are unaware of each other and use inconsistent terminology.
Approach: They propose avenues for future NLP research on automated fact checking . they highlight the use of evidence as an important distinguishing factor .
Outcome: The proposed methods unify the task formulations and methodologies across papers and authors.
Your Co-Workers Matter: Evaluating Collaborative Capabilities of Language Models in Blocks World (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on how large language model agents collaborate with humans in equal roles emphasize the importance of coordination and communication.
Approach: They propose to use chain-of-thought prompts to evaluate different collaboration perspectives, from independent to more complex, dependent tasks.
Outcome: The proposed model significantly improves the evaluation metric.
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams (2026.findings-eacl)

Copied to clipboard

Challenge: a growing need for tools that support legal education, especially in under-resourced languages such as Romanian . we evaluate the capabilities of large language models and vision-language models in legal education .
Approach: They evaluate the capabilities of Large Language Models and Vision-Language Models in Romanian driving law through textual and visual question-answering tasks.
Outcome: The proposed model improves retrieval performance and QA accuracy in Romanian driving tests.
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing research has not explored the joint task of emotion detection and explanatory span identification in e-commerce reviews.
Approach: They propose a joint task unifying Emotion detection and Opinion Trigger extraction (EOT) which explicitly models the relationship between causal text spans (opinion triggers) and affective dimensions (emotion categories).
Outcome: The proposed framework surpasses zero-shot and chain-of-thought techniques across e-commerce domains.
Digital Socrates: Evaluating LLMs through Explanation Critiques (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can provide reasoned explanations, but the nature and quality of those explanations are still poorly understood.
Approach: They propose to define a task of explanation critiquing and train an open-source automatic critique model using this data.
Outcome: The proposed model can provide high-quality, nuanced evaluations without expensive API calls or human annotations.
Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Recent studies show that Large language models struggle with handling long token sequences due to limited training context size.
Approach: They propose a single-stage continual pretraining method to equip LLMs with long context modeling capabilities.
Outcome: The proposed method outperforms existing methods on 4 language modeling benchmarks.
DeepRTL2: A Versatile Model for RTL-Related Tasks (2025.findings-acl)

Copied to clipboard

Challenge: Integration of large language models into electronic design automation has been a key driver in eDA.
Approach: They propose a family of large language models that unifies generation- and embedding-based tasks related to RTL.
Outcome: The proposed model achieves state-of-the-art performance across all evaluated tasks.
LegalViz: Legal Text Visualization by Text To Diagram Generation (2025.naacl-long)

Copied to clipboard

Challenge: Graphviz provides diagrams for legal documents that are easy to understand and understand . a novel dataset of 23 languages and 7,010 cases of legal document and visualization pairs is proposed .
Approach: They propose a dataset of legal diagrams using DOT graph description language of Graphviz.
Outcome: The proposed dataset outperforms existing models including GPTs in 23 languages and 7,010 cases of legal document and visualization pairs.
Expected Validation Performance and Estimation of a Random Variable’s Maximum (2021.findings-emnlp)

Copied to clipboard

Challenge: In this paper, we analyze three statistical estimators for expected validation performance . Often researchers only report the performance of the best-found model during a hyperparameter search .
Approach: They analyze three estimators for expected validation performance to compare models . they find that the estimator with the smallest variance has the largest bias .
Outcome: The proposed model has the highest variance and the estimator with the smallest variance has the largest bias.
Evaluating LLMs for Targeted Concept Simplification for Domain-Specific Texts (2024.emnlp-main)

Copied to clipboard

Challenge: Simplifying the entire text makes it understandable but sometimes removes important details.
Approach: They propose a simplification task for rewriting text to help readers comprehend text containing unfamiliar concepts and introduce a dataset of 22k definitions from 13 academic domains paired with a difficult concept within each definition.
Outcome: The proposed model outperforms open-source and commercial models on the task and human judges prefer explanations over simplifications of the difficult concept.
Event Representation with Sequential, Semi-Supervised Discrete Variables (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods for event modeling take discrete, external knowledge into account . obtaining fully accurate structured knowledge can be difficult .
Approach: They propose a method that takes partially-observed sequences of discrete, external knowledge into account.
Outcome: The proposed method outperforms baselines and state-of-the-art in script induction and converges faster.
PLUG: Leveraging Pivot Language in Cross-Lingual Instruction Tuning (2024.acl-long)

Copied to clipboard

Challenge: Instruction tuning has advanced large language models (LLMs) but its application in lower-resource languages faces challenges due to the imbalanced foundational abilities of LLMs across different languages.
Approach: They propose a pivot language guided generation approach that utilizes a high-resource language as the pivot to enhance instruction tuning in lower-resourced languages.
Outcome: The proposed approach improves instruction-following abilities of LLMs by 29% on average compared to directly responding in the target language alone.
Large Language Models Are No Longer Shallow Parsers (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models (LLMs) have reshaped the field of natural language processing (NLP) however, fundamental NLP tasks that involve linguistic analysis still play essential roles in the field.
Approach: They propose to use constituency parsing to improve performance of LLMs on deep syntactic parse trees to prompt LLM chunking, filter out low-quality chunks and add remaining chunks to prompts to instruct LLM for parser.
Outcome: The proposed approach improves LLMs' performance on constituency parsing on English and Chinese benchmark datasets.
LILA: A Unified Benchmark for Mathematical Reasoning (2022.emnlp-main)

Copied to clipboard

Challenge: Towards evaluating and improving AI systems in this domain, we propose a mathematical reasoning benchmark based on 23 diversetasks .
Approach: They propose a mathematical reasoning benchmark that includes 23 diverse tasks . they extend the benchmark by collecting task instructions and solutions in the form of Python programs .
Outcome: The proposed model improves on multi-tasking while the best performing model only achieves 60.40%.
Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Temporal reasoning is a vital component of human communication and understanding, yet remains an underexplored area within the context of Large Language Models (LLMs).
Approach: They propose to use 3 prompting strategies to evaluate 8 different LLMs across 6 datasets and 2 Code Generation LMs to perform the analysis.
Outcome: The proposed models perform better on NLP tasks than the standard models on the same dataset.
Probing Contextual Language Models for Common Ground with Visual Representations (2021.naacl-main)

Copied to clipboard

Challenge: Contextual language models have attracted great interest in probing what is encoded in their representations.
Approach: They propose a probing model that evaluates how effective are text-only representations in distinguishing between matching and non-matching visual representations.
Outcome: The proposed model outperforms text-only language models in instance retrieval, but underperform humans.
Do Zombies Understand? A Choose-Your-Own-Adventure Exploration of Machine Cognition (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in LLMs have sparked a debate on whether they understand text.
Approach: They propose two working definitions for understanding which explicitly acknowledge the question of consciousness and draw connections with a rich literature in philosophy, psychology and neuroscience.
Outcome: The proposed models achieve impressive results on various benchmarks, seeming to generalize to unseen tasks and domains.
Rewarding What Matters: Step-by-Step Reinforcement Learning for Task-Oriented Dialogue (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing RL methods focus on generation tasks while neglecting dialogue state tracking (DST) for understanding.
Approach: They propose a method that integrates RL into both understanding and generation tasks by introducing step-by-step rewards throughout the token generation.
Outcome: The proposed approach achieves state-of-the-art results on three widely used datasets.
GupShup: Summarizing Open-Domain Code-Switched Conversations (2021.emnlp-main)

Copied to clipboard

Challenge: Abstractive summarization is the process of generating a condensed version of a given conversation while preserving the most salient aspects.
Approach: They propose to use a dataset to analyze code-switched conversations in Hindi and English to summarize them.
Outcome: The proposed dataset contains over 6,800 code-switched conversations and their corresponding human-annotated summaries in English (En) and Hi-En.
Transformers to Learn Hierarchical Contexts in Multiparty Dialogue for Span-based Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to embedding in multiparty dialogues are poor for span-based question answering (QA)
Approach: They propose a novel approach to transformers that learns hierarchical representations in multiparty dialogue.
Outcome: The proposed model improves on the FriendsQA dataset by 3.8% and 1.4% over the two state-of-the-art models.
Improving Large Language Models in Event Relation Logical Prediction (2024.acl-long)

Copied to clipboard

Challenge: Event relation extraction tasks require rigorous logical reasoning and semantic comprehension, a challenge for narrative understanding and reasoning.
Approach: They propose three approaches to endow LLMs with event relation logic to generate more coherent answers across different scenarios.
Outcome: The proposed approach improves on a set of ERE tasks and provides insights for future work.
PropSegmEnt: A Large-Scale Corpus for Proposition-Level Segmentation and Entailment Recognition (2023.findings-acl)

Copied to clipboard

Challenge: Existing systems for Natural Language Inference (NLI) only recognize textual entailment relations on sentence-level . however, even a simple sentence often contains multiple propositions, i.e. distinct units of meaning conveyed by the sentence .
Approach: They propose a system to recognize whether one text is textually entailed by another . they use a corpus of over 45K propositions annotated by human raters to study the textual entailment relation of each proposition in a sentence individually.
Outcome: The proposed dataset can be used to understand the compositionality of NLI labels.
Coherence boosting: When your pretrained language model is not paying enough attention (2022.acl-long)

Copied to clipboard

Challenge: Long-range semantic coherence remains a challenge in automatic language generation and understanding.
Approach: They propose a procedure that increases a model’s focus on a long context by distributional analyses of generated ordinary text and dialog responses.
Outcome: The proposed procedure increases the model's focus on a long context.
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for video comment art are constrained by their limited modalities and insufficient categories, hindering creativity in video-based comment art creation.
Approach: They propose a benchmark that integrates video and text modalities to evaluate MLLMs’ abilities to compose video Comment art.
Outcome: The proposed framework integrates video and text modalities to evaluate MLLMs’ abilities to compose video comment art.
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have raised concerns regarding their intrinsic values.
Approach: They propose a psychologically grounded five-factor value system for Large Language Models that integrates psychological principles with cutting-edge AI priorities.
Outcome: The proposed value system meets standard psychological criteria, improves LLM safety prediction, and enhances Llm alignment, when compared to the canonical Schwartz’s values.
On General Language Understanding (2023.findings-emnlp)

Copied to clipboard

Challenge: a recent paper suggests that the evidence underspecifies the understanding of large language models.
Approach: They propose to use a "general language understanding" benchmark to examine what it could mean in machines.
Outcome: The proposed model can be used to ground questions of the adequacy of benchmarking methods.
CHEER: Centrality-aware High-order Event Reasoning Network for Document-level Event Causality Identification (2023.acl-long)

Copied to clipboard

Challenge: Recent studies focus on building a document-level graph for cross-sentence reasoning, but ignore important causal structures.
Approach: They propose a document-level event causality identification model which annotates central events and incorporates event centrality information into the reasoning network.
Outcome: The proposed model performs high-order reasoning while considering event centrality.
Extracting Biomedical Entities from Noisy Audio Transcripts (2024.lrec-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is particularly affected by noise, often termed the ASR-NLP gap.
Approach: They propose a dataset to bridge the ASR-NLP gap in the biomedical domain by extracting adverse drug reactions and mentions of entities from the Brief Test of Adult Cognition by Telephone (BTACT) exam.
Outcome: The proposed method can clean 2,000 clean and noisy recordings and eliminate errors using zero-shot and few-shot methods.
A Review of Prominent Paradigms for LLM-Based Agents: Tool Use, Planning (Including RAG), and Feedback Learning (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models have been used for planning, tool use, and feedback learning . inconsistent taxonomy and complexity of workflows create challenges .
Approach: They propose a unified taxonomy to review and discuss the three paradigms . they define environments/tasks, common LLM-profiled roles and universally applicable workflows based on prior work .
Outcome: The proposed taxonomy compares LMPR implementations and workflow usage across paradigms . large language models have human-like reasoning capabilities, the authors say .
On the Calibration of Large Language Models and Alignment (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models are becoming more popular and are proving to be reliable . however, their reliability is often understudied due to their uncertainty and complex structure .
Approach: They conduct a systematic examination of the calibration of aligned language models throughout the entire construction process including pretraining and alignment training.
Outcome: The results shed light on whether popular large language models are well-calibrated and how the training process influences model calibration.
Bridging Context Gaps: Enhancing Comprehension in Long-Form Social Conversations Through Contextualized Excerpts (2025.coling-main)

Copied to clipboard

Challenge: a recent rise in polarization has led to a rise in the use of loud and extreme voices in public spaces.
Approach: They propose ways to parse and convey information from small-group recorded conversations . they show that LLMs can provide socially relevant context to improve comprehension .
Outcome: The proposed models improve comprehension, readability, and empathy in small-group conversations . the proposed models struggle with capturing key social aspects, the authors show .
MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Social media is a key platform for emotional expression, yet deep learning lacks flexibility and interpretability.
Approach: They propose to use Chinese social media to train interpretable mental health instruction datasets to test models' ability to explain their decisions.
Outcome: The proposed models outperform deep learning and LLMs on three mental health downstream tasks and demonstrate their potential for clinical applications.
Numbers Matter! Bringing Quantity-awareness to Retrieval Systems (2024.findings-emnlp)

Copied to clipboard

Challenge: Quantitative information is important for understanding documents and interpreting them.
Approach: They propose two quantity-aware ranking techniques that rank both quantity and textual content . they use available retrieval systems to incorporate quantity information into queries .
Outcome: The proposed methods can rank both quantity and textual content, either jointly or independently.
Figure Me Out: A Gold Standard Dataset for Metaphor Interpretation (2020.lrec-1)

Copied to clipboard

Challenge: Metaphor comprehension and understanding is a complex cognitive task that requires interpreting metaphors by grasping the interaction between the meaning of their target and source concepts.
Approach: They propose an automatic retrieval approach to annotate verb-noun metaphors in text . they validated their approach by annotating around 1,500 metaphors from tweets .
Outcome: The proposed method reduces the workload on annotators and maintains consistency . it can be used to interpret verb-noun metaphoric expressions in tweets .
ESCoT: Towards Interpretable Emotional Support Dialogue Systems (2024.acl-long)

Copied to clipboard

Challenge: Emotion-focused and strategy-driven chain-of-thought (ESCoT) is a new paradigm for emotional support dialogues.
Approach: They propose an emotional support response generation scheme to improve interpretability . they generate a dataset and develop a model to generate dialogue responses with better interpretability.
Outcome: The proposed scheme can generate dialogue responses with better interpretability.
Get Confused Cautiously: Textual Sequence Memorization Erasure with Selective Entropy Maximization (2025.coling-main)

Copied to clipboard

Challenge: Existing methods for erasure of memorized text fail to unlearn large numbers of memorizable samples without jeopardizing model utility.
Approach: They propose a method that allows LLMs to memorize and recite some training sequences verbatim . they propose an entropy-based loss method that is shown to be more stable .
Outcome: The proposed method improves model utility and accuracy while preserving model ability in language generation and understanding.
The Bull and the Bear: Summarizing Stock Market Discussions (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 7888 reddit posts and 400 posts is used to summarize stock market topics.
Approach: They curate discussions on social media platforms and construct an abstractive summarization dataset.
Outcome: The proposed dataset consists of 7888 Reddit posts and summaries for 400 posts . it is robustly evaluated and will be made publicly available .
Securing Multi-turn Conversational Language Models From Distributed Backdoor Attacks (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) have acquired the ability to handle longer context lengths and understand nuances in text, expanding their dialogue capabilities beyond a single utterance.
Approach: They propose a decoding time defense that scales linearly with the input sequence length and reduces the backdoor to as low as 0.35%.
Outcome: The proposed framework is generalizable, compatible with any trigger in an adversary’s toolbox in a plug-and-play manner.
Fora: A corpus and framework for the study of facilitated dialogue (2024.acl-long)

Copied to clipboard

Challenge: a new study of facilitated dialogues focuses on the sharing of personal experience . social media is a popular method of civic engagement but lacks the tools to analyze it .
Approach: They compile 262 facilitated conversations hosted with partner organizations . they taxonomize personal sharing behaviors and facilitation strategies in the corpus .
Outcome: The proposed framework can be used to analyze facilitated dialogues and parse spoken conversations . the data can be applied to other fields, including civic use in governance and social science .
Large Language Models Are Poor Clinical Decision-Makers: A Comprehensive Benchmark (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on evaluating large language models in close-ended QA tasks, but many clinical decisions involve answering open-ended questions without pre-set options.
Approach: They construct a benchmark to better understand large language models in the clinic . they use existing datasets to evaluate LLMs in clinical situations .
Outcome: The proposed model outperforms human experts in multiple medical tasks.
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue? (2024.findings-emnlp)

Copied to clipboard

Challenge: Emphasis is a crucial component in human communication, which indicates speaker’s intention and implication beyond pure text in dialogue.
Approach: They propose a benchmark dataset with annotated dialogue samples capturing the implications of emphasis.
Outcome: The proposed evaluation pipeline achieves high correlation with human scoring and commercial LLMs perform better than open-source LLM.
SMASH Corpus: A Spontaneous Speech Corpus Recording Third-person Audio Commentaries on Gameplay (2020.lrec-1)

Copied to clipboard

Challenge: Developing a spontaneous speech corpus is important for spoken language research . a corpus of spontaneous speech is needed to develop these techniques .
Approach: They propose to use Japanese male commentators' spontaneous speech to construct a SMASH corpus . they use transcriptions and topic tags to annotate the commentaries and report some results .
Outcome: The proposed corpus includes spontaneous speech of two Japanese male commentators . the authors report that the annotations yielded a better corpus than the previous methods .
FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents (2024.emnlp-main)

Copied to clipboard

Challenge: FinDVer is a benchmark to evaluate the explainable claim verification capabilities of LLMs . financial documents are typically long, intricate and dense, and they include both quantita and numerical reasoning.
Approach: They propose a benchmark to evaluate the explainable claim verification capabilities of LLMs . they assess 25 LLM systems under long-context and RAG settings .
Outcome: The proposed benchmark can be used to evaluate the explainable claim verification capabilities of LLMs in financial documents.
Not All Languages Are Created Equal in LLMs: Improving Multilingual Capability by Cross-Lingual-Thought Prompting (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) demonstrate impressive multilingual capability, but their performance varies substantially across different languages.
Approach: They propose a generic template prompt that stimulates cross-lingual and logical reasoning skills to enhance task performance across languages.
Outcome: The proposed method improves multilingual capability across languages and covers high-resource and low-resourced languages.
MoE-SLU: Towards ASR-Robust Spoken Language Understanding via Mixture-of-Experts (2024.findings-acl)

Copied to clipboard

Challenge: Spoken language understanding (SLU) is a crucial task in task-oriented dialogue systems.
Approach: They propose an ASR-Robust SLU framework based on the mixture-of-experts technique to generate additional transcripts from clean transcripts and use it to weigh the representations of the generated transcripts, ASR transcripts .
Outcome: The proposed framework achieves state-of-the-art on three benchmark SLU datasets.
Leibniz: Theory-of-Mind Driven Neuro-Symbolic Logical Reasoning via Multi-Agent Collaboration (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for logical reasoning with large language models suffer from insufficient rule semantic grounding and weak rule application mechanisms.
Approach: They propose a theory-of-mind driven neuro-symbolic reasoning framework that integrates natural language and symbolic representations throughout the reasoning process.
Outcome: The proposed model surpasses state-of-the-art models in reasoning accuracy and flexibility.
CULEMO: Cultural Lenses on Emotion - Benchmarking LLMs for Cross-Cultural Emotion Understanding (2025.acl-long)

Copied to clipboard

Challenge: Existing emotion benchmarks rely on keyword-based emotion recognition, overlooking cultural dimensions required for emotion understanding.
Approach: They propose a benchmark to evaluate culturally-aware emotion prediction across six languages.
Outcome: The proposed benchmark evaluates state-of-the-art LLMs on culture-aware emotion prediction and sentiment analysis tasks.
FactLens: Benchmarking Fine-Grained Fact Verification (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown impressive capability in language generation and understanding, but their tendency to hallucinate and produce factually incorrect information remains a key limitation.
Approach: They propose a benchmark to evaluate fine-grained fact verification where claims are broken down into smaller sub-claims for individual verification.
Outcome: The proposed model enables more precise identification of inaccuracies, improved transparency, and reduced ambiguity in evidence retrieval.
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions (2024.emnlp-main)

Copied to clipboard

Challenge: Using a method to identify next-token neurons, we find that some attention heads recognize contexts relevant to predicting a token and activate a downstream token-predicting neuron accordingly.
Approach: They propose a method to identify next-token neurons and determine the upstream attention heads responsible for their activity in LLMs.
Outcome: The proposed method identifies next-token neurons, finds prompts that highly activate them, and determines the upstream attention heads responsible.
PunMemeCN: A Benchmark to Explore Vision-Language Models’ Understanding of Chinese Pun Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Pun memes combine wordplay with visual elements to create humor, irony, or other rhetorical effects.
Approach: They propose a benchmark to assess Chinese pun memes' processing capabilities across three progressive tasks: pun meme detection, sentiment analysis, and chat-driven meme response.
Outcome: The proposed model can detect pun memes, analyze sentiments, and respond to chats, while ignoring homophone wordplay.
TELeR: A General Taxonomy of LLM Prompts for Benchmarking Complex Tasks (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that conversational Large Language Models (LLMs) can perform ill-defined complex tasks with different prompt types/styles and different degrees of detail.
Approach: They propose a general taxonomy that can be used to design prompts with specific properties to perform a wide range of complex tasks.
Outcome: The proposed taxonomy will allow future benchmarking studies to report specific categories of prompts used as part of the study, enabling meaningful comparisons across different studies.
Specialist or Generalist? Instruction Tuning for Specific NLP Tasks (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies have shown that instruction tuning can be a data-efficient method for transforming large language models into generalist models, but their performance lags behind specialist models trained exclusively for specific tasks.
Approach: They propose to incorporate broadcoverage generalist instruction tuning into large language models to build a specialist model by incorporating task specificity and skill requirements.
Outcome: The proposed method improves model performance when task coverage is broad and when training data is limited.
From Observation to Understanding: Front-Door Adjustments with Uncertainty Calibration for Enhancing Egocentric Reasoning in LVLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods that adapt LVLMs to egocentric tasks overlook critical agent-environment interactions, limiting their ability to perform egoic reasoning.
Approach: They propose a zero-shot paradigm to enhance egocentric reasoning by simulating human causal reasoning by formalizing ego-centric reasoning using a structural causal model.
Outcome: The proposed method improves egocentric reasoning abilities on six tasks.
Towards IP Intelligence: Benchmarking Large Language Models on Intellectual Property Knowledge and Practice (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets and benchmarks focus only on patents or cover limited aspects of the IP field, lacking alignment with real-world scenarios.
Approach: They propose a bilingual IP task taxonomy and a large-scale bilingual benchmark to evaluate LLMs in real-world IP practice.
Outcome: The proposed model achieves only 75.8% accuracy, indicating room for improvement . open-source IP and law-oriented models lag behind closed-source general-purpose models .
LC-Eval: A Bilingual Multi-Task Evaluation Benchmark for Long-Context Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated sophisticated capabilities, including the ability to process and comprehend extended contexts.
Approach: They propose a bilingual, multi-task evaluation benchmark designed to evaluate long-context understanding in English and Arabic.
Outcome: The proposed benchmark targets context lengths ranging from 4k to over 128k tokens.
Multilingual Topic Classification in X: Dataset and Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Social media platforms such as X (Twitter), Snapchat and Instagram provide an environment for content creation and information sharing.
Approach: They propose a multilingual dataset featuring tweet topic classification in four languages . they leverage X-Topic to perform cross-linguistic and multilingual analysis .
Outcome: The proposed dataset includes topics in four languages and is useful for cross-linguistic analysis and the development of robust multilingual models.
Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties (2024.emnlp-main)

Copied to clipboard

Challenge: Emergent In-context Learning on Videos induces in-contact learning over video and text . eILeV-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions.
Approach: They implement Emergent In-context Learning on Videos (EILeV) that induces in-contact learning over video and text by capturing key properties of pre-training data.
Outcome: The proposed training paradigm outperforms off-the-shelf VLMs in few-shot video narration for novel, rare actions.
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Large multimodal foundation models perceive objects as indivisible, overlooking the components that constitute them.
Approach: They propose a novel benchmark for large multimodal foundation models comprising hand-labeled part segmentation annotations and task-oriented instructions to evaluate their performance.
Outcome: The proposed benchmark improves performance of current models in understanding and executing part-level tasks within everyday contexts.
Reflections & Resonance: Two-Agent Partnership for Advancing LLM-based Story Annotation (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for story annotation require a meticulous and resourceintensive effort, but the advent of advanced computational tools like GPT-4 can streamline the process and mitigate common limitations.
Approach: They propose a multi-agent system that generates tailored prompts for a large language model and provides feedback to refine the initial prompts.
Outcome: The proposed system significantly improves the model's reconstruction accuracy and confidence, demonstrating that dynamic interaction between agents significantly boosts the annotation process's precision and efficiency.
Rethinking Text-based Protein Understanding: Retrieval or LLM? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment.
Approach: They propose a retrieval-enhanced method which significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios.
Outcome: The proposed method significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios.
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search (2026.findings-acl)

Copied to clipboard

Challenge: Recent work has demonstrated the promise of orchestrating large language models (LLMs) within evolutionary and agentic optimization systems.
Approach: They present a large-scale study of LLM-guided evolutionary search . they find strong LLMs behave as local refiners, producing frequent improvements . weaker LLM optimizers exhibit large semantic drift, they say .
Outcome: The results highlight the importance of trajectory analysis for understanding and improving LLM-based optimization systems.
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to test-time scaling are limited due to the quality of candidate responses.
Approach: They propose a new metric to quantify the relative improvement of self-refinement beyond majority voting.
Outcome: The proposed method achieves state-of-the-art performance across five benchmarks over other methods.
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks (2026.acl-long)

Copied to clipboard

Challenge: Generating synthetic datasets via large language models (LLMs) has emerged as promising approach to improve LLM performance.
Approach: They propose three mitigation strategies to mitigate bias inheritance in LLMs by analyzing real and LLM-augmented data.
Outcome: The proposed methods can work differently on different tasks and biases.
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing statistical methods for evacuation decision prediction fail to capture complex and diverse behavioral logic of different individuals.
Approach: They propose a Large Language Model (LLM)-based framework that integrates behavioral theories and models to streamline the Chain-of-Thought reasoning and integrates with memory-based Reinforcement Learning module to provide accurate evacuation decision prediction and understanding.
Outcome: The proposed framework improves on three post-wildfire survey datasets with strong cross-event generalizability over existing models.
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing prompting methods for multimodal large language models lack fine-grained perception across disparate images . existing methods fail to integrate perception and reasoning, causing problems with general multi-image reasoning tasks.
Approach: They propose a generalized prompting method that integrates perception and reasoning . they evaluate the method on open-source and closed-source MLLMs .
Outcome: The proposed method shows competitive performance across tasks and improves in challenging scenarios.
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding (2025.acl-long)

Copied to clipboard

Challenge: 3D visual grounding models localize entities in a scene referred to by natural language text . recent studies focused on LLM-based scaling of 3DVG datasets, but these do not capture the full range of potential prompts which could be specified in the English language.
Approach: They propose a framework for linguistically analyzing 3DVG prompts and introduce a diagnostic dataset for evaluating 3D visual grounding methods against a diverse set of language patterns.
Outcome: The proposed framework scales up and tests against a representative set of prompts in the english language.
MemeIntel: Explainable Detection of Propagandistic and Hateful Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for label detection and explanation generation have been limited in understanding complex issues . identifying propaganda and hate in memes is essential for combating misinformation and minimizing harm .
Approach: They propose an explanation-enhanced dataset for propaganda memes in Arabic and hateful memes on English to solve these tasks.
Outcome: The proposed model outperforms the current state-of-the-art in label detection and explanation generation.
UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing datasets address understanding and generation in isolation, limiting the performance of unified vision large language models.
Approach: They propose a dataset that facilitates mutual enhancement between multimodal understanding and generation.
Outcome: The proposed framework integrates diverse visual and textual inputs and outputs, enabling comprehensive cross-modal reasoning and precise text-to-image alignment.
REVIVING YOUR MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation methods assess performance after LLMs are fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors.
Approach: They propose a framework for identifying unintended side effects using sparse model diffing.
Outcome: The proposed framework can detect unintended side effects without fine-tuning data . it achieves 95% accuracy in predicting side effects, aligning with known benchmarks .
Bias Dynamics in BabyLMs: Towards a Compute-Efficient Sandbox for Democratising Pre-Training Debiasing (2026.findings-acl)

Copied to clipboard

Challenge: Pre-trained language models (LMs) have grown substantially in both societal adoption and training costs.
Approach: They propose to use low-cost proxy models to democratise pre-model debiasing research by using small and mutable corpora.
Outcome: The proposed model can approximate bias acquisition and learning dynamics of larger models despite their reduced size.
TounsiBench: Benchmarking Large Language Models for Tunisian Arabic (2025.emnlp-main)

Copied to clipboard

Challenge: a dataset of Tunisian Arabic instructions and prompts is used to evaluate LLMs' ability to understand and generate responses in Tunisia . we assess the quality, correctness, relevance, and dialectal adherence of LLM responses .
Approach: They propose a benchmark for evaluating the capabilities of large language models in Tunisian Arabic . they use a dataset of Tunisia Arabic instructions and prompts to evaluate their models .
Outcome: The proposed model can judge quality, correctness, relevance, and dialectal adherence . the model can also generate a leaderboard for the Tunisian Arabic language .
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations (2025.emnlp-main)

Copied to clipboard

Challenge: a recent study evaluated the extent to which SLMs encode nuanced syntactic and conceptual features . acoustic and phonetic features are shallow, but the extent of nuance is unclear .
Approach: a new study evaluates contextual syntactic and semantic features in transformer-based speech language models . authors compare SLMs to linguistic competence assessments for large language models.
Outcome: a new study compares SLMs with linguistic competence assessments to assess speech recognition and understanding . the results show that SLM models encode grammatical features more robustly than conceptual ones .
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are a powerful tool for high-performance inference serving.
Approach: They focus on system-aware KV infrastructure for serving LLMs . they analyze cross-behavior co-design affinity and behavior-objective links .
Outcome: The proposed key-value (KV) cache is crucial for low-latency, high-throughput LLM inference serving.
A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Multimodal mathematical Reasoning (MMR) has attracted increasing attention for its ability to solve mathematical problems involving both textual and visual modalities.
Approach: They review the theoretical frameworks of multimodal reasoning and examine the challenges they face in visual math tasks.
Outcome: The proposed models can solve problems involving both textual and visual modalities.
How Long Reasoning Chains Influence LLMs’ Judgment of Answer Factuality (2026.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly adopted as scalable judges for open-ended generation, yet how they form judgments remains insufficiently understood.
Approach: They show that exposing reasoning influences LLM-based judgment . they also show that reasoning fluency and factuality critically shape judgment outcomes .
Outcome: Empirical results show that the presence of reasoning significantly alters judgment behavior . stronger judges exhibit more selective behavior and achieve higher judgment accuracy .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations