Papers with Multimodality

300 papers
PAI-Diffusion: Constructing and Serving a Family of Open Chinese Diffusion Models for Text-to-image Synthesis on the Cloud (2024.acl-demos)

Copied to clipboard

Challenge: Existing diffusion models fail to address the challenges of generating high-quality images from textual descriptions due to its large vocabulary size and complex character relationships.
Approach: They propose a framework that integrates Chinese diffusion models with Alibaba Cloud's Platform for AI and enables the generation of contextually relevant images.
Outcome: The proposed framework integrates with Alibaba Cloud’s Platform for AI, providing accessible and scalable solutions.
AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process.
Approach: This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process.
Outcome: This tutorial provides an overview of recent advances in AI-assisted tools and models that support and enhance the scientific research process.
NLP Lean Programming Framework: Developing NLP Applications More Effectively (N18-5)

Copied to clipboard

Challenge: NLPf is a framework for creating custom natural language processing models and pipelines by utilizing common software development build systems.
Approach: They propose a framework for creating custom NLP models and pipelines by utilizing common software development build systems.
Outcome: This framework allows developers to train and integrate domain-specific NLP pipelines into their applications seamlessly.
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding (2026.eacl-srw)

Copied to clipboard

Challenge: Recent studies improve visual contrastive decoding (VCD) by constructing more informative auxiliary views.
Approach: They propose to construct an object-aligned auxiliary view that disrupts unsupported tokens and produces a stronger contrast signal.
Outcome: Empirically, the proposed method shows consistent gains on two popular object hallucination benchmarks across two MLLMs.
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios (2024.acl-demos)

Copied to clipboard

Challenge: OpenVNA is an open-source framework for analyzing the behavior of multimodal language understanding systems under noisy conditions.
Approach: They propose to use OpenVNA to analyze behavior of multimodal language understanding systems under noisy conditions.
Outcome: The proposed framework provides high flexibility and extensibility, enabling customization with user-defined noise types and models.
MERaLiON-AudioLLM: Advancing Speech and Language Understanding for Singapore (2025.acl-demo)

Copied to clipboard

Challenge: MERaLiON-AudioLLM is the first general-purpose audio-based large language model for multitask learning.
Approach: They introduce MERaLiON-AudioLLM, a general-purpose audio-based large language model for multitask learning with a focus on Singlish understanding.
Outcome: The proposed model exhibits strong generalization across a diverse set of tasks . it is a leading solution for region-specific AI applications.
LAVIS: A One-stop Library for Language-Vision Intelligence (2023.acl-demo)

Copied to clipboard

Challenge: a new open-source library for language-vision research and applications is available for free.
Approach: They introduce LAVIS, an open-source deep learning library for LAnguage-VISion research and applications.
Outcome: The proposed library is open-source and highly extensible and configurable.
Learning-based Composite Metrics for Improved Caption Evaluation (P18-3)

Copied to clipboard

Challenge: Existing image captioning metrics focus on linguistic aspects and do not match human judgements at sentence-level.
Approach: They propose to incorporate lexical and semantic metrics as features to capture adequacy and fluency of captions at different linguistic levels.
Outcome: The proposed framework captures adequacy and fluency of captions at different linguistic levels.
Aligning Multilingual Word Embeddings for Cross-Modal Retrieval Task (D19-66)

Copied to clipboard

Challenge: Existing methods to learn multimodal multilingual embeddings for text and image retrieval tasks are limited to English.
Approach: They propose a new approach to learn multimodal multilingual embeddings for matching images and captions in two languages by combing two existing objective functions and adapting alignment between existing languages.
Outcome: The proposed model achieves state-of-the-art in retrieval and caption-caption tasks while adapting existing language alignments.
Tutorial on Multimodal Machine Learning (2022.naacl-tutorials)

Copied to clipboard

Challenge: Multimodal machine learning is a challenging but crucial area with numerous applications in multimedia, affective computing, robotics, finance, HCI, and healthcare.
Approach: This tutorial will describe an updated taxonomy on multimodal machine learning synthesizing its core technical challenges and major directions for future research.
Outcome: The proposed taxonomy synthesizes the core technical challenges and major directions for future research.
NeurST: Neural Speech Translation Toolkit (2021.acl-demo)

Copied to clipboard

Challenge: a toolkit for speech translation is available for free and provides step-by-step recipes for feature extraction, data preprocessing, distributed training, and evaluation.
Approach: They propose to use NeurST to facilitate speech translation research for NLP researchers . they show experimental results for different benchmark datasets which can be regarded as reliable baselines .
Outcome: The proposed framework provides reliable benchmarks for speech translation research.
Exploring Hyperbolic Hierarchical Structure for Multimodal Rumor Detection (2025.findings-emnlp)

Copied to clipboard

Challenge: rumor detection models often assume a simplistic one-to-one alignment between modalities . authors present a method that preserves hierarchical, non-linear relationships .
Approach: They propose a method that uses hyperbolic geometry to preserve hierarchical relationships . it decomposes image and text content into three levels and embeds them in hyperbolical space .
Outcome: The proposed method preserves hierarchical relationships rather than representing them at a flat semantic level.
ALTER: Augmentation for Large-Table-Based Reasoning (2025.naacl-long)

Copied to clipboard

Challenge: Recent studies have focused on the use of large language models (LLMs) for table-based reasoning, but most approaches struggle with scalability when applied to large tables.
Approach: They propose a framework to harness latent augmentation potential in tabular data . they use only a small subset of relevant data from the table to supplement it with schema .
Outcome: The proposed framework outperforms all other approaches and exhibits robustness and efficiency against perturbations in large-table scenarios.
Entity resolution for noisy ASR transcripts (D19-3)

Copied to clipboard

Challenge: Domain-agnostic Automatic Speech Recognition systems often mistranscribe domain-specific words and phrases.
Approach: They propose a method for handling ASR errors in named entities, specifically person names, for a voice-based collaboration assistant.
Outcome: The proposed method improves accuracy by 40.8% on a voice-based collaboration assistant.
Does Simultaneous Speech Translation need Simultaneous Models? (2022.findings-emnlp)

Copied to clipboard

Challenge: Simultaneous speech translation (SimulST) systems strive for high output quality but also low latency.
Approach: They propose to train SimulST offline without additional training or adaptation . they also show offline training achieves similar or better quality compared to offline training .
Outcome: The proposed model can serve both offline and simultaneous applications without additional training or adaptation.
PaddleSpeech: An Easy-to-Use All-in-One Speech Toolkit (2022.naacl-demo)

Copied to clipboard

Challenge: PaddleSpeech is an open-source speech toolkit that supports speech-to-text and text-to speech tasks.
Approach: They describe the design philosophy and core architecture of PaddleSpeech to support several essential speech-to-text and text-to speech tasks.
Outcome: The proposed framework achieves competitive or state-of-the-art performance on various speech datasets and implements the most popular methods.
N-Best ASR Transformer: Enhancing SLU Performance using Multiple ASR Hypotheses (2021.acl-short)

Copied to clipboard

Challenge: Spoken Language Understanding systems parse spoken utterances into semantic structures like dialog acts and slots.
Approach: They propose to use concatenated N-best ASR alternatives to represent utterances . they propose to employ a simpler utteration representation with no special delimiter .
Outcome: The proposed model outperforms the prior state-of-the-art model on DSTC2 dataset.
The Emergence of High-Level Semantics in a Signaling Game (2024.starsem-1)

Copied to clipboard

Challenge: a symbol grounding problem has been raised in recent years in AI . we show that neural agents can communicate high-level semantic concepts .
Approach: They propose to use an adversarial agent to train neural agents in a signaling game . they show that the agents can communicate high-level semantic concepts rather than low-level features .
Outcome: The proposed method can learn to communicate high-level semantic concepts . it also produces an appropriate training signal when no other method is available .
MULSUM: A Multimodal Summarization System with Vis-Aligner and Diversity-Aware Image Selection (2026.eacl-long)

Copied to clipboard

Challenge: Existing systems that condense text and images into concise, faithful digests are inefficient and require large fusion transformers.
Approach: They propose a framework that uses image embeddings to generate a visually informed text summary and a Diversity-Aware Image Selector to maximize images-relevance to the summary.
Outcome: The proposed framework outperforms baselines on automatic metrics such as ROUGE and human evaluation shows that selected images act as explanatory evidence rather than ornamental add-ons.
End-to-End Emotion-Cause Pair Extraction with Graph Convolutional Network (2020.coling-main)

Copied to clipboard

Challenge: Emotion-cause pair extraction (ECPE) aims to extract emotion expressions and their corresponding causes in a document simultaneously.
Approach: They propose to model pair-level contexts so that to capture dependency information among local neighborhood candidate pairs.
Outcome: The proposed model extracts emotion-cause pairs and their causes from documents . it is based on a benchmark Chinese emotion-case pair extraction corpus .
Mining Tweets that refer to TV programs with Deep Neural Networks (D19-55)

Copied to clipboard

Challenge: opinion mining is a popular natural language processing technique, but a problem is robustness for user-generated texts . a recent study shows that a model that handles context can extract the opinion target with 90% accuracy .
Approach: They propose a model that handles context in many natural language processing areas to solve a problem of extracting opinion references from text.
Outcome: Experiments on tweets that refer to television programs show the proposed model can extract opinion references with more than 90% accuracy.
Tab-Cleaner: Weakly Supervised Tabular Data Cleaning via Pre-training for E-commerce Catalog (2023.acl-industry)

Copied to clipboard

Challenge: Existing methods for analyzing textual attributes in product catalogs are not effective on structured tabular data since they are trained on free-form natural language texts.
Approach: They propose a model to handle error detection over tabular data following a pre-training paradigm.
Outcome: The proposed model improves on a real-world Amazon Product Catalog table by 16% over state-of-the-art methods and by 11% on PR AUC over attribute value validation task.
Towards Speech to Speech Machine Translation focusing on Indian Languages (2023.eacl-demo)

Copied to clipboard

Challenge: SSMT is a web application for translating videos from one language to another by cascading multiple language modules.
Approach: They introduce an SSMT pipeline for translating videos from one language to another by cascading multiple language modules.
Outcome: The proposed system can get 3.5+ MOS score for English to Hindi using human intervention.
HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical Understanding (2025.coling-main)

Copied to clipboard

Challenge: Object categories are typically organized into a multi-granularity taxonomic hierarchy . traditional uni-modal approaches focus primarily on image features, revealing limitations in complex scenarios.
Approach: They propose a framework that combines vision-language models with a deeper exploitation of the hierarchy.
Outcome: The proposed framework shows significant improvements on 11 diverse visual recognition benchmarks.
SMARTAVE: Structured Multimodal Transformer for Product Attribute Value Extraction (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for product attribute value extraction are noisy and incomplete with missing values for most retailers.
Approach: They propose a Structure Mltimodal trAnsformeR for producT Attribute Value Extraction which jointly encodes the structured product information from multiple modalities.
Outcome: The proposed method outperforms state-of-the-art methods on two multimodal product datasets.
Providing Semantic Knowledge to a Set of Pictograms for People with Disabilities: a Set of Links between WordNet and Arasaac: Arasaac-WN (2020.lrec-1)

Copied to clipboard

Challenge: Pictograms are a tool that is increasingly used by people with cognitive or communication disabilities.
Approach: They propose a database that links WordNet and Arasaac to link pictograms to semantic knowledge.
Outcome: The proposed database links pictograms with WordNet and Arasaac to create language-independent prototypes.
MusicAgent: An AI Agent for Music Understanding and Generation with Large Language Models (2023.emnlp-demo)

Copied to clipboard

Challenge: MusicAgent integrates numerous music-related tools and an autonomous workflow to address user requirements.
Approach: a new system is built to integrate music-related tools and an autonomous workflow . the system is based on large language models (LLMs) that can be used to organize and decompose requests .
Outcome: the proposed system integrates numerous music-related tools and an autonomous workflow to address user requirements.
ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery (2026.acl-demo)

Copied to clipboard

Challenge: a new approach to natural-language research questions requires manual effort to generate an annotation schema and label the corpus.
Approach: They propose a natural-language search tool that takes a question and a corpus to produce a schema and db with a web interface that lets steer and revise the extraction.
Outcome: The proposed model yields outputs that support real-world analysis in law and computational biology.
Improving Image Captioning via Predicting Structured Concepts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on image captioning ignore the relationship between concepts . current methods for image caption generation ignore this relationship .
Approach: They propose a structured concept predictor to predict concepts and their structures . they integrate these predictions into captioning to enhance visual signals .
Outcome: The proposed approach improves image captioning performance by using semantic concepts as a bridge between images and texts.
Through the Looking Glass: Common Sense Consistency Evaluation of Weird Images (2025.naacl-srw)

Copied to clipboard

Challenge: Existing methods to measure image common sense inconsistentness are difficult to implement because of their complexity.
Approach: They propose a visual commonsense model that leverages large vision-language models to extract atomic facts from images and a compact attention-pooling classifier to fine-tune it over encoded atomic fact.
Outcome: The proposed method outperforms existing methods on the WHOOPS! and WEIRD datasets while maintaining a compact attention-pooling classifier over encoded atomic facts.
Revealing Single Frame Bias for Video-and-Language Learning (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for video-and-language learning use multiple frames as inputs.
Approach: They propose to use single-frame models for video-and-language learning to investigate temporality in video- and language tasks.
Outcome: The proposed model does not take into account temporal information on video-and-language tasks.
Open-Ended Visual Question Answering by Multi-Modal Domain Adaptation (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to visual question answering (VQA) are not suitable for real-world applications.
Approach: They propose a supervised multi-modal domain adaptation method for visual question answering in images that exploits supervised domain adaptation.
Outcome: The proposed method outperforms state-of-the-art methods on the benchmark VQA 2.0 and VizWiz datasets.
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials (2025.findings-naacl)

Copied to clipboard

Challenge: Recent advances in vision-language models have accelerated research into models capable of advanced reasoning based on images.
Approach: They propose a method that leverages vision-language models to convert charts into table format . they use Large Language Model (LLM) for reasoning to extract only the essential information .
Outcome: The proposed method extracts only the elements necessary for chart reasoning without the need for additional annotations or datasets.
TURING: an Accurate and Interpretable Multi-Hypothesis Cross-Domain Natural Language Database Interface (2021.acl-demo)

Copied to clipboard

Challenge: Existing text-to-SQL semantic parsers cannot achieve high accuracy in cross-database setting . TURING is a NLDB system that can be used to democratize data-driven insights for non-technical users .
Approach: They propose a TURING system that provides high-precision natural language explanations of SQL queries in a beam.
Outcome: The proposed system achieves 75.1% execution accuracy and 78.3% top-5 beam execution accuracy on the Spider validation set.
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving (2026.findings-eacl)

Copied to clipboard

Challenge: Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance.
Approach: They propose a Spatial Comprehension-Infused Symbolic Reasoning Framework to integrate spatial representations into structured symbolic reasoning chains.
Outcome: The proposed framework outperforms existing models in vision-intensive mathematical problems.
MMPE: A Multi-Modal Interface using Handwriting, Touch Reordering, and Speech Commands for Post-Editing Machine Translation (2020.acl-demos)

Copied to clipboard

Challenge: a shift from traditional translation to post-editing (PE) of machine-translated text can save time and reduce errors, but it also affects the design of translation interfaces.
Approach: They propose a prototype that combines traditional input modes with pen, touch, and speech modalities for post-editing of machine-translated (MT) they propose to use these modalités to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
Outcome: The proposed interfaces can be used to cross out or hand-write new text, drag and drop words for reordering, or use spoken commands to update the text in place.
NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails (2023.emnlp-demo)

Copied to clipboard

Challenge: NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.
Approach: They propose to add programmable guardrails to LLMs that are user-defined, independent of the underlying LLM, and interpretable.
Outcome: The proposed approach can be used with several LLM providers to develop controllable and safe LLM applications using programmable rails.
MedEBench: Diagnosing Reliability in Text-Guided Medical Image Editing (2025.findings-emnlp)

Copied to clipboard

Challenge: Text-guided image editing has seen significant progress in natural image domains, but its application in medical imaging remains limited.
Approach: a new benchmark is designed to diagnose reliability in text-guided medical image editing. a clinically grounded evaluation framework measures Editing Accuracy, Context Preservation, and Visual Quality.
Outcome: a new benchmark is designed to diagnose reliability in medical image editing.
Do Vision-Language Models Understand Compound Nouns? (2024.naacl-short)

Copied to clipboard

Challenge: Open-vocabulary vision-language models (CLIP) are emerging as a promising new paradigm for text-to-image retrieval.
Approach: They propose a benchmark to evaluate the effectiveness of open-vocabulary vision-language models (CLIP) for text-to-image retrieval using contrastive loss.
Outcome: The proposed framework improves CN understanding of CLIP by 8.25% on Compun.
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing datasets are too challenging for direct model learning or suffer from misalignment between text and images.
Approach: They propose a pipeline that leverages GPT-4 and GPT4V to generate geometry problems with aligned text and images, facilitating model learning.
Outcome: The proposed pipeline generates 4.9K geometry problems with aligned text and images, facilitating model learning.
Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions (2025.naacl-industry)

Copied to clipboard

Challenge: Streaming automatic speech recognition models use high power consumption to improve usability and accuracy.
Approach: They propose to optimize on-device speech recognition models by adjusting component energy sensitivities based on their specific energy sensitities to reduce power consumption.
Outcome: The proposed approach achieves up to 47% lower energy usage while preserving comparable model accuracy and improving real-time performance compared to leading methods.
Reducing cohort bias in natural language understanding systems with targeted self-training scheme (2023.acl-industry)

Copied to clipboard

Challenge: In deep learning models, it is hard to capture all the variations of the language that different users can use.
Approach: They propose a framework that uses four active learning strategies to identify important samples coming from new users and a self training phase where a teacher model is trained from the first phase to expand the training data with relevant cohort utterances.
Outcome: The proposed framework reduces the bias related to new customers in a digital voice assistant system by using two phases: a fixing phase and a self training phase.
Evaluating Multimodal Generative AI with Korean Educational Standards (2025.naacl-short)

Copied to clipboard

Challenge: Current benchmarks focus on English, overlooking the linguistic diversity worldwide and offering limited insights into low-resource languages like Korean.
Approach: They propose to use Korean national educational tests to evaluate AI systems using a benchmark dataset.
Outcome: The proposed benchmarks evaluate models in less-explored languages and open-source code and dataset builder will be fully open-sourced.
Challenging America: Modeling language in longer time scales (2022.findings-naacl)

Copied to clipboard

Challenge: a dominant approach to solving NLP tasks is pre-training a large neural language model and fine-tuning the model for specific tasks.
Approach: They propose a challenge to train and fine-tune large Transformer models for historical texts . they pre-trained a RoBERTa model from scratch from the historical texts and evaluate them on benchmarks .
Outcome: The proposed ML task is based on OCR-ed clippings from the Chronicling America portal.
Learning to Represent Image and Text with Denotation Graph (2020.emnlp-main)

Copied to clipboard

Challenge: Recent advances in learning representations of visual and language information have been a problem with many applications.
Approach: They propose to extract visual expressions from images aligned with linguistic expressions that describe the images to learn representations from implicit expressions.
Outcome: The proposed representations lead to stronger empirical results on downstream tasks of cross-modal image retrieval, referring expression, and compositional attribute-object recognition.
Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning (2020.emnlp-main)

Copied to clipboard

Challenge: Observable changes in the scene are reflected in captions, but actions are also linked to social aspects such as intentions, effects, and attributes that describe the agent.
Approach: They propose to generate captions from videos that describe latent aspects of the human agent's actions.
Outcome: The proposed model can be used to describe latent aspects of human actions in video clips and answer questions about videos.
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs (2024.findings-naacl)

Copied to clipboard

Challenge: Visual language models (VLMs) are achieving increasingly strong performance on multimodal tasks.
Approach: They propose to transfer reasoning capabilities from large-language models to VLMs by constructing a 20x larger dataset and a larger dataset to improve general reasoning capabilities.
Outcome: The proposed model outperforms larger models without an upstream OCR system while keeping inference time constant.
ARQA: A Benchmark for Grounded Table–Text QA in Enterprise Annual Reports (2026.eacl-industry)

Copied to clipboard

Challenge: Existing QA benchmarks focus on retrieval or single-modality reasoning . annual reports are a company's definitive record of performance .
Approach: They propose an annual report QA benchmark that compares QAs with lookups, arithmetics, and insights.
Outcome: The proposed benchmarks show strong factual retrieval but persistent weaknesses in grounded arithmetic and causal reasoning.
The Role of Data Curation in Image Captioning (2024.eacl-long)

Copied to clipboard

Challenge: Existing image captioning models treat all samples equally, neglecting mismatched data . Several other techniques have relied on curriculum learning strategies to adapt learning to the difficulty of the task.
Approach: They propose to actively curate difficult samples in datasets using curriculum learning strategies to improve captioning models.
Outcome: The proposed methods outperform existing models on the Flickr30K and COCO datasets.
Neural Naturalist: Generating Fine-Grained Image Comparisons (D19-1)

Copied to clipboard

Challenge: a dataset of 41k sentences describes fine-grained differences between photographs of birds . human observers are adept at making fine-grain comparisons, but sometimes require aid in distinguishing visually similar classes.
Approach: They propose a model that generates comparative language from a dataset of 41k sentences describing fine-grained differences between photographs of birds.
Outcome: The proposed model can explain differences in visual embedding space using natural language . it evaluates the results with humans who must use the descriptions to distinguish real images .
VGA: Vision GUI Assistant - Minimizing Hallucinations through Image-Centric Fine-Tuning (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing Large Vision-Language Models (VLMs) often overly rely on internal text-based knowledge while neglecting visual inputs.
Approach: They propose a model that balances attention image and text to enhance interpretation and reduce hallucinations by using a visual input.
Outcome: The proposed model improves interpretation and reduces hallucinations by balancing attention image and text to enhance interpretation and reduction of hallucinosity.
UniEDU: Toward Unified and Efficient Large Multimodal Models for Educational Tasks (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing research has focused on plain text, while real-world K-12 scenarios often involve multimodal data.
Approach: They propose a unified language and vision assistant called UniEDU for educational applications . it excels across multiple educational tasks while maintaining strong generalization capabilities . authors propose to use UniEDu for industry-scale deployment .
Outcome: The proposed model excels across multiple educational tasks while maintaining strong generalization capabilities.
StarFlow: Generating Structured Workflow Outputs From Sketch Images (2026.eacl-long)

Copied to clipboard

Challenge: Despite being widely used, building workflows can be complex, often requiring manual configuration through low-code platforms or visual programming tools.
Approach: They propose a framework for generating structured workflow outputs from sketches using vision-language models to automate the process.
Outcome: The proposed framework outperforms large vision-language models in the task of generating structured workflow outputs from sketches and diagrams.
DiaBiz – an Annotated Corpus of Polish Call Center Dialogs (2022.lrec-1)

Copied to clipboard

Challenge: DiaBiz is a large corpus of phone conversations from different business domains . it contains nearly 410 hours of recordings and over 3 million words of transcribed speech.
Approach: They introduce DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations . it is a multimodal, multi-modal corpor of 4036 phone conversations from nine different domains .
Outcome: The corpus of 4036 phone conversations in Poland is 410 hours long and contains over 3 million words of transcribed speech.
SlideGuard: AI-Driven Evaluation of Graduate Student Presentation Materials (2026.acl-demo)

Copied to clipboard

Challenge: Effective communication is a core objective of graduate education in AI and machine learning (ML).
Approach: They propose an evaluation agent that assesses slide decks against a framework of expert-defined criteria using a visual language model.
Outcome: The evaluation agent detects the majority of expert-identified issues on a dataset of 150 annotated slide decks and shows strong results on structural and visual criteria and known limitations on subjective dimensions such as research quality.
IndiFoodVQA: Advancing Visual Question Answering and Reasoning with a Knowledge-Infused Synthetic Data Generation Pipeline (2024.findings-eacl)

Copied to clipboard

Challenge: Large Vision Language Models lack domain-specific data for reasoning on complex problems.
Approach: They propose to use explicit knowledge-infused questions, answers, and reasons to answer and reason upon the questions.
Outcome: The proposed model improves by 25% over the baseline model.
OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance (2026.acl-demo)

Copied to clipboard

Challenge: OpenGlass is an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance . cloud MLLM assistants offer strong visual understanding but often require uploading first-person visual data .
Approach: They propose an open-source system for low-latency multimodal visual assistance . they use an ESP32-based glasses-side unit to capture visual context .
Outcome: The proposed system captures visual context while a nearby device performs local MLLM inference and speech output.
Improve Vision Language Model Chain-of-thought Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Current training recipes often rely on datasets dominated by short annotations with limited rationales, hindering the models' ability to generalize to tasks requiring comprehensive reasoning.
Approach: They propose a two-stage post-training strategy that augments short answers with CoT reasoning generated by GPT-4o, enhancing the VLM's CoT capabilities through fine-tuning.
Outcome: The proposed strategy enhances the model's CoT capabilities through fine-tuning and reinforcement learning.
Unified Multimodal Interleaved Document Representation for Retrieval (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods focus on textual content, ignoring the fact that documents can contain multiple modalities.
Approach: They propose a method that holistically embeds documents interleaved with multiple modalities . they use vision-language models that combine text, images, and tables into a unified format .
Outcome: The proposed method outperforms baselines on textual and multimodal queries.
Exploring Decomposition for Table-based Fact Verification (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing research focuses on fact verification based on unstructured text, but structured data is becoming more prevalent.
Approach: They propose to decompose complex statements into simpler subproblems to improve table-based verification by a weakly supervised parser.
Outcome: The proposed method achieves state-of-the-art accuracy on the TabFact benchmark.
ITERATE: Image-Text Enhancement, Retrieval, and Alignment for Transmodal Evolution with LLMs (2025.coling-main)

Copied to clipboard

Challenge: a new framework for visual annotation of text-based questions is needed to improve performance . obtaining corresponding images through manual annotation often entails high costs .
Approach: They propose a framework that uses visual modality to enhance the performance of text-based questions.
Outcome: The proposed framework improves the alignment between text and images by using search engines or web scraping techniques.
mTVR: Multilingual Moment Retrieval in Videos (2021.acl-short)

Copied to clipboard

Challenge: mTVR is a multilingual video moment retrieval dataset with 218K queries in English and Chinese . Various datasets have been proposed or adapted for the task, but they are all created for a single language (English).
Approach: They propose a multilingual video moment retrieval dataset with 218K queries from 21.8K TV show video clips.
Outcome: The proposed model outperforms strong monolingual baselines while using fewer parameters.
CAST: Cross-modal Alignment Similarity Test for Vision Language Models (2025.coling-main)

Copied to clipboard

Challenge: Vision Language Models (VLMs) are typically evaluated with Visual Question Answering tasks which assess a model’s understanding of scenes.
Approach: They propose to use visual question answering (VQA) to assess a model's understanding of scenes to probe for self-consistency across modalities.
Outcome: The proposed test does not focus on objective accuracy but rather on whether VLMs are internally consistent in their outputs.
Multimodal Topic-Enriched Auxiliary Learning for Depression Detection (2020.coling-main)

Copied to clipboard

Challenge: Existing studies on depression detection rely on textual and visual content to determine whether a human being is depressed or non-depressed.
Approach: They propose a multimodal topic-enriched Auxiliary Learning approach that captures topic information from texts and images for depression detection.
Outcome: The proposed approach improves the performance of the primary task by using topic information from text and images.
An Effective Span-based Multimodal Named Entity Recognition with Consistent Cross-Modal Alignment (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to name entity recognition rely on word-based sequence labeling and align image and text at inconsistent semantic levels.
Approach: They propose a span-based method which achieves a more consistent multimodal alignment from the perspectives of information-theoretic and cross-modal interaction.
Outcome: Experiments on two datasets show that SMNER outperforms the state-of-the-art methods.
SCO-VIST: Social Interaction Commonsense Knowledge-based Visual Storytelling (2024.eacl-long)

Copied to clipboard

Challenge: Visual storytelling aims to automatically generate a coherent story based on a given image sequence.
Approach: They propose a framework that represents the image sequence as a graph with objects and relations that includes human action motivation and its social interaction commonsense knowledge.
Outcome: The proposed framework produces stories superior across multiple metrics in terms of visual grounding, coherence, diversity, and humanness, per both automatic and human evaluations.
Increasing Coverage and Precision of Textual Information in Multilingual Knowledge Graphs (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate knowledge graphs are unable to handle non-English textual information.
Approach: They propose a task of automatic Knowledge Graph Completion to bridge the gap between English and non-English textual information.
Outcome: The proposed method bridges the gap between the quantity and quality of textual information between English and non-English languages.
Ask No More: Deciding when to guess in referential visual dialogue (C18-1)

Copied to clipboard

Challenge: Using a task-oriented visual dialogue model, we add a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
Approach: They augment a task-oriented visual dialogue model with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
Outcome: The proposed model can be enhanced with a decision-making component that decides whether to ask a follow-up question to identify a target referent in an image, or to stop the conversation to make a guess.
CLeLfPC: a Large Open Multi-Speaker Corpus of French Cued Speech (2022.lrec-1)

Copied to clipboard

Challenge: Cued Speech is a visual communication system developed for deaf people to complement speechreading at the phonetic level with hands.
Approach: They describe a visual communication mode that uses handshapes in different placements near the face and mouth movements to make the phonemes of spoken language look different from each other.
Outcome: The proposed system is based on 4 hours of audio and video recordings of 23 participants.
Multimodal Emoji Prediction (N18-2)

Copied to clipboard

Challenge: Emojis are small images that are commonly included in social media text messages.
Approach: They propose a multimodal approach that is able to predict emojis in Instagram posts by using both text and image.
Outcome: The proposed model incorporates both text and image to improve accuracy .
DrugEHRQA: A Question Answering Dataset on Structured and Unstructured Electronic Health Records For Medicine Related Queries (2022.lrec-1)

Copied to clipboard

Challenge: a new question answering dataset is being developed for electronic health records . structured tables and unstructured notes can be duplicated, contradictory or provide additional context .
Approach: They develop a question-answer-matching dataset using structured tables and unstructured notes from an EHR.
Outcome: The proposed model is based on a model with a modality selection network . it uses the prediction of a RAT-SQL to choose between EHR tables and clinical notes .
Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for music question answering do not systematically evaluate reasoning across tracks.
Approach: They propose a dataset and benchmark for multi-track comparative question answering . they construct 36,519 comparative QA items over 12,173 track pairs .
Outcome: The proposed dataset and benchmark for multi-track comparative question answering is based on the Jamendo-QA dataset.
RealTalk-CN: A Realistic Chinese Speech Task-Oriented Dialogue Benchmark with Cross-Modal Analysis (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in speech large language models have enabled end-to-end spoken interactions, but their robustness in real-world applications remains unclear.
Approach: They propose a multi-turn, multi-domain speech–text TOD dataset for Chinese users . it contains 5.4k dialogues with annotations for dialogue states, disfluency types, speaker characteristics .
Outcome: The proposed model can be used to evaluate speech large language models in real-world scenarios . the proposed model is based on 5.4k real human-to-human dialogues with annotations .
A Regularization-based Transfer Learning Method for Information Extraction via Instructed Graph Decoder (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for information extraction (IE) focus on training task-specific models, while common knowledge among different IE tasks is not explicitly modeled.
Approach: They propose a regularization-based transfer learning method for IE via an instructed graph decoder which decodes various complex structures into a graph uniformly based on corresponding instructions.
Outcome: The proposed method can learn common knowledge from existing datasets and transfer it to a new dataset with new labels.
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages (2025.findings-acl)

Copied to clipboard

Challenge: Music information retrieval (MIR) is a field that aims at developing computational tools for processing, organizing, and accessing music data.
Approach: They propose a framework that aligns music modalities with multilingual text in a shared representation space.
Outcome: Experiments show CLaMP 3 performs state-of-the-art on multiple MIR tasks . it surpasses baselines and shows excellent generalization in multimodal and multilingual contexts .
DentalGPT: Incentivizing Multimodal Reasoning in Dentistry (2026.findings-acl)

Copied to clipboard

Challenge: Current multimodal large language models (MLLMs) show limited understanding of dental images.
Approach: They propose a dental-specialized multimodal large language model trained via staged multimodal alignment and reinforcement learning.
Outcome: The proposed model outperforms state-of-the-art models on disease classification and dental VQA tasks.
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization (2026.acl-long)

Copied to clipboard

Challenge: Existing speech codecs struggle to balance high-quality reconstruction with semantically rich representations, limiting their effectiveness in both generative and understanding tasks.
Approach: They propose a neural speech codec with semantic-acoustic dual-stream quantization that disentangles semantic and acousian modeling into two dedicated streams.
Outcome: The proposed codec outperforms state-of-the-art speech tokenizers in auto-propagating text-to-speech models.
Dynamic Routing Transformer Network for Multimodal Sarcasm Detection (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for multimodal sarcasm detection rely on fixed architectures to capture cross-modal incongruity.
Approach: They propose a method that uses dynamic paths to activate different routing transformer modules with hierarchical co-attention adapting to cross-modal incongruity.
Outcome: The proposed method is compared to state-of-the-art methods on a public dataset.
Cross-Lingual Cross-Modal Consolidation for Effective Multilingual Video Corpus Moment Retrieval (2022.findings-naacl)

Copied to clipboard

Challenge: Existing multilingual video corpus moment retrieval methods are based on a two-stream structure.
Approach: They propose a multilingual video corpus moment retrieval task that uses a two-stream structure to generate a query-visual similarity and a subtitle stream exploits the query-subtitle similarity.
Outcome: The proposed method improves accuracy on a large-scale video corpus moment retrieval dataset.
Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table Perturbation (2022.acl-long)

Copied to clipboard

Challenge: Existing Text-to-SQL parsers are vulnerable to perturbations in NL questions . we propose the Adversarial Table Perturbation (ATP) as a new attacking paradigm .
Approach: They propose to use the Adversarial Table Perturbation to measure robustness of Text-to-SQL parsers against adversarial perturbations.
Outcome: The proposed approach outperforms baseline methods in robustness evaluations on ADVETA and can be used in future projects.
Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality (2022.emnlp-main)

Copied to clipboard

Challenge: Recent visuolinguistic pre-trained models fail miserably on the Winoground dataset, which challenges models to match paired images and English captions.
Approach: They propose to annotate a Winoground dataset that challenges visuolinguistic models to match paired images and English captions with items constructed to overlap lexically but differ in meaning.
Outcome: The proposed dataset challenges models to match paired images and English captions with items constructed to overlap lexically but differ in meaning.
\mathsf{Con Instruction}: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities (2025.acl-long)

Copied to clipboard

Challenge: Existing attacks communicate instruction through text, accompanied by a toxic image or audio . a novel gray-box attack method generates adversarial images or audio to convey harmful instructions to MLLMs .
Approach: They propose a gray-box attack method that generates adversarial images or audio to convey specific harmful instructions to MLLMs by following non-textual instruction.
Outcome: The proposed method achieves highest success rates on visual and audio-language models . larger models are more susceptible toCon Instruction, compared to their underlying models - the results will be released .
Measuring the Diversity of Automatic Image Descriptions (C18-1)

Copied to clipboard

Challenge: a lack of diversity in automatic image description systems is a general problem in natural language generation . authors use established metrics to evaluate system performance on the head of the vocabulary . automatic image descriptions are difficult because of the unbounded range of variation in natural languages .
Approach: They propose to frame automatic image description as a word recall task to quantify the production of generic sentences as 'undiversity' they propose to use established metrics to evaluate the diversity of the output .
Outcome: The proposed metrics evaluate the diversity of sentences generated by state-of-the-art systems on a MS COCO dataset.
Content-Specific Humorous Image Captioning Using Incongruity Resolution Chain-of-Thought (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods for generating humorous captions are generic and do not capture the content of images.
Approach: They propose a framework that generates content-specific resolutions from fine details extracted from an image and integrates logit bias and negative sampling to suppress the output of generic resolutions.
Outcome: The proposed framework generates humorous captions tailored to the content of specific input images.
Automatically Estimating Textual and Phonemic Complexity for Cued Speech: How to See the Sounds from French Texts (2024.lrec-main)

Copied to clipboard

Challenge: Cued Speech (CS) is a visual communication system developed for people with hearing loss to complement speech reading at the phonetic level.
Approach: They propose a method to phonemize written corpora so that each word is aligned with the corresponding CS key(s) this method is part of a wider project aimed at creating an augmented reality system displaying a virtual coding hand where the user will be able to choose a text upon its complexity for cueing.
Outcome: The proposed method is part of a wider project aimed at creating an augmented reality system displaying a virtual coding hand where the user can choose a text upon its complexity for cueing.
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild (2025.findings-naacl)

Copied to clipboard

Challenge: EgoSpeak predicts when an agent should begin speaking based on egocentric streaming video.
Approach: They propose a framework for real-time speech initiation prediction in egocentric streaming video by modeling the conversation from the camera wearer's first-person perspective.
Outcome: The proposed framework outperforms random and silence-based baselines in real time and highlights the importance of multimodal input and context length in effectively deciding when to speak.
ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities (2023.findings-eacl)

Copied to clipboard

Challenge: a vision-language benchmark for human activity planning is designed for humans . the task is easy for humans, but challenging for SOTA deep learning models .
Approach: They propose a vision-language benchmark for human activity planning that extends Charades with intents and builds on a multi-choice question test set.
Outcome: The proposed benchmark evaluates the ability of systems to anticipate and plan human actions in a multimodal visionlanguage setting.
GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing Document-level machine translation systems struggle to handle discourse-level phenomena such as pronoun resolution, lexical cohesion, and ellipsis.
Approach: They propose a graph-based document-level machine translation framework that leverages Large Language Models to model translation flow and discourse structure.
Outcome: The proposed framework outperforms commercial and closed systems in eight languages and six domains.
With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models and Vision Language Model (VLMs) have demonstrated aptitude as potential substitutes for human participants in psycholinguistic experiments.
Approach: They examine whether large language models and vision language models implicitly understand sound-based phenomena via orthography and imagery alone.
Outcome: The proposed models demonstrate sound symbolism and ability to "hear" using language and vision modules.
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Current speech evaluation systems rely on specialized systems for individual audio characteristics and poor correlation between automatic methods and human preferences.
Approach: They propose a unified evaluation framework for Large Audio Models as a Judge, AudioJudge . they propose specialized judges that can be prompted to perform audio characteristic detection tasks .
Outcome: The proposed method improves performance across audio characteristic detection and human preference simulation tasks.
Retrieval, Analogy, and Composition: A framework for Compositional Generalization in Image Captioning (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches fail to generalize well to concepts that are not observed during training.
Approach: They propose a framework that revolves around probing several similar image caption training instances and performing analogical reasoning over relevant entities in retrieved prototypes.
Outcome: The proposed framework improves on the widely used image captioning benchmarks and on composition-related evaluation metrics.
Modeling Empathetic Alignment in Conversation (2024.naacl-long)

Copied to clipboard

Challenge: Empathy requires perspective-taking and is not explicitly modelled in NLP .
Approach: They propose a new approach to recognizing alignment in empathetic speech, grounded in Appraisal Theory, and use reddit to study emotional conversations to examine alignment.
Outcome: The proposed approach can recognize appraisals and alignments in empathetic speech, and mental health professionals engage with substantially more emotional alignment.
Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge (2026.acl-long)

Copied to clipboard

Challenge: Existing text-to-video models struggle to accurately simulate real-world physics and dynamic entity interactions.
Approach: They propose a framework that integrates graph-structured temporal knowledge into video latent diffusion models to enhance compositional generation and interaction fidelity.
Outcome: The proposed framework enhances compositional generation and interaction fidelity by integrating graph-structured temporal knowledge into video latent diffusion models.
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description (2025.findings-naacl)

Copied to clipboard

Challenge: Existing 3D facial emotion modeling models are constrained by limited emotion classes and insufficient datasets.
Approach: They propose a 3D facial emotion modeling dataset that spans a wide spectrum of human emotions . they use large language models to generate a diverse array of textual descriptions .
Outcome: Emo3D is an extensive dataset that spans human emotions with images and 3D blendshapes.
Image Caption Generation for News Articles (2020.coling-main)

Copied to clipboard

Challenge: Existing work on news-image captioning requires a joint understanding of image and text.
Approach: They propose a Transformer model that integrates text and image modalities and attends to textual features from visual features in generating a caption.
Outcome: The proposed model outperforms the state-of-the-art model and improves the quality of news-image captions.
Automatic Speech Interruption Detection: Analysis, Corpus, and System (2024.lrec-main)

Copied to clipboard

Challenge: Interruption detection is a new but challenging task in the field of speech processing.
Approach: They propose to define automatic speech interruption detection and build a specialized corpus to analyze interrupted conversations.
Outcome: The proposed system can detect interruptions in speech with promising results . it can be used to ensure speaking turns are respected during official political debates .
CoNAN: A Complementary Neighboring-based Attention Network for Referring Expression Generation (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to identify and describe complex daily scenes are limited by ambiguity.
Approach: They propose a Complementary Neighboring-based Attention Network that utilizes visual differences between the target object and its highly-related neighbors as complementary features.
Outcome: The proposed expression outperforms state-of-the-art models on a dataset of 3D objects.
Farewell Freebase: Migrating the SimpleQuestions Dataset to DBpedia (C18-1)

Copied to clipboard

Challenge: Existing datasets for question answering over knowledge graphs lack answer triples from Freebase . a defunct knowledge graph makes it difficult to build "real-world" question answering systems .
Approach: They propose a benchmark dataset for simple question answering over knowledge graphs that maps SimpleQuestions entities and predicates from Freebase to DBpedia.
Outcome: The proposed dataset provides simple yet strong baselines with and without neural networks.
Automatic Speech Recognition System-Independent Word Error Rate Estimation (2024.lrec-main)

Copied to clipboard

Challenge: Word error rate (WER) is a metric used to evaluate the quality of transcriptions produced by Automatic Speech Recognition systems.
Approach: They propose a hypothesis generation method for ASR system-dependent WER estimation . they use phonetically similar or linguistically more likely alternative words to generate hypotheses .
Outcome: The proposed method outperforms baseline estimators on in-domain data and out-of-domain on Switchboard and CALLHOME.
Entheos: A Multimodal Dataset for Studying Enthusiasm (2021.findings-acl)

Copied to clipboard

Challenge: Enthusiasm is an important part of engaging communication.
Approach: They propose a multimodal dataset for studying enthusiasm composed of video, audio, and text.
Outcome: The proposed model shows that pitch, loudness, and discourse relation parsing are important in distinguishing enthusiastic communication.
Ihquin tlahtouah in Tetelahtzincocah: An annotated, multi-purpose audio and text corpus of Western Sierra Puebla Nahuatl (2025.naacl-long)

Copied to clipboard

Challenge: a corpus of audio and annotated transcriptions of an endangered Nahuatl is presented . data made available in this corpus are useful for ASR, spelling normalization, and word-level language identification.
Approach: They present a corpus of audio and annotated transcriptions of an endangered Nahuatl in Mexico . the data are useful for ASR, spelling normalization, and word-level language identification .
Outcome: The corpus is made available for use in ASR, spelling normalization, and word-level language identification tasks.
Understanding Social Media Cross-Modality Discourse in Linguistic Space (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on how images are structured with texts to form coherent meanings in human cognition have not addressed the problem.
Approach: They propose a concept of cross-modality discourse which defines how human readers couple image and text understandings.
Outcome: The proposed model shows that trendy encoders based on multi-head attention are unable to understand cross-modality discourse and modeling texts at the output layer helps yield the-state-of-the-art results.
Multimodal Named Entity Disambiguation for Noisy Social Media Posts (P18-1)

Copied to clipboard

Challenge: Social media posts often contain unstructured text or images, making opinion mining challenging.
Approach: They propose a new task for multimodal social media captions with named entities annotated and linked to external knowledge bases.
Outcome: The proposed model outperforms state-of-the-art text-only NED models . it predicts correct entities in knowledge graph embeddings space, showing its efficacy and potentials .
VoiSeR: A New Benchmark for Voice-Based Search Refinement (2021.eacl-main)

Copied to clipboard

Challenge: a new study shows that voice-based search systems are challenging to support in the context of the user intent of voice searches . support for voice-driven search, exploration, and refinement is a fundamental aspect of voice assistants .
Approach: They propose to use crowdsourcing to collect voice-based search refinements . they use 10,000 search refinement utterances to annotate a search intent .
Outcome: The proposed dataset shows that voice-based search refinements can support most common tasks . the study shows that the proposed dataset can support research in conversational query understanding .
Opinion Summarization by Weak-Supervision from Mix-structured Data (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for opinion summarization of multiple reviews lack reference summaries . OAs and ISs are often mismatched between review input and summary .
Approach: They propose a method to generate mixed-structured synthetic training data for opinion summarization.
Outcome: The proposed method outperforms existing methods on Yelp, Amazon and RottenTomatos datasets.
AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text Columns (2026.acl-long)

Copied to clipboard

Challenge: Existing tabular data synthesis methods fail to account for cross-modal heterogeneity of real-world tables, where structured continuous and discrete attributes coexist with unstructured long-text columns.
Approach: They propose a framework that synergistically trains an LLM-based text generator and a deep-learning-based non-textual generator to quantify cross-modal semantic alignment.
Outcome: The proposed framework outperforms state-of-the-art frameworks in fidelity, diversity, and task utility.
The Speed-Vel Project: a Corpus of Acoustic and Aerodynamic Data to Measure Droplets Emission During Speech Interaction (2022.lrec-1)

Copied to clipboard

Challenge: Conversations and professional interactions are associated with increased risk of SARS-CoV-2 exposure . however, it is unclear to what extent speech properties influence droplets emission .
Approach: They propose to measure velocity and direction of airflow, the number and size of droplets spread during conversation in french.
Outcome: The results will allow future simulation studies to predict the transport, dispersion and evaporation of droplets emitted under different speech conditions.
Cascade versus Direct Speech Translation: Do the Differences Still Make a Difference? (2021.acl-long)

Copied to clipboard

Challenge: a gap between direct approaches to speech translation (ST) and traditional cascade solutions has gradually decreased . a recent study found that the subtle differences observed in their behavior are not sufficient for humans neither to distinguish them nor to prefer one over the other.
Approach: They compare state-of-the-art systems representative of the two paradigms . they find subtle differences observed in their behavior are not sufficient .
Outcome: The proposed system is compared with state-of-the-art systems representative of the two paradigms.
AnRe: Analogical Replay for Temporal Knowledge Graph Forecasting (2025.acl-long)

Copied to clipboard

Challenge: Temporal Knowledge Graphs (TKGs) are vital for event prediction, yet current methods face limitations.
Approach: They propose a training-free Analogical Replay reasoning framework that uses LLMs to extract historical contexts and generate analogical reasoning examples as contextual inputs.
Outcome: The proposed model outperforms existing training-free methods on four benchmarks.
Phonetically Balanced Code-Mixed Speech Corpus for Hindi-English Automatic Speech Recognition (L18-1)

Copied to clipboard

Challenge: a phonetic balance in code-mixed Hindi-English corpus has been created . code-switching is a common phenomenon in multilingual and bilingual communities .
Approach: They propose to create a phonetically balanced read speech corpus of code-mixed Hindi-English . they use a method to select sentences that contain triphones lower in frequency than a threshold .
Outcome: The proposed corpus is phonetically balanced with a large code-mixed reference corpus.
Programmable Annotation with Diversed Heuristics and Data Denoising (2022.coling-1)

Copied to clipboard

Challenge: Neural natural language generation and understanding models require massive amounts of annotated data to be competitive.
Approach: They propose a data programming framework that can jointly construct labeled data for language generation and understanding tasks by allowing annotators to modify an automatically-inferred alignment rule set between sequence labels and text.
Outcome: The proposed framework generates high-quality data within a 1.48 BLEU and 6.42 slot F1 of 100% human-labeled data with just 100 labeled data samples outperforming benchmark annotation frameworks and other semi-supervised approaches.
MSEarth: A Multimodal Benchmark for Earth Science Phenomenon Discovery with MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing datasets often rely on synthetic data or figure-caption pairs, failing to capture the depth and complexity of geoscientific reasoning.
Approach: They propose a multimodal scientific dataset and benchmark curated from open-access publications.
Outcome: MSEarth features over 289K figures with captions enriched by contextual discussions and reasoning from original papers.
Learning Translations via Images with a Massively Multilingual Image Dataset (P18-1)

Copied to clipboard

Challenge: Existing datasets for learning translations of words are limited to a few high-resource languages and unrealistically easy settings.
Approach: They propose a large-scale multilingual corpus of images labeled with the word they represent to facilitate translation research.
Outcome: The proposed method improves on an unsupervised technique that has been limited to a few languages and unrealistic settings.
Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: End-to-end speech-to speech (S2S) dialogue systems face key challenges in incorporating external knowledge into their models.
Approach: They propose a framework that directly retrieves relevant textual knowledge from speech queries.
Outcome: The proposed framework improves the performance of end-to-end speech-tospeech dialogue systems while achieving higher retrieval efficiency.
Think Visually: Question Answering through Virtual Imagery (P18-1)

Copied to clipboard

Challenge: Existing models of geometric reasoning are based on visual representations of objects and objects, but they are not based in symbols or words.
Approach: They propose a new deep network architecture that specializes in answering questions that admit latent visual representations and learns to generate and reason over such representations.
Outcome: The proposed model can generate and reason over latent visual representations and is validated by two synthetic benchmarks.
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator (2025.naacl-long)

Copied to clipboard

Challenge: Domain Large Language Models (LLMs) are developed for domain-specific tasks based on general LLMs, but it still requires professional knowledge to facilitate the expertise for some domain- specific tasks.
Approach: They propose a pipeline to solve domain-specific calculation problems with KIPG . they use it to extract key variables and calculate outcomes dependent on domain knowledge .
Outcome: The proposed pipeline solves domain-specific calculation problems more effectively . it generates knowledge-intensive programs according to the domain- specific documents .
Open-ended Knowledge Tracing for Computer Science Education (2022.emnlp-main)

Copied to clipboard

Challenge: Knowledge tracing (KT) is a method used to estimate student mastery of concepts/skills/knowledge components from their responses to questions and to predict future performance.
Approach: They propose a student knowledge-guided code generation approach that combines program synthesis methods with student knowledge tracing methods to solve the OKT problem.
Outcome: The proposed method is based on a student knowledge-guided code generation approach and validates on coding questions.
NL2pSQL: Generating Pseudo-SQL Queries from Under-Specified Natural Language Questions (D19-1)

Copied to clipboard

Challenge: Existing studies focus on generating SQL codes from natural language questions . however, questions cover more diverse tasks including table manipulation or performance issues .
Approach: They propose a task to generate pSQL codes from natural language questions . they define two new metrics suitable for the task, Canonical-BLEU and SQL-BLUE .
Outcome: The proposed task generates well-formed queries on under-specified database issues.
Visually Guided Generative Text-Layout Pre-training for Document Intelligence (2024.naacl-long)

Copied to clipboard

Challenge: Prior work shows that pre-training techniques can boost the performance of visual document understanding (VDU) . Xu et al., 2020;; Gu e t al, 2021;; Appalaraju e al. 2022)
Approach: They propose a visually guided generative text-layout pre-training method that optimizes hierarchical language and layout modeling objectives to generate interleaved text and layout sequences.
Outcome: The proposed model can process word-intensive documents of any length and achieves competitive performance over baselines on VDU tasks.
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances have extended DPO to multimodal scenarios, achieving strong performance.
Approach: They propose to use a sentence-level preference optimization technique to optimize individual sentences for more precise preference optimization without additional models or parameters.
Outcome: Experiments show that Adaptive Sentence-level Preference Optimization significantly improves the alignment of multimodal models.
Aligning Complex Knowledge Graph Question Answering as Knowledge-Aware Constrained Code Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing frameworks that generate LF using Large Language Models (LLMs) in a few-shot setting are limited due to little exposure to the LF during pre-training.
Approach: They propose a framework that aligns the LF generation as code generation that incorporates LF-specific constraints.
Outcome: The proposed framework surpasses all few-shot baselines on KQA Pro by 21%.
Asking More Informative Questions for Grounded Retrieval (2024.findings-naacl)

Copied to clipboard

Challenge: Existing approaches to question generation for interactive retrieval have constrained answer spaces, limiting the amount of information a model can gain in a single turn.
Approach: They propose a method that incorporates presupposition handling into question selection and belief updates.
Outcome: The proposed method increases accuracy over the past state-of-the-art by 14% while resulting in 48% more efficient games in human evaluations.
Geo-Aware Image Caption Generation (2020.coling-main)

Copied to clipboard

Challenge: Standard image caption generation systems do not take contextual information or world knowledge into account.
Approach: They propose to build an image-specific representation of the geographic context and adapt the caption generation network to produce appropriate geographic names in the image descriptions.
Outcome: The proposed system achieves significant improvements on a dataset that contains contextualized captions and geographic metadata and improves BLEU, ROUGE, METEOR and CIDEr scores.
Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization Challenge (2025.findings-naacl)

Copied to clipboard

Challenge: Using a small dataset, we propose that audio of tabletop role-playing games (TTRPGs) could serve as a challenge for speaker diarization systems.
Approach: They propose that audio of tabletop role-playing games (TTRPGs) could serve as a challenge for speaker diarization systems.
Outcome: The proposed system can pick the speaker and determine that impersonating is just that.
From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models consisting of multiple steps of visual and language processing are limited in the visual and visual processing community . a visual reasoner is a plug-and-play approach that can be used to improve VLMs' reasoning abilities.
Approach: They propose a least-to-most visual reasoning paradigm that divides a question into sub-questions and invokes external tools for resolving sub-questions.
Outcome: The proposed method can improve four VLMs on four VQA benchmarks.
FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) lack the capacity to handle multimodal inputs effectively.
Approach: They introduce a reference-free and fine-grained evaluation metric that measures the faithfulness of the generated free-form answers from large vision-language models.
Outcome: The proposed metric measures the faithfulness of free-form answers from large vision-language models.
An Empirical Investigation of Bias in the Multimodal Analysis of Financial Earnings Calls (2021.naacl-main)

Copied to clipboard

Challenge: Existing research focuses on textual elements of financial disclosures but ignores the rich acoustic features in the executives’ speech.
Approach: They propose to use a multimodal approach that leverages the verbal and vocal cues of speakers in financial disclosures to predict volatility and risk.
Outcome: The proposed models outperform existing models in the financial realm but still underrepresent the diverse communities spanning demographics, gender, and native speech.
EDIS: Entity-Driven Image Search over Multimodal Web Content (2023.emnlp-main)

Copied to clipboard

Challenge: Existing image retrieval methods require large datasets and a large candidate set.
Approach: They propose a news-domain dataset for cross-modal image search with 1 million web images . they propose combining multimodal image-text pairs with a million candidates .
Outcome: The proposed dataset challenges state-of-the-art methods with dense entities and the large-scale candidate set.
PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles (L18-1)

Copied to clipboard

Challenge: Existing tools to extract structured textual content from PDFs are essential to enable scientific text mining.
Approach: They propose a PDF-to-XML textual content extraction tool that extracts structured textual contents from scientific articles in PDF format.
Outcome: The proposed tool extracts structured textual content from scientific articles in PDF format while preserving both the textual contents and layout details of the input PDF document.
Leaner and Faster: Two-Stage Model Compression for Lightweight Text-Image Retrieval (2022.naacl-main)

Copied to clipboard

Challenge: Existing text-image approaches use pre-trained vision-language representations for text retrieval . however, these models pose non-trivial memory requirements and substantial indexing time .
Approach: They propose a framework to compress large pre-trained dual-encoders for lightweight text-image retrieval.
Outcome: The proposed model performs better on Flickr30K and MSCOCO benchmarks than the original full model on mobile devices.
Text encoders bottleneck compositionality in contrastive vision-language models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multimodal models are often unable to reason about simple spatial relations or attribute attachments.
Approach: They first curate CompPrompts, a set of increasingly compositional image captions that VL models should be able to capture . then train text-only recovery probes that aim to reconstruct captions from single-vector text representations produced by several VL model.
Outcome: The proposed model can reconstruct captions from single-vector text representations produced by several models on a broader range of scenes compared to previous models.
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating the code understanding and generation capacities of Large Language Models are insufficient . existing benchmarks focus on a narrow range of popular programming languages and specific tasks .
Approach: They propose an execution-based, multilingual, multitask evaluation benchmark for LLMs . they evaluate coding performance from three dimensions: length, difficulty, efficiency .
Outcome: The proposed benchmark covers 43 programming languages and eight coding tasks.
Bridge Video and Text with Cascade Syntactic Structure (C18-1)

Copied to clipboard

Challenge: Using LSTM-CSS, we construct basic syntactic structure by completing syntastic structure.
Approach: They propose a video captioning approach that progressively completes syntactic structure by a conditional random field to construct basic syntaktic structure.
Outcome: The proposed method produces natural sentences with 42.3% and 28.5% accuracy compared to state-of-the-art methods.
Decolonising Speech and Language Technology (2020.coling-main)

Copied to clipboard

Challenge: Indigenous peoples are increasingly unable to go on without speech and language technologies, says a researcher . a postcolonial approach to computational methods for supporting language vitality is needed, says the researcher - lil'watul Lorna Williams .
Approach: They propose to examine colonising discourses in speech and language technology and propose a postcolonial approach to computational methods for supporting language vitality.
Outcome: The paper reviews colonising discourses in speech and language technology and suggests new ways of working with Indigenous communities.
Open-Vocabulary Federated Learning with Multimodal Prototyping (2024.naacl-long)

Copied to clipboard

Challenge: Existing studies assume the label space of training data and test data is identical.
Approach: They propose a framework for adaptation to a federated learning (FL) query that uses arbitrary unknown classes.
Outcome: The proposed framework exploits the knowledge learned from seen classes and robustifies the adapted framework to unseen categories.
Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training (2023.acl-long)

Copied to clipboard

Challenge: Empirical results show that CCLM significantly outperforms the prior state-of-the-art with an average absolute improvement of over 10%.
Approach: They introduce a pre-training framework that unifies cross-lingual and cross-modal pre-trained models with shared architectures and objectives.
Outcome: The proposed framework outperforms the state-of-the-art in two multi-lingual datasets and two multilingual image-text retrieval datasets.
Speech Corpus of Ainu Folklore and End-to-end Speech Recognition for Ainu Language (2020.lrec-1)

Copied to clipboard

Challenge: Ainu is an unwritten language spoken by Ainus, a minority of whom are critically endangered by UNESCO . a project of automatic speech recognition (ASR) for the Ainous language is being developed .
Approach: They propose to use automatic speech recognition for the Ainu language to help preserve its language archives.
Outcome: The proposed system improves word and phone recognition accuracy in speaker-open conditions.
Visual Pivoting Unsupervised Multimodal Machine Translation in Low-Resource Distant Language Pairs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies show that neural MT achieves much worse translation quality than statistical MT with a small number of corpora.
Approach: They propose a visual pivoting method for alignment between distant language pairs . they first construct a dataset and then apply it to pre-training and fine-tuning .
Outcome: The proposed method outperforms baselines on DLPs and close language pairs.
Attention as Grounding: Exploring Textual and Cross-Modal Attention on Entities and Relations in Language-and-Vision Transformer (2022.findings-acl)

Copied to clipboard

Challenge: Existing work has focused on what is captured by multi-modal architectures.
Approach: They propose a multi-modal transformer that learns syntactic and semantic representations about entities and relations grounded in objects at the level of masked self-attention and cross-modal attention.
Outcome: The proposed model learns syntactic and semantic representations about objects and relations cross-modally and unimodally.
FOAM: A Follower-aware Speaker Model For Vision-and-Language Navigation (2022.naacl-main)

Copied to clipboard

Challenge: Existing speaker-follower models are follower-agnostic and fail to take state of follower into account.
Approach: They propose a speaker-follower model that is constantly updated given follower feedback . they optimize the speaker and obtain its training signals by evaluating the follower on labeled data .
Outcome: The proposed model outperforms strong baseline models on room-to-room and room-across-room datasets.
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) are hardly comprehensively evaluated for their cognitive abilities.
Approach: They propose to evaluate high-level cognitive abilities of Large Vision-Language Models (LVLMs) using images with rich semantics.
Outcome: The proposed evaluation benchmark consists of 251 images along with comprehensive annotations.
Improving End-to-End Speech Processing by Efficient Text Data Utilization with Latent Synthesis (2023.findings-emnlp)

Copied to clipboard

Challenge: Latent Synthesis is an efficient textual data utilization framework for end-to-end speech processing models . labeled speech data are scarcer and more expensive for collection compared to textual ones .
Approach: They propose a textual data utilization framework for E2E speech processing models . they train a latent synthesizer to convert textual information into an intermediate latent representation .
Outcome: The proposed framework improves on low-resource speech recognition and spoken language understanding tasks.
TRANSIENTTABLES: Evaluating LLMs’ Reasoning on Temporally Evolving Semi-structured Tables (2025.naacl-long)

Copied to clipboard

Challenge: a recent study shows that large language models are limited in their ability to reason over time due to static datasets.
Approach: They present a dataset that includes 3,971 questions derived from over 14,000 tables . they introduce a template-based question-generation pipeline that harnesses LLMs to refine questions .
Outcome: The proposed model improves on the TRANSIENTTABLES dataset . it demonstrates that the model can reason over time, even when it is not static .
Constructing a Culinary Interview Dialogue Corpus with Video Conferencing Tool (2022.lrec-1)

Copied to clipboard

Challenge: Existing interview dialogue corpora are based on news interviews which serve the purpose of information broadcasting or entertainment.
Approach: They propose an interview dialogue corpus in the culinary domain in which interviewers play an active role to elicit culinary knowledge from the cooking expert.
Outcome: The proposed corpus consists of 308 interview dialogues, each about 13 minutes long, which add up to a total of 69,000 utterances.
Beyond Additive Fusion: Learning Non-Additive Multimodal Interactions (2022.findings-emnlp)

Copied to clipboard

Challenge: Multimodal fusion addresses the problem of analyzing spoken words in the multimodal context, including visual expressions and prosodic cues.
Approach: They propose to use multimodal fusion to separate unimodal, bimodal, and trimodal interactions in a multimodal model.
Outcome: The proposed model separates unimodal, bimodal, and trimodal interactions while not degrading predictive performance.
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation (2026.acl-long)

Copied to clipboard

Challenge: Literature review tables are essential for summarizing and comparing collections of scientific papers.
Approach: They propose to generate a database of literature review tables from a pool of papers and to model retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators.
Outcome: The proposed method improves over strong baselines while the absolute scores remain modest, underscoring the task’s difficulty.
MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Recent tabular question answering models only answer questions over a single table . multi-table operations often result in tabular outputs .
Approach: They propose a model that answers questions over multiple tables and generalizes to generate tabular answers.
Outcome: The proposed model outperforms state-of-the-art single table QA models on a multi-table QA setting.
Automated Parsing of Interlinear Glossed Text from Page Images of Grammatical Descriptions (2020.lrec-1)

Copied to clipboard

Challenge: linguistic typology is a subfield of linguistics which studies the design features of human language and the distribution of such features across the languages of the world.
Approach: They propose to parse interlinear glossed text from scanned grammars to make them machinereadable.
Outcome: The proposed technology achieves high precision and recall in the identification of examples sentences in IGT format.
Dynamic Programming in Rank Space: Scaling Structured Inference with Low-Rank HMMs and PCFGs (2022.naacl-main)

Copied to clipboard

Challenge: Hidden Markov Models (HMMs) and Probabilistic Context-Free Grammars (PCFGs) are widely used structured models.
Approach: They use tensor rank decomposition to reduce computational complexities for a subset of FGGs subsuming HMMs and PCFGs.
Outcome: The proposed model performs better on HMM modeling and unsupervised PCFG parsing than previous work.
Voice Builder: A Tool for Building Text-To-Speech Voices (L18-1)

Copied to clipboard

Challenge: a text-to-speech voice building tool is available for low-resourced languages . the tool allows researchers to run voice training experiments and listen to the resulting voice .
Approach: They propose an opensource text-to-speech (TTS) voice building tool that focuses on simplicity, flexibility, and collaboration.
Outcome: The proposed tool can help improve TTS research especially for low-resourced languages . it can be used to run voice training experiments and listen to the resulting synthesized voice .
Improving Knowledge Graph Completion with Generative Hard Negative Mining (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for knowledge graph completion (KGC) use generative methods with a self-information-enhanced training strategy to generate high-quality negatives.
Approach: They propose to leverage a sequence-to-sequence architecture to generate high-quality hard negatives from the same decoding distributions as the anchor.
Outcome: The proposed method produces high-quality negatives with good hardness and diversity on three KGC benchmarks.
OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment (2023.acl-long)

Copied to clipboard

Challenge: Speech Recognition often gets stuck in the lack of new domain utterances when training a model of new-domain speech.
Approach: They propose a training system Open-modality Speech Recognition that enables zero-shot modality transfer . they use multi-modal alignment in phoneme space to maintain multi-modality alignment .
Outcome: The proposed system achieves zero-shot modality transfer compared to existing methods . it achieves state-of-the-art performance on audio-visual speech recognition and lip-reading with 2.7% and 25.0%, respectively.
TeleMelody: Lyric-to-Melody Generation with a Template-Based Two-Stage Method (2022.emnlp-main)

Copied to clipboard

Challenge: a new lyric-to-melody generation system bridges the gap between lyrics and melodies . previous generation systems lack paired data and lack of control on generated melodie.
Approach: They develop a lyric-to-melody generation system with music template to bridge the gap between lyrics and melodies.
Outcome: The proposed system bridges the gap between lyrics and melodies by using music template.
MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQL (2025.coling-main)

Copied to clipboard

Challenge: Existing multi-hop retrieval of open-domain text-to-SQL tasks is not applicable due to the tendency to retrieve tables similar to those already retrieved but irrelevant to the question.
Approach: They propose a multi-hop table retrieval with removal task to retrieve unretrieved tables from open-domain text-to-SQL databases.
Outcome: The proposed method improves performance 5.7% over the previous state-of-the-art methods on open-domain text-to-SQL datasets.
Writing by Memorizing: Hierarchical Retrieval-based Medical Report Generation (2021.acl-long)

Copied to clipboard

Challenge: Existing methods for medical image analysis use predefined template databases or ignore hierarchical nature of medical report generation.
Approach: They propose a hierarchical retrieval mechanism to extract both report and sentence-level templates for clinically accurate report generation.
Outcome: The proposed model extracts both report and sentence-level templates for clinically accurate report generation.
Hope ‘The Paragraph Guy’ explains the rest : Introducing MeSum, the Meme Summarizer (2024.findings-emnlp)

Copied to clipboard

Challenge: a lack of large datasets for supervised learning and resource-intensive vision language models have hindered the development of meme comprehension.
Approach: They propose a framework to bridge the gap between meme comprehension and vision language models by using a multimodal dataset.
Outcome: The proposed framework outperforms existing methods in the meme comprehension test.
WhyAct: Identifying Action Reasons in Lifestyle Vlogs (2021.emnlp-main)

Copied to clipboard

Challenge: Existing systems for action recognition rely on pattern memorization and do not understand the action.
Approach: They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Outcome: The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question Answering (2022.acl-long)

Copied to clipboard

Challenge: Existing retrieval methods for knowledge base question answering are either heuristic or interwoven with the reasoning, causing reasoning on the partial subgraphs.
Approach: They propose a subgraph retrieval framework that decouples the retrieval from the subsequent reasoning process and trains subgraphs for easier reasoning.
Outcome: The proposed framework improves retrieval and QA performance over existing methods.
CateEA: Enhancing Entity Alignment via Implicit Category Supervision (2025.coling-main)

Copied to clipboard

Challenge: Existing Entity Alignment methods neglect the inherent semantic information of entities, limiting alignment precision and robustness.
Approach: They propose to combine implicit category information into multi-modal representations by generating pseudo-category labels from entity embeddings and integrating them into a multi-task learning framework.
Outcome: Experiments on benchmark datasets show that CateEA outperforms state-of-the-art methods in various settings.
VISREAS: Complex Visual Reasoning with Unanswerable Questions (2024.findings-acl)

Copied to clipboard

Challenge: Logic2Vision is a visual question-answering dataset that validates question authenticity with the corresponding image and then reasoning over it.
Approach: They propose a compositional visual question-answering dataset, VisReas, that consists of answerable and unanswerable visual queries . they use visual genome scene graphs to generate the query and the reasoning steps to generate it.
Outcome: The proposed model outperforms generative models and the existing classification models and outperformed existing models.
Learning Event Graph Knowledge for Abductive Reasoning (2021.acl-long)

Copied to clipboard

Challenge: Existing models for abductive reasoning based on formal logic lack commonsense knowledge and effective reasoning mechanism.
Approach: They propose a narrative text-based abductive reasoning task NLI with a latent variable to capture commonsense knowledge from event graph for guiding the abductive reasoning task.
Outcome: The proposed model outperforms baseline methods on the abductive reasoning task.
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis (2026.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs) are expected to be trained to act as agents in various real-world environments, but access to real systems is often restricted and manually built sandboxes are hard to scale.
Approach: They propose an automated framework for scalable tool-interaction environments via programmatic synthesis that synthesizes 191 environments and about 7K scenarios and applies them to Supervised Fine-Tuning and Reinforcement Learning for Qwen3 series models.
Outcome: The proposed framework significantly improves LLMs’ ability to solve tasks in complex environments involving multi-turn, multi-tool interactions.
Python is Not Always the Best Choice: Embracing Multilingual Program of Thoughts (2024.emnlp-main)

Copied to clipboard

Challenge: Program of Thoughts (PoT) is an approach characterized by its executable intermediate steps, which ensure the accuracy of the logical calculations in the reasoning process.
Approach: They propose a task and model agnostic approach which harnesses strength and diversity from various languages to achieve better performance across all tasks.
Outcome: The proposed approach outperforms Python Self-Consistency in almost all tasks and models and achieves comparable or superior performance on ChatGPT.
A CLARIN Transcription Portal for Interview Data (2020.lrec-1)

Copied to clipboard

Challenge: a transcription portal for audio files based on automatic speech recognition (ASR) is implemented in the CLARIN resources research network and intended for use by non-technical scholars.
Approach: They propose a transcription portal for audio files based on automatic speech recognition in various languages.
Outcome: The proposed transcription portal is implemented in the CLARIN resources research network and intended for use by non-technical scholars.
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment (2024.acl-long)

Copied to clipboard

Challenge: Recent Large Multimodal Models (LMMs) focus on visual knowledge-dimension alignment, but ignore visual knowledge.
Approach: They propose a cognitive visual-language mapper that integrates visual-linguistic knowledge alignment with a fine-grained knowledge Adapter.
Outcome: The proposed model significantly improves LMMs on knowledge-based visual question answering (VQA) it also improves the performance of other models, including GPT-4V and Gemini-Pro.
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Object hallucination has been an Achilles’ heel which hinders the broader applications of large vision-language models (LVLMs).
Approach: They propose a logical closed loop-based framework for Object Hallucination Detection and Mitigation that uses logical consistency probing to raise questions with logical correlations to determine hallucinations.
Outcome: The proposed method can be applied to all existing LVLMs and is effective and general.
MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering (2020.findings-emnlp)

Copied to clipboard

Challenge: MMFT-BERT is a multimodal fusion transformer that decomposes input modalities into different BERT instances with similar architectures, but variable weights.
Approach: They propose a multimodal fusion transformer with BERT encodings to solve Visual Question Answering (VQA) .
Outcome: The proposed method achieves SOTA results on the TVQA dataset and TVQA-Visual, an isolated diagnostic subset of TVQA, which strictly requires the knowledge of visual (V) modality based on a human annotator’s judgment.
Language Technology Programme for Icelandic 2019-2023 (2020.lrec-1)

Copied to clipboard

Challenge: a new national language technology programme for Icelandic is described . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Approach: They describe a new national language technology programme for Icelandic . the programme aims to make Icelandic usable in communication and interactions in the digital world .
Outcome: The proposed programme aims to make Icelandic usable in communication and interactions in the digital world.
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs (2026.acl-long)

Copied to clipboard

Challenge: despite significant progress, full-duplex SLMs are constrained by severe modality interference, authors say . modality interferes with acoustic and semantic modeling, making them unintelligent and unnatural . authors propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers .
Approach: They propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel.
Outcome: The proposed method significantly advances the state of the art on full-duplex benchmarks . it decouples conflicting modalities in deep layers while preserving cross-modality coherence .
ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis (2023.acl-long)

Copied to clipboard

Challenge: Multimodal sentiment analysis aims to predict the sentiment of video content.
Approach: They propose a framework that performs contrastive representation learning and contrastive feature decomposition to enhance the representation of multimodal information.
Outcome: The proposed framework outperforms baseline methods on CH-SIMS, MOSI and MOSEI datasets on a range of metrics.
Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models (2025.findings-naacl)

Copied to clipboard

Challenge: Using large vision-language models to understand cultural contexts is a critical area of research.
Approach: They conduct a thorough evaluation of multimodal models at different scales, focusing on their alignment with cultural values.
Outcome: The proposed models show that they exhibit sensitivity to cultural values but their performance is highly context-dependent.
Open-Domain Sign Language Translation Learned from Online Video (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work on sign language translation has focused mainly on data collected in controlled environments or domains, which limits its applicability to real-world settings.
Approach: They propose to use sign search as a pretext task and fusion of mouthing and handshape features to improve sign language translation in real-world settings.
Outcome: The proposed techniques produce consistent and large improvements over baseline models based on prior work.
XplainLLM: A Knowledge-Augmented Dataset for Reliable Grounded Explanations in LLMs (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have achieved remarkable success in natural language tasks, yet understanding their reasoning processes remains a significant challenge.
Approach: They propose a dataset that includes 24204 instances where each instance interprets the LLM’s reasoning behavior using knowledge graphs and graph attention networks (GAT).
Outcome: The proposed explanation framework reduces hallucinations and improves grounded explanation generation in large language models.
Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing (2020.findings-emnlp)

Copied to clipboard

Challenge: BRIDGE is a powerful sequential architecture for cross-modal semantic parsing . BRidege captures cross-modal dependencies between natural language questions and relational databases .
Approach: They propose a sequential architecture that captures cross-modal dependencies between questions and relational databases in cross-DB semantic parsing.
Outcome: The proposed architecture performs well on the well-studied Spider benchmark (65.5% dev, 59.2% test).
Word Graph Guided Summarization for Radiology Findings (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on introducing salient word information to general text summarization framework to guide selection of key content in radiology findings.
Approach: They propose a method for automatic impression generation using word graphs and a Word Graph guided Summarization model to capture critical words and their relations.
Outcome: The proposed method is validated on two datasets, OPENI and MIMIC-CXR.
Collaborative Reasoning on Multi-Modal Semantic Graphs for Video-Grounded Dialogue Generation (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for video-grounded dialogue generation do not allow information from different modalities to complement each other.
Approach: They propose a video-grounded dialogue generation model that integrates video data into pre-trained language models to allow information from different modalities to complement each other.
Outcome: The proposed model outperforms state-of-the-art models on automatic and human evaluations on two public datasets.
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have shown promising results on multimodal tasks, but remain prone to hallucinations due to their reliance on a single modality or memorizing training data without properly grounding their outputs.
Approach: They propose a training-free, tri-layer contrastive decoding with watermarking that uses a watermark-related question to identify a pivot layer and apply tri-layered contrastive coding to generate the final output.
Outcome: The proposed method reduces hallucinations and generates more visually grounded responses.
Digital Voicing of Silent Speech (2020.emnlp-main)

Copied to clipboard

Challenge: Using electromyography, we can convert silently mouthed words into audible speech . prior work focused on training speech synthesis models from vocalized data .
Approach: They propose a method of training on silent EMG by transferring audio targets from vocalized to silent signals and propose voicing task using muscle sensor measurements.
Outcome: The proposed method greatly improves intelligibility of audio generated from silent EMG compared to baseline that only trains with vocalized data.
ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-Playing (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models are increasingly utilized as role-playing agents to simulate personas in interactive settings.
Approach: They propose a role-playing agent trained to explicitly ground responses in individual identity.
Outcome: The proposed agent can generate persona-consistent responses in long-context dialogues while maintaining general instruction-following capabilities.
RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human Feedback (2025.coling-main)

Copied to clipboard

Challenge: Existing methods to mitigate hallucinations generate erroneous or fabricated information.
Approach: They propose a rank-response-based model that annotates pair-reponses and trains alignment algorithms to improve the correspondence between images and text.
Outcome: The proposed model outperforms the DPO method and outperfies existing methods on two MLLMs of different sizes and four widely used benchmarks.
An Empirical Study of Frame Selection for Text-to-Video Retrieval (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for text-to-video retrieval select a subset of frames to represent video content . current methods only explore video contents while ignoring relevancy to texts .
Approach: They propose to use a subset of frames to represent video content for TVR . they analyze six different frame selection methods to determine their effectiveness .
Outcome: The proposed method improves retrieval efficiency without sacrificing visual details . the proposed method explores the video contents while ignoring relevancy to texts .
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs (2024.lrec-main)

Copied to clipboard

Challenge: Existing work on document visual question answering fails to capture the differences and correlations between elements of a document and associated questions.
Approach: They propose a document-visual question-answering challenge that exploits element-level semantics and employs hierarchical Graph structures to capture differences and correlations between elements.
Outcome: The proposed model surpasses the state-of-the-art method and large language model in terms of Exact Match (EM) metric, demonstrating exceptional effectiveness.
CFSum Coarse-to-Fine Contribution Network for Multimodal Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing multimodal summarization models ignore the contribution of visual modalities . we propose a novel contribution network to consider different contributions of images .
Approach: They propose a Coarse-to-Fine contribution network for multimodal summarization to consider different contributions of images for summarizing.
Outcome: The proposed system outperforms baselines on the visual and textual modalities.
Unified Thinker: A General Reasoning Core for Image Generation (2026.acl-long)

Copied to clipboard

Challenge: generative models struggle with logic-intensive instruction following, exposing a persistent reasoning–execution gap.
Approach: They propose a task-agnostic reasoning architecture for general image generation . they propose pixel-level feedback to ground the Thinker's policy in pixel feedback .
Outcome: The proposed system significantly improves image reasoning and generation quality.
Scaling Laws for Code: Every Programming Language Matters (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on language-agnostic settings, neglecting the inherently multilingual nature of modern software development.
Approach: They propose a proportion-dependent scaling law that prioritizes high-utility languages . they propose PLs to have varying effects during pre-training that affect model performance .
Outcome: The proposed scaling law is based on 1000+ experiments across multiple languages and models.
ACT-Thor: A Controlled Benchmark for Embodied Action Understanding in Simulated Environments (2022.coling-1)

Copied to clipboard

Challenge: embodied AI tasks require a strong understanding of verbs and their corresponding actions.
Approach: They propose a controlled benchmark for embodied action understanding using a simulated environment and a visual feature extractor.
Outcome: The proposed benchmark achieves 81.4% accuracy and high inter-annotator agreement . the proposed model falls behind human models in a zero-shot scenario .
Putting Natural in Natural Language Processing (2023.findings-acl)

Copied to clipboard

Challenge: human language is firstly spoken and only secondarily written.
Approach: aaron carroll: human language is firstly spoken and only secondarily written . carroll says the field of NLP has overwhelmingly focused on processing written language . he says the focus is on a subset of human language which is convenient to work with .
Outcome: the ACL 2023 theme track urges the community to check the reality of the progress in NLP .
AI Knows Where You Are: Exposure, Bias, and Inference in Multimodal Geolocation with KoreaGEO (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks show coarse granularity, linguistic bias, and a neglect of multimodal privacy risks.
Approach: They propose a benchmark for visual-language models that analyzes social photos to assess location privacy risks.
Outcome: The proposed benchmarks show coarse granularity, linguistic bias, and neglect of privacy risks.
Stereotypes and Smut: The (Mis)representation of Non-cisgender Identities by Text-to-Image Models (2023.findings-acl)

Copied to clipboard

Challenge: Initial studies have pointed to the potential for harm due to predictive bias, reflecting and potentially reinforcing cultural stereotypes.
Approach: They conduct a survey among non-cisgender individuals and interviews to establish which harms affected individuals anticipate, and how they would like to be represented.
Outcome: The results show that certain non-cisgender identities are consistently (mis)represented as less human, more stereotyped and more sexualised.
Task Arithmetic can Mitigate Synthetic-to-Real Gap in Automatic Speech Recognition (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for speech recognition suffer from the synthetic-to-real gap . existing methods suffer from this distributional shift due to acoustic mismatches .
Approach: They propose to use task arithmetic to fine-tune an ASR model on synthetic data to mitigate the synthetic-to-real gap.
Outcome: The proposed method shows an improvement of 10.03% over baselines on the SLURP dataset.
Common Voice: A Massively-Multilingual Speech Corpus (2020.lrec-1)

Copied to clipboard

Challenge: Common Voice is a massively-multilingual collection of transcribed speech intended for speech technology research and development.
Approach: They propose to use Mozilla’s DeepSpeech Speech-to-Text toolkit to perform multilingual automatic speech recognition experiments.
Outcome: The proposed corpus is the largest in the public domain for speech recognition, both in terms of hours and languages.
LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles (2025.findings-emnlp)

Copied to clipboard

Challenge: Figure captions are crucial for helping readers understand and remember a figure’s key message.
Approach: They propose a dataset for personalized figure caption generation with multimodal figure profiles that provide inputs and profiles for each figure .
Outcome: The proposed dataset provides inputs and profiles for personalized figure caption generation with multimodal figure profiles.
Learning to Imagine: Visually-Augmented Natural Language Generation (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for natural language generation are pre-trained on text-only corpora, resulting in visual commonsense.
Approach: They propose a method that makes pre-trained language models learn to imagine for visually-augmented natural language generation.
Outcome: The proposed method is compatible with Transformer-based architecture.
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
Enhancing Image-to-Text Generation in Radiology Reports through Cross-modal Multi-Task Learning (2024.lrec-main)

Copied to clipboard

Challenge: Image-to-text generation relies on independent models for image understanding and natural language generation, which often exhibit a semantic gap between visual and textual information.
Approach: They propose a multi-task learning framework to leverage both visual and non-imaging data for generating radiology reports.
Outcome: The proposed framework improves performance over single-task baselines across language generation metrics and mitigates overfitting in auxiliary tasks.
TR-Rules: Rule-based Model for Link Forecasting on Temporal Knowledge Graph Considering Temporal Redundancy (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models suffer from temporal redundancy when leveraged under dynamic settings.
Approach: They propose a temporal knowledge graph extrapolation method which solves temporal redundancy issues by using cyclic rules to capture more information lurking in TKGs.
Outcome: The proposed model captures more information lurking in TKGs, and also mines and properly leverages acyclic rules, which has not been explored by existing models.
Joint Verification and Reranking for Open Fact Checking Over Tables (2021.acl-long)

Copied to clipboard

Challenge: Existing research into structured data has focused on textual data and the closed-domain setting is not reflective of real-world fact checking tasks.
Approach: They propose a joint reranking-and-verification model which fuses evidence documents in the verification component and a heuristic retrieval baseline.
Outcome: The proposed model achieves comparable performance to the closed-domain state-of-the-art on the TabFact dataset and significantly improves over a heuristic retrieval baseline.
A Survey of Automatic Text Summarization Using Graph Neural Networks (2022.coling-1)

Copied to clipboard

Challenge: Abstractive ATS involves generating factually correct and fluent sentences.
Approach: They provide an overview of the use of graph neural networks (GNNs) for automatic text summarization.
Outcome: The proposed model is based on a set of graph neural networks (GNNs) that are used to generate a concise, correct and fluent summary of a given text.
Program Structure-aware Language Models: Targeted Software Testing beyond Textual Semantics (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models for test case generation have improved branch coverage via prompt-engineered mutations, limiting their effectiveness for discovering subtle bugs and security vulnerabilities.
Approach: They propose a program structure-aware LLM framework that integrates code property graphs and code semantics to condition test case generation on execution branches.
Outcome: Experiments on real-world projects show that GLMTest improves branch accuracy from 27.4% to 50.2% on TestGenEval benchmark compared with state-of-the-art LLMs, i.e., Claude-Sonnet-4.5 and GPT-4o-mini.
Probing Logical Reasoning of MLLMs in Scientific Diagrams (2025.emnlp-main)

Copied to clipboard

Challenge: logical reasoning is key to real-world applications like science education, environmental monitoring, and medical diagnostics.
Approach: They construct visual questions that follow seven structured templates with progressively more complex reasoning involved.
Outcome: The proposed models perform logical inferences based on visual information.
IGSQL: Database Schema Interaction Graph Based Neural Model for Context-Dependent Text-to-SQL Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models on context-dependent text-to-SQL task focus on utilizing historic user inputs.
Approach: They propose a database schema interaction graph encoder to utilize historic user inputs.
Outcome: The proposed model outperforms previous state-of-the-art models on two datasets . it also outperformed existing models on the benchmark SParC and CoSQL datasets.
A Survey on Multimodal Disinformation Detection (2022.coling-1)

Copied to clipboard

Challenge: Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation.
Approach: They propose to tackle online multimodal offensive content using different modalities and combinations thereof.
Outcome: The proposed approach combines factuality and harmfulness in a framework that can be used for multiple modalities and combinations of modality.
Multimodal Reasoning with Multimodal Knowledge Graph (2024.acl-long)

Copied to clipboard

Challenge: Multimodal reasoning with large language models (LLMs) often suffers from hallucinations and the presence of deficient or outdated knowledge within LLMs.
Approach: They propose a multimodal reasoning method that leverages multimodal knowledge graphs to learn rich and semantic knowledge across modalities.
Outcome: The proposed method outperforms state-of-the-art models on multimodal question answering and multimodal analogy reasoning tasks while training on only a small fraction of parameters.
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing work on temporal sentence grounding rely on expensive video-query paired annotations . despite this, there are no ground-truth annotations in the current work .
Approach: They propose to use paired video-query and segment boundary annotations to generate temporal sentence grounding without training.
Outcome: The proposed model outperforms existing unsupervised methods and beats supervised ones on two challenging datasets.
Generating Pedagogically Meaningful Visuals for Math Word Problems: A New Benchmark and Analysis of Text-to-Image Models (2025.findings-acl)

Copied to clipboard

Challenge: Math word problems (MWPs) describe mathematical scenarios through text, requiring learners to interpret both linguistic and numerical information to derive mathematical expressions.
Approach: They propose a framework for generating pedagogically meaningful visuals from MWP text descriptions using a pre-defined visual language and a design space grounded in interviews with math teachers.
Outcome: The proposed framework illustrates the core mathematical relationships in math word problems.
Advancing Zero-shot Text-to-Speech Intelligibility across Diverse Domains via Preference Alignment (2025.acl-long)

Copied to clipboard

Challenge: Existing zero-shot text-to-speech systems struggle in challenging scenarios such as tongue twisters, repeated words, code-switching, and cross-lingual synthesis.
Approach: They propose a dataset that leverages preference alignment techniques to improve performance . they also extend the Direct Preference Optimization framework to accommodate diverse TTS architectures .
Outcome: The proposed dataset improves intelligibility, similarity, and audio quality for multiple models across domains.
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models (2025.acl-long)

Copied to clipboard

Challenge: Existing RAG frameworks rely on Automatic Speech Recognition to process speech input, which discards crucial audio information and increases computational overhead.
Approach: They propose a retrieval augmented generation framework with native, end-to-end audio support that integrates audio and text into a unified knowledge representation.
Outcome: The proposed framework can perform 10x faster than current pipelines while delivering 10x acceleration.
Fast-and-Frugal Text-Graph Transformers are Effective Link Predictors (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods that encode textual and structural information for inductive link prediction are frugal and fast at training and inference time.
Approach: They propose a Transformer-based framework that unifies textual and structural information for inductive link prediction in text-attributed knowledge graphs by encoding ego-graphs (1-hop neighbourhoods).
Outcome: The proposed framework can achieve superior performance on three popular datasets and reduce the reliance on resource-intensive encoders.
From Charts to Code: A Hierarchical Benchmark for Multimodal Models (2026.acl-long)

Copied to clipboard

Challenge: Chart2Code is a new benchmark for evaluating the natural language to chart code generation capabilities of large multimodal models.
Approach: They introduce Chart2Code, a new benchmark for evaluating the natural language to chart code generation capabilities of large multimodal models.
Outcome: The proposed benchmark is the first to scale task complexity while capturing diverse scenarios.
Less is More: Mitigating Multimodal Hallucination from an EOS Decision Perspective (2024.acl-long)

Copied to clipboard

Challenge: Extensive research has shed light on the origins of multimodal hallucinations, including the inability of vision encoders to represent finegrained visual details and model reliance on inherent parametric knowledge such as language priors and statistical biases.
Approach: They propose to use EOS to terminate generation of large multimodal models by comparing the generated text with the image to mitigate multimodal hallucinations.
Outcome: The proposed method significantly improves the hallucination performance of Large Multimodal Models without additional data or knowledge.
Generating Contextual Images for Long-Form Text (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in Text-to-Image models require short prompts that describe both the content and style of the target image.
Approach: They propose to use Large Language Models (LLMs) and Text-to-Image Models to synthesize relevant visual imagery from generic long-form text.
Outcome: The proposed models can generate high-quality images from short prompts that describe both the content and style of the target image.
Cross-Modality Relevance for Reasoning on Language and Vision (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to learn and reason over language and vision data for downstream tasks such as visual question answering (VQA) and natural language for visual reasoning (NLVR)
Approach: They propose a cross-modality relevance module that is used in an end-to-end framework to learn the relevance representation between components of various input modalities under supervision of a target task.
Outcome: The proposed approach shows competitive performance on two different language and vision tasks using public benchmarks and improves the state-of-the-art published results.
The Lexometer: A Shiny Application for Exploratory Analysis and Visualization of Corpus Data (2022.lrec-1)

Copied to clipboard

Challenge: Lexometer is a data science application that integrates data analysis and visualization functions into an easy-to-use graphical user interface.
Approach: They propose a Shiny application that integrates data analysis and visualization functions into an easy-to-use graphical user interface.
Outcome: The Lexometer integrates numerous data analysis and visualization functions into an easy-to-use graphical user interface.
VoxpopuliTTS: a large-scale multilingual TTS corpus for zero-shot speech generation (2025.coling-main)

Copied to clipboard

Challenge: Existing multilingual TTS datasets are limited in speech generation fields due to lack of quality data.
Approach: They propose to use 30,000 hours of high-quality speech data across 3 languages . they filter out low-quality text-text pairs and concatenate short transcripts .
Outcome: The proposed dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR.
Generating Coherent Sequences of Visual Illustrations for Real-World Manual Tasks (2024.acl-long)

Copied to clipboard

Challenge: Large Vision/Language Models (LVLMs) are less capable of generating accompanying image sequences.
Approach: They propose a method that integrates a Latent Diffusion Model (LDM) with an LLM to generate captions to maintain semantic coherence of the sequence.
Outcome: The proposed method is preferred by humans in 46.6% of the cases against 26.6% for the second best method.
Finnish Dialect Identification: The Effect of Audio and Text (2021.emnlp-main)

Copied to clipboard

Challenge: Finnish is a language with multiple dialects that differ in accent, morphological forms and lexical choice.
Approach: They propose an approach to automatically detect the dialect of a speaker based on a transcript and transcript with audio recording in a dataset consisting of 23 different dialects.
Outcome: The proposed method achieves 57% accuracy, compared to 85% accuracy for text and audio.
Automatic Speech Recognition Datasets in Cantonese: A Survey and New Dataset (2022.lrec-1)

Copied to clipboard

Challenge: In this paper, we address the problem of data scarcity for the Hong Kong Cantonese language . due to the popularization of deep learning, ASR technology has led to a significant improvement in recognizing many languages.
Approach: They propose to use a dataset to analyze the data available for the Hong Kong Cantonese language . they use zh-HK as a source and a state-of-the-art ASR model to build a powerful model .
Outcome: The proposed model improves on the biggest existing dataset, Common Voice zh-HK.
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work has adapted vision-and-language models to generative tasks like image captioning.
Approach: They propose an extension to LXMERT with training refinements to generate images from text.
Outcome: The proposed model can generate images from pieces of text while still being comparable to existing models.
How to Understand “Support”? An Implicit-enhanced Causal Inference Approach for Weakly-supervised Phrase Grounding (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Weakly-supervised Phrase Grounding (WPG) largely ignore the implicit phrase-region matching relations, rendering it arduous to explore the semantic nature of phrases.
Approach: They propose an Implicit-Enhanced Causal Inference approach to address the challenges of modeling the implicit relations and highlighting them beyond the explicit.
Outcome: The proposed approach outperforms the state-of-the-art baselines on an implicit-enhanced dataset.
MultiSubs: A Large-scale Multimodal and Multilingual Dataset (2022.lrec-1)

Copied to clipboard

Challenge: a large-scale multimodal and multilingual dataset is used to facilitate research on visual grounding of words to images in their contextual usage in language.
Approach: They propose a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language.
Outcome: The proposed dataset will facilitate research on visual grounding of words in their contextual usage in language.
CART: A Generative Cross-Modal Retrieval Framework With Coarse-To-Fine Semantic Modeling (2025.acl-long)

Copied to clipboard

Challenge: Cross-modal retrieval tasks are used to retrieve data from one modality or another based on a query from another modality.
Approach: They propose a generative cross-modal retrieval framework based on coarse-to-fine semantic modeling . they propose combining K-Means and RQ-VAE to discretize multimodal data into token sequences that support autoregressive generation.
Outcome: The proposed framework achieves excellent performance and efficiency in multimodal retrieval tasks.
Multi-modal Stance Detection: New Datasets and Model (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for stance detection for pure texts have limited results to multi-modal content.
Approach: They propose a multi-modal stance detection framework that leverages target information to learn multi-modal stance features from textual and visual modalities.
Outcome: The proposed framework achieves state-of-the-art in multi-modal stance detection on five datasets based on Twitter .
Caption Enriched Samples for Improving Hateful Memes Detection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for classifying memes are difficult to perform, with human accuracy only about 85% . recent state-of-the-art models perform considerably less accurately, achieving up to 64.73% accuracy.
Approach: They propose to use an off-the-shelf caption generator to capture the first image and overlayed text.
Outcome: The proposed tool improves classification accuracy for unimodal and multimodal models . the proposed tool can be used to model the contrast between image content and overlayed text .
Towards relation extraction from speech (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for extracting relations from speech have been neglected due to the nature of speech.
Approach: They propose a listening information extraction task that uses speech to extract relation extraction from speech . they use a text-to-speech system and crowd-sourced native English speakers to test the task .
Outcome: The proposed task extracts semantic relationships from speech data using a new model . the proposed task is more challenging than the existing method due to the characteristics of speech .
The MERSA Dataset and a Transformer-Based Approach for Speech Emotion Recognition (2024.acl-long)

Copied to clipboard

Challenge: Existing models for speech emotion recognition lack a comprehensive dataset to design accurate models.
Approach: They propose to use a multimodal dataset to build a model that integrates pre-trained wav2vec 2.0 and BERT to learn hidden representations from fused representations of speech and text.
Outcome: The proposed model predicts emotions on dimensions of arousal, valence, and dominance . it achieved competitive results on the MSP-PODCAST dataset .
Opinions in Interactions : New Annotations of the SEMAINE Database (2022.lrec-1)

Copied to clipboard

Challenge: a new method for the detection of opinions in interactions is proposed . a dataset of dyadic interactions is annotated continuously in two affective dimensions related to the emotions .
Approach: They propose to annotate opinions over a multimodal corpus of dyadic interactions . they use a d-acting algorithm to annnotate the opinions of a speaker .
Outcome: The proposed method allows to obtain a precise annotation regarding the opinion of a speaker.
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results .
Approach: They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach .
Outcome: The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts.
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks (2022.lrec-1)

Copied to clipboard

Challenge: Audiobook readers play with their voices to emphasize some text passages, highlight discourse changes or significant events, or in order to make listening easier and entertaining.
Approach: They propose to modify the narrator’s voice to fit the context of the story, such as the character who is speaking, using voice conversion.
Outcome: The proposed method improves the quality of the voice conversion system and the speaker similarity.
Boosting Multi-modal Keyphrase Prediction with Dynamic Chain-of-Thought in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Multi-modal keyphrase prediction (MMKP) aims to produce concise, informative phrases that capture the essence of cross-modal inputs.
Approach: They propose to use vision-language models to generate conclusive phrases using multiple modalities of input information.
Outcome: The proposed methods outperform existing methods on absence and unseen scenarios and overestimate model capability due to overlap in training tests.
GNN-SL: Sequence Labeling Based on Nearest Examples via GNN (2023.findings-acl)

Copied to clipboard

Challenge: Existing sequence labeling algorithms can be decomposed into two parts .
Approach: They propose a graph neural networks sequence labeling (GNN-SL) that augments the vanilla SL model output with similar tagging examples retrieved from the whole training set.
Outcome: The proposed model performs well on three sequence labeling tasks.
The Revolution of Multimodal Large Language Models: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have led to the development of multimodal large language model.
Approach: They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies.
Outcome: The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities.
Information Screening whilst Exploiting! Multimodal Relation Extraction with Feature Denoising and Multimodal Topic Modeling (2023.acl-long)

Copied to clipboard

Challenge: Existing research on multimodal relation extraction (MRE) faces internal-information over-utilization and external-information under-exploitation.
Approach: They propose a framework that implements internal-information screening and external-information exploiting to address these challenges.
Outcome: The proposed framework outperforms the current best model on the benchmark dataset.
Pay Attention to Implicit Attribute Values: A Multi-modal Generative Framework for AVE Task (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to extract attribute values from product descriptions are incomplete and noisy due to the tedious nature of this task.
Approach: They propose a framework to extract attributes from product descriptions to acquire implicit attributes in addition to the explicit ones.
Outcome: The proposed framework outperforms existing methods on the extraction of implicit attribute values while achieving comparable performance for the explicit ones.
Finding needles in a haystack: Sampling Structurally-diverse Training Sets from Synthetic Data for Compositional Generalization (2021.emnlp-main)

Copied to clipboard

Challenge: Recent research shows that automatic generation of synthetic utterance-program pairs can alleviate the first problem, but its potential for the second has thus far been under-explored.
Approach: They propose to generate synthetic utterance-program pairs for improving compositional generalization in semantic parsing by using structurally-diverse examples.
Outcome: The proposed approach leads to dramatic improvements in compositional generalization and moderate improvements in the traditional i.i.d setup.
Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents (2025.acl-long)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) are developing but lack external feedback . there is no clear on how to select reward models for agents .
Approach: They propose a benchmark to evaluate agent reward modeling ability in MLLMs . they use multiple dimensions and real-world agent scenarios evaluation .
Outcome: The proposed benchmark evaluates agent performance in multimodal large language models . it covers perception, planning, and safety with 7 scenarios and is highly difficult and high-quality .
Ask in Any Modality: A Comprehensive Survey on Multimodal Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) suffer from hallucinations and outdated knowledge due to their reliance on static training data.
Approach: They review training strategies, robustness enhancements, loss functions, and agent-based approaches and outline open challenges and future directions to guide research in this evolving field.
Outcome: The proposed model improves accuracy and accuracy while integrating external dynamic information for improved factual grounding.
TellWhisper: Tell Whisper Who Speaks When (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches decouple temporal modeling and speaker modeling when addressing 'when' and 'who' . a new framework that couples temporal structure with speaker dynamics is proposed to address these limitations .
Approach: They propose a framework that couples temporal and speaker identity within the speech encoder . they propose TS-RoPE, a time-speaker rotary positional encoding that partitions Query/Key channels into temporal, speaker subspaces and applies region-specific rotations to align "when" and "who" cues in selfattention.
Outcome: The proposed framework couples temporal structure with speaker dynamics in speech encoder . it uses frame-level speaker activity to estimate speaker-activity estimates .
DisCo_Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec (2026.acl-long)

Copied to clipboard

Challenge: DisCo-Speech is a zero-shot controllable text-to-speech framework . standard codecs entangle timbre and prosody, which hinders independent control in continuation-based LMs.
Approach: They propose a disentangled speech codec and an LM-based generator to solve this problem . they propose fusion and reconstruction that merges content and prosody into unified tokens .
Outcome: DisCo-Speech achieves competitive voice cloning and superior zero-shot prosody control.
MemeInterpret: Towards an All-in-One Dataset for Meme Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing research has not explored meme captioning's decomposition into subtasks or its connections to other CMU tasks.
Approach: a new meme corpus is built upon the Facebook Hateful Memes dataset . it contains meme captions, corresponding surface messages and relevant background knowledge .
Outcome: a new corpus of meme captions and surface messages unifies three major categories of CMU tasks for the first time.
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to mitigating vision-knowledge conflict in Large Language Models (MLLMs) are not effective and can be further scaled.
Approach: They propose a framework to generate inputs to simulate and evaluate vision-knowledge conflict in Multimodal Large Language Models (MLLMs) using original images and 1,122 high-quality question-answer pairs, they propose 'a diagnostic benchmark'
Outcome: The proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in Multimodal Large Language Models (MLLMs).
Fintan - Flexible, Integrated Transformation and Annotation eNgineering (2020.lrec-1)

Copied to clipboard

Challenge: Fintan is a platform for converting heterogeneous linguistic resources to RDF.
Approach: They introduce Fintan for converting heterogeneous linguistic resources to RDF with its modular architecture, workflow management and visualization features.
Outcome: The Fintan platform is designed to transform linguistic resources to graphs and graphs.
Bridging Intuitive Associations and Deliberate Recall: Empowering LLM Personal Assistant with Graph-Structured Long-term Memory (2025.findings-acl)

Copied to clipboard

Challenge: Large language models (LLMs)-based personal assistants struggle to capture entity relationships and handle multiple intents effectively.
Approach: They propose a graph-structured memory framework that mimics human cognitive processes and an event-centric memory graph.
Outcome: The proposed framework outperforms retrieval and QA methods across long-term dialogue benchmarks and enables more human-like memory systems.
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation (2025.findings-acl)

Copied to clipboard

Challenge: Traditional video topic segmentation methods struggle to discern topical transitions . supervised approaches have improved performance on video action or scene segmentation .
Approach: They propose a new task for video topic segmentation that enhances multimodality alignment and fusion by exploring different architectures using Cross-Attention and Mixture of Experts.
Outcome: The proposed model improves on educational videos, in the form of lectures . it combines cross-attention and mixture of experts to strengthen multimodality alignment and fusion .
UniCM: A Unified Consistency Model For Efficient Multimodal Generation and Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Consistency models (CMs) have shown promise in the efficient generation of both image and text.
Approach: They propose to use a discrete token for both image and text generation to achieve a unified denoising perspective.
Outcome: The proposed model outperforms SD3 on GenEval and Image Reward while being 1.5 faster at long-sequence generating speed.
CP-BCS: Binary Code Summarization Guided by Control Flow Graph and Pseudo Code (2023.emnlp-main)

Copied to clipboard

Challenge: Current work on understanding assembly code is oriented towards generating function names, which involve numerous abbreviations that make them confusing.
Approach: They propose a control flow graph and pseudo code guided binary code summarization framework to learn the comprehensive binary function execution behavior and logic semantics.
Outcome: The proposed framework improves the efficiency of reverse engineering on 3 different binary optimization levels for 3 different computer architectures.
Soundwave: Less is More for Speech-Text Alignment in LLMs (2025.acl-long)

Copied to clipboard

Challenge: Existing end-to-end speech large language models rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth.
Approach: They propose a training strategy and a novel architecture to address representation space gap and sequence length inconsistency in speech and text.
Outcome: The proposed model outperforms other advanced speech LLMs in speech translation and AIR-Bench speech tasks with only a fraction of the training data.
MONAQ: Multi-Objective Neural Architecture Querying for Time-Series Analysis on Resource-Constrained Devices (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent efforts in hardware-aware neural architecture search (NAS) automate architecture discovery for specific platforms; however, none focus on general time-series analysis with edge deployment.
Approach: They propose a framework that reformulates NAS into ***M***ulti-***O***bjective ***N***eural ***A***rchitecture ***Q***uerying tasks.
Outcome: Experiments on 15 datasets show that the proposed framework outperforms both handcrafted models and NAS baselines while being more efficient.
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos (2025.findings-emnlp)

Copied to clipboard

Challenge: a new multimodal dataset of stand-up comedies is proposed to improve humor detection . the dataset is the biggest available for this type of task, and the most diverse .
Approach: They propose a method to enhance the automatic laughter detection based on Audio Speech Recognition errors.
Outcome: The proposed method improves existing models of humor detection by using audio speech recognition errors.
Who Wrote When? Author Diarization in Social Media Discussions (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for author diarization are unable to detect stylistic shifts in a text .
Approach: They propose a framework that integrates pre-trained neural representations of writing style with author-conditional encoder-decoder diarization.
Outcome: The proposed framework is able to attribute comments in online discussions to individual authors.
UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook (2025.acl-long)

Copied to clipboard

Challenge: Existing neural audio codecs are not capable of handling multi-domain audio data . et al., 2023) integrate speech modality with text-based large language models .
Approach: They propose a unified audio codec with a single codebook to support multi-domain audio data . they propose combining a mix-of-experts strategy and a partitioned domain-adaptive codebook method .
Outcome: The proposed codec outperforms existing codecs on acoustic and semantic representation capabilities.
Med-VRAgent: A Framework for Medical Visual Reasoning-Enhanced Agents (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Language Models (VLMs) have shown strong performance in tasks like radiology report generation but struggle with hallucinations, vague descriptions, Inconsistent logic and poor localization.
Approach: They propose a framework for medical visual reasoning based on Visual Guidance and Self-Reward paradigms and Monte Carlo Tree Search to improve the model's visual reasoning capabilities.
Outcome: The proposed framework outperforms existing models on multiple medical VQA benchmarks.
Aligned Multi-View Scripts for Universal Chart-to-Code Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for chart-to-code generation are largely Python-centric, limiting practical use and overlooking a critical source of supervision.
Approach: They propose a chart-to-code generation tool that converts a graph image into an executable plotting script.
Outcome: The proposed method outperforms existing systems and is competitive with proprietary systems.
PunMemeCN: A Benchmark to Explore Vision-Language Models’ Understanding of Chinese Pun Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Pun memes combine wordplay with visual elements to create humor, irony, or other rhetorical effects.
Approach: They propose a benchmark to assess Chinese pun memes' processing capabilities across three progressive tasks: pun meme detection, sentiment analysis, and chat-driven meme response.
Outcome: The proposed model can detect pun memes, analyze sentiments, and respond to chats, while ignoring homophone wordplay.
Information Flow Routes: Automatically Interpreting Language Models at Scale (2024.emnlp-main)

Copied to clipboard

Challenge: Current state-of-the-art language models (LMs) are built on top of the Transformer architecture.
Approach: They propose to build graphs where nodes correspond to token representations and edges to computations . they show that attention heads and subword merging heads are important .
Outcome: The proposed model can analyze behavior for specific types of predictions, or different domains.
MolCA: Molecular Graph-Language Modeling with Cross-Modal Projector and Uni-Modal Adapter (2023.emnlp-main)

Copied to clipboard

Challenge: Language Models (LMs) have demonstrated impressive molecule understanding ability on 1D text-related tasks, but lack 2D graph perception, a critical ability of human professionals in comprehending molecules’ topological structures.
Approach: They propose to combine a cross-modal projector and a uni-modal adapter to enable an LM to understand both text- and graph-based molecular contents via a Q-Former.
Outcome: The proposed model outperforms the baselines on tasks such as molecule captioning, IUPAC name prediction, and molecule-text retrieval.
TransferCVLM: Transferring Cross-Modal Knowledge for Vision-Language Modeling (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent large vision-language multimodal models pre-trained with huge amount of image-text pairs show remarkable performances in downstream tasks.
Approach: They propose a method of efficient knowledge transfer that integrates pre-trained uni-modal models into a combined vision-language model without pre-training . they propose to fine-tune the model and transfer multimodal knowledge from a teacher vision-linguistic model to the CVLM for each task application.
Outcome: The proposed method outperforms existing vision-language models in downstream tasks.
We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus more on end-to-end performance, but neglect the underlying principles of knowledge acquisition and generalization.
Approach: They propose a benchmark specifically designed to explore the problem-solving principles by decomposing 6.5K visual math problems into 10.9K step-level questions for evaluation.
Outcome: The proposed benchmark covers 6.5K visual math problems and 10.9K step-level questions spanning 5 layers of knowledge granularity and 67 hierarchical knowledge concepts.
SGG-R 3: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for scene graph generation lack task-specific structured reasoning and sparse, long-tailed relation distributions.
Approach: They propose a structured reasoning framework that integrates task-specific Chain-of-Thought and reinforcement learning with group sequence policy optimization to achieve unbiased scene graph generation.
Outcome: The proposed framework achieves superior performance on two benchmarks.
Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Recent approaches to optimize communication topology rely on single-sample policy gradients with absolute rewards.
Approach: They propose a topology optimization framework that integrates Group Relative Policy Optimization.
Outcome: The proposed topology optimization framework outperforms state-of-the-art methods on reasoning and code generation benchmarks.
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for expert parallelism inference suffer from a significant efficiency bottleneck . existing methods fail to address information heterogeneity and modality dynamics .
Approach: They propose a training-free inference framework that scales experts without training . they propose an Entropy-Weighted Load mechanism to quantify the semantic value of visual tokens .
Outcome: Experiments show that MACS outperforms existing methods on multimodal benchmarks.
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have shown promise in MER, but their internal decision-making mechanisms under modality conflict and missingness remain underexplored.
Approach: They propose a multimodal large language model that can detect and control modality conflicts and missing subsets by a lightweight mechanism that detects and controls modality conflict.
Outcome: The proposed framework improves performance across settings, showing it can handle conflict and missing behaviors.
How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads (2025.emnlp-main)

Copied to clipboard

Challenge: Despite advances in Large Vision Language Models, a gap remains in their interpretability and performance.
Approach: They identify the Optical Character Recognition Head (OCR Head) heads that are more efficient at recognizing text from images.
Outcome: The Optical Character Recognition Head (OCR Head) is identified as the most efficient head for recognizing text from images.
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains.
Approach: They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective.
Outcome: The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs.
CamoQuery: Language-Guided Reasoning Camouflaged Object Segmentation (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for camouflaged object segmentation are limited to vision-only mask prediction under fixed task assumptions.
Approach: They propose a language-guided reasoning camouflaged object segmentation task that generates an intent-consistent segmentation mask from an image and an implicit query text instruction.
Outcome: The proposed task can generate an intent-consistent segmentation mask from a camouflaged image and an implicit query text instruction.
Differentiated Vision: Unveiling Entity-Specific Visual Modality Requirements for Multimodal Knowledge Graph (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to extract features from images of entities overlook varying relevance of visual information across entities.
Approach: a new model integrates structural and multimodal information of entities into a multimodal knowledge graph . a model evaluates the necessity of visual modality for each entity based on its attributes .
Outcome: The proposed model improves on existing methods by adjusting visual data to different entity types.
Attention as Selector: Unlocking VLM Attention for Long Document Page Retrieval (2026.acl-long)

Copied to clipboard

Challenge: Existing page-level retrieval methods lack query–page interaction before similarity scoring . Existing methods require large-scale datasets to align visual and textual embeddings .
Approach: They propose a retrieval framework that utilizes attention mechanisms inside VLMs for page selection.
Outcome: The proposed retrieval framework outperforms embedding-based retrieval methods on four long-document benchmarks.
Collaborative Rational Speech Act: Pragmatic Reasoning for Multi-Turn Dialog (2025.emnlp-main)

Copied to clipboard

Challenge: Existing extensions of Rational Speech Act face challenges in scaling to multi-turn, collaborative scenarios.
Approach: They propose a Rational Speech Act extension that optimizes a gain function adapted from rate-distortion theory to model multi-turn dialog by optimizing a model gain . they demonstrate the effectiveness of CRSA on referential games and template-based doctor–patient dialogs in the medical domain.
Outcome: The proposed model yields more consistent, interpretable, and collaborative behavior than baselines, paving the way for more pragmatic and socially aware language agents.
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant (2024.emnlp-main)

Copied to clipboard

Challenge: Existing knowledge-aware text-based visual question answering methods are based on textual entities in images.
Approach: They propose a visual text entity linking module that harnesses a state-of-the-art visual text recognition engine and the power of a large multimodal model to perform visual text-entity linking.
Outcome: The proposed approach surpasses the previous best approach by 23.3% on an absolute scale and establishes a new state of the art.
Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on simple image-text interactions, overlooking complex visual formats like charts.
Approach: They propose a semi-automatic framework for generating evaluation samples through multi-modal keypoint extraction, knowledge graph construction, and qa pair synthesis.
Outcome: The proposed framework generates 4,738 question-answering pairs across 8 domains from real-world documents.
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture (2025.findings-emnlp)

Copied to clipboard

Challenge: Long-context Large Language Models (MLLMs) are critical for video understanding and image analysis.
Approach: They propose a hybrid architecture that integrates Mamba and Transformer blocks . they introduce data construction methods that capture both temporal and spatial dependencies .
Outcome: The proposed model achieves competitive results across various benchmarks while maintaining high throughput and low memory consumption.
Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Media (2025.findings-emnlp)

Copied to clipboard

Challenge: Self-disclosure can provide psychological comfort but can also pose privacy concerns . a lack of high-quality corpora, analysis, and methods for detection is limiting research .
Approach: They construct a high-quality text-image corpus on Chinese multimodal social media platforms . they analyze the distribution of self-disclosure types, modality preferences, user intent .
Outcome: The proposed corpus analyzes self-disclosure behaviors on Chinese social media platforms . it fine-tunes five multimodal large language models to enhance self-discovery detection .
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Currently, vision-Language Models are optimized for direct visual question-answering tasks.
Approach: They propose a visual-language-based VLM that prioritizes reasoning within the perception process.
Outcome: The proposed model outperforms existing models and domain-specific open-source models in the chemical domain.
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for I-MCoT fail to capture dynamic needs of vision-language models . existing methods rely on attention signals, which are unreliable under severe granularity imbalance between brief textual query and informative image.
Approach: They propose a framework that integrates specially selected visual evidence into the context of Vision-Language Models (VLMs) they propose 'AIM-CoT' to improve evidence selection and insertion triggering .
Outcome: Experiments across three benchmarks and four backbones demonstrate the proposed framework’s consistent superiority.
Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset (2025.findings-emnlp)

Copied to clipboard

Challenge: Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets.
Approach: They propose to construct a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding.
Outcome: The proposed dataset covers ten culturally significant domains covering all Arab countries and includes two evaluation benchmarks (PEARL and PEARL-LITE) and a specialized subset (PearL-X).
ProcVQA: Benchmarking the Effects of Structural Properties in Mined Process Visualizations on Vision–Language Model Performance (2025.findings-emnlp)

Copied to clipboard

Challenge: Vision-Language Models have shown impressive capabilities and notable failures in data visualization understanding tasks.
Approach: They propose a benchmark to analyze how specific properties within a visualization type affect VLM performance.
Outcome: The proposed benchmark examines how specific properties affect VLM performance . it shows that models exhibit steep drops on multi-hop reasoning and extraction errors increase with edge density .
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction (2025.acl-long)

Copied to clipboard

Challenge: Existing methods focus on replicating dialogues in textual form, neglecting the role’s voice traits as a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios.
Approach: They propose a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency.
Outcome: The proposed model exhibits role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses.
D2CS - Documents Graph Clustering using LLM supervision (2025.findings-emnlp)

Copied to clipboard

Challenge: Document clustering does not inherently ensure thematic consistency.
Approach: They propose a framework that constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters.
Outcome: The proposed framework constructs a similarity graph over document embeddings and applies iterative graph-based clustering algorithms to partition the corpus into initial clusters.
AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA (2026.findings-acl)

Copied to clipboard

Challenge: Existing audio question answering benchmarks emphasize sound event classification or caption-grounded queries.
Approach: They propose a large-scale, real-world audio question answering benchmark to evaluate audio reasoning beyond surface-level acoustic recognition.
Outcome: The proposed model achieves 32.13% accuracy while demonstrating comprehension of audio . state-of-the-art models perform poorly, with average accuracy below 8.86%.
Benchmarking Deflection and Hallucination in Large Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks overlook conflicts between visual and textual evidence and the importance of generating deflections when incomplete knowledge is retrieved.
Approach: They propose a dynamic curation pipeline that preserves benchmark difficulty over time . they propose 'vlm-DeflectionBench' benchmark to probe model behaviour under conflicting evidence .
Outcome: The proposed benchmarks overlook conflicts between visual and textual evidence and are prone to obsolescence . the proposed benchmark is based on 2,775 samples spanning diverse retrieval settings .
Beyond Chunks and Graphs: Retrieval-Augmented Generation through Triplet-Driven Thinking (2026.findings-acl)

Copied to clipboard

Challenge: Retrieval-augmented generation (RAG) is critical for reducing hallucinations and incorporating external knowledge into Large Language Models (LLMs).
Approach: They propose a framework that leverages an LLM to decompose questions into searchable triplets with placeholders.
Outcome: Empirical results show that T2RAG outperforms state-of-the-art multi-round and Graph RAG methods while reducing retrieval costs by up to 45%.
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing vision-language models lack spatial reasoning capability, despite their ability to comprehend spatial arrangements and model structural relations.
Approach: They propose a benchmark to evaluate vision-language models' spatial perception, structural understanding, and reasoning capabilities by minimizing reliance on domain-specific knowledge.
Outcome: The proposed benchmark is based on 1,100 carefully curated real-world images with high spatial complexity.
SPOTTER: A Framework for Investigating Convention Formation in a Visually Grounded Human-Robot Reference Task (2024.lrec-main)

Copied to clipboard

Challenge: Existing research has shown that conventions arise in repeated interactions over the same task, leading to a decrease in utterance length while maintaining informative content.
Approach: They propose to elicit conventions for members of an inner circle of well-known individuals in common ground, as opposed to individuals from an outer circle, who are unfamiliar.
Outcome: The proposed game platform elicits conventions for familiar and unfamiliar individuals in human-robot interaction.
NBDESCRIB: A Dataset for Text Description Generation from Tables and Code in Jupyter Notebooks with Guidelines (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for Jupyter Notebooks focus on generating cell-level descriptions from code snippets or table outputs independently.
Approach: They propose a task to generate personalized cell-level descriptions using code, tables, and user-written guidelines in Jupyter Notebooks.
Outcome: The proposed task combines code, tables, and user-written guidelines with personalized descriptions to evaluate the performance of existing models.
Text-to-Multimodal Retrieval with Bimodal Input Fusion in Shared Cross-Modal Transformer (2024.lrec-main)

Copied to clipboard

Challenge: Multimodal video retrieval systems are needed for multimodal content retrieval . multimodal video search systems are sub-optimal for multi-modal content representations .
Approach: They propose a model that learns retrieval cues for the textual query from multiple modalities and a shared embedding space with task-specific contrastive loss functions.
Outcome: The proposed model outperforms state-of-the-art methods on the MSR-VTT and YouCook2 datasets and shows significant improvements from baseline.
CPR-RAG: Clinical Prior-Regularized Retrieval for Anatomy-Aware 3D CT Report Generation (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to grounding radiology reports from 3D volumetric data are limited due to visual-semantic ambiguity and lack of "normal" context.
Approach: They propose a model-agnostic retrieval-augmented generation framework that integrates clinical priors into the retrieval process.
Outcome: The proposed model improves clinical efficacy across state-of-the-art models.
Multi-Frequency Contrastive Decoding: Alleviating Hallucinations for Large Vision-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies attribute object hallucinations to linguistic priors and data biases . MFCD method removes hallucinian distribution in the original output distribution .
Approach: They propose a method that removes the hallucination distribution in the original output distribution . they propose MFCD to mitigate hallucinism in large visual-language models .
Outcome: The proposed method reduces hallucination distributions without training or external tools . the proposed method can be applied to various LVLMs without modifying model architecture or training .
Predicting Implicit Arguments in Procedural Video Instructions (2025.acl-long)

Copied to clipboard

Challenge: Prior SRL benchmarks often miss implicit arguments, leading to incomplete understanding.
Approach: They propose a dataset that necessitates inferring implicit and explicit arguments from contextual information in multimodal cooking procedures.
Outcome: The proposed dataset achieves a 17% relative improvement in F1-score for what-implicit and a 14.7% improvement for where/with-implicative semantic roles over GPT-4o.
GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL (2026.acl-long)

Copied to clipboard

Challenge: despite growing interest in NL2GQL, benchmarking progress has been constrained by the lack of resources that are simultaneously large-scale, cross-domain, and cross-dialect.
Approach: They propose a framework that integrates NL2SQL-to-NL2GQL conversion with graph-native data generation.
Outcome: The proposed framework supports execution-based evaluation on Cypher and ISO-GQL, covering hundreds of graph databases and over 20k natural language questions for each dialect.
Unraveling Spontaneous Speech Dimensions for Cross-Corpus ASR System Evaluation for French (2024.lrec-main)

Copied to clipboard

Challenge: 'spontaneous speech' is a catch-all term used for situations like speaking with a friend, being interviewed on radio/TV or giving a lecture.
Approach: They propose to use four dimensions to describe spontaneous speech variation in automatic speech recognition systems.
Outcome: The proposed system can be used to predict the WER of speech recognition systems on face-to-face interactions.
IMOL: Incomplete-Modality-Tolerant Learning for Multi-Domain Fake News Video Detection (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for fake news video detection focus on a specific domain and assume multiple modalities.
Approach: They propose an incomplete-modality-tolerant learning framework for fake news video detection . they use cross-modal consistency to reconstruct missing modalities and transferable knowledge through cross-sample reasoning .
Outcome: The proposed framework improves performance and robustness of multi-domain fake news video detection while generalizing to unseen domains under incomplete modality conditions.
HintsOfTruth: A Multimodal Checkworthiness Detection Dataset with Real and Synthetic Claims (2025.acl-long)

Copied to clipboard

Challenge: Identifying checkworthy claims is the first step, but detection methods struggle with content that is (1) multimodal, (2) from diverse domains, and (3) synthetic.
Approach: They propose a dataset for multimodal checkworthiness detection with 27K real-world and synthetic image/claim pairs.
Outcome: The proposed dataset compares lightweight text-based encoders to multimodal models but only focus on claim-like content.
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap .
Approach: They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models.
Outcome: The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting .
AROMA: Augmented Reasoning Over a Multimodal Architecture for Virtual Cell Genetic Perturbation Modeling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for virtual cell genetic perturbation modeling suffer from unconstrained reasoning, uninterpretable predictions, and retrieval signals that are weakly aligned with regulatory topology.
Approach: They propose an Augmented Reasoning Over a Multimodal Architecture for virtual cell genetic perturbation modeling.
Outcome: The proposed model outperforms existing methods across multiple cell lines and remains robust under zero-shot evaluation on unseen cells.
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks rely heavily on text-based evaluation and largely ignore paralinguistic cues such as prosody, emotion, and speaker traits.
Approach: They propose a speech-native benchmark for evaluating instruction-following S2S models with explicit assessment of both semantic understanding and paralinguistic expression.
Outcome: The proposed system enables more natural, robust, and human-aligned speech agents.
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts (2025.emnlp-main)

Copied to clipboard

Challenge: Understanding and reasoning over text within visual contexts poses a significant challenge for Vision-Language Models.
Approach: They propose a benchmark for Korean Reading and rEasoning in Text-rich VQA Attuned to diverse visual contexts to address this challenge.
Outcome: The proposed benchmark is tailored for Korean reading and rEasoning in text-rich VQA attuned to diverse visual contexts.
HCFD: A Benchmark for Audio Deepfake Detection in Healthcare (2026.findings-acl)

Copied to clipboard

Challenge: a new task for detecting codec-fakes under pathological speech conditions is presented . we focus on codec based synthetic speech since neural codec decoding is a core building block in speech generation pipelines.
Approach: They propose a new task for detecting codec-fakes under pathological speech conditions . they focus on codec based synthetic speech since neural codec decoding is a core building block in speech pipelines .
Outcome: The proposed framework outperforms speech-based models on Healthcare CodecFake . it achieves the strongest performance on the task across clinical conditions and codecs .
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity (2026.findings-acl)

Copied to clipboard

Challenge: TableVista evaluates multimodal table reasoning under visual and structural complexity . current models struggle to maintain reasoning consistency when structural complexity combined with visually integrated presentations.
Approach: They propose a benchmark for evaluating multimodal table reasoning under visual and structural complexity.
Outcome: The proposed model performs poorly on visual and structural complexity.
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound (2026.acl-long)

Copied to clipboard

Challenge: Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos.
Approach: They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues.
Outcome: The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor .
CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations (2026.findings-acl)

Copied to clipboard

Challenge: Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment.
Approach: They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation.
Outcome: The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations.
Mind’s Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of multimodal large language models (MLLMs) have demonstrated compelling visual understanding in recent years.
Approach: They propose a multimodal large language model with eight visuo-cognitive tasks inspired by classic human intelligence tests organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation.
Outcome: The proposed frameworks are based on eight visuo-cognitive tasks inspired by human intelligence tests and organized under a novel A–R–T taxonomy: Abstraction, Relation, and Transformation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations