Papers by Tanmoy Chakraborty

59 papers
A Survey on Multimodal Disinformation Detection (2022.coling-1)

Copied to clipboard

Challenge: Recent years have witnessed the proliferation of offensive content online such as fake news, propaganda, misinformation, and disinformation.
Approach: They propose to tackle online multimodal offensive content using different modalities and combinations thereof.
Outcome: The proposed approach combines factuality and harmfulness in a framework that can be used for multiple modalities and combinations of modality.
Harmonizing Code-mixed Conversations: Personality-assisted Code-mixed Response Generation in Dialogues (2024.findings-eacl)

Copied to clipboard

Challenge: blending multiple languages within a single conversation presents a formidable challenge, given the wide-ranging variations influenced by individual speaking styles and cultural backgrounds.
Approach: They propose a novel approach to harness the Big Five personality traits acquired in an unsupervised manner from code-mixed conversations to bolster the performance of response generation.
Outcome: The proposed approach enhances contextual relevance and performance of the proposed model by combining personality traits with dialogue context.
POSIX: A Prompt Sensitivity Index For Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are sensitive to minor variations in prompts, such as spelling errors, alteration of wording or the prompt template.
Approach: They propose a PrOmpt Sensitivity IndeX to measure prompt sensitivity . they use this to compare prompt sensitability of various open source LLMs .
Outcome: The proposed method can measure and compare prompt sensitivity of open source LLMs.
HIT - A Hierarchically Fused Deep Attention Network for Robust Code-mixed Language Representation (2021.findings-acl)

Copied to clipboard

Challenge: linguistics and morphology of resource-short code-mixed texts remain a key challenge in text processing.
Approach: They propose a hierarchical transformer-based framework that captures the semantic relationship among words and hierarchically learns sentencelevel semantics using a fused attention mechanism.
Outcome: The proposed framework improves on one European and five Indic languages on four NLP tasks on eleven datasets.
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving (2026.findings-eacl)

Copied to clipboard

Challenge: Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance.
Approach: They propose a Spatial Comprehension-Infused Symbolic Reasoning Framework to integrate spatial representations into structured symbolic reasoning chains.
Outcome: The proposed framework outperforms existing models in vision-intensive mathematical problems.
Parallel Communities Across the Surface Web and the Dark Web (2025.findings-emnlp)

Copied to clipboard

Challenge: Sense of Community is a social motivation that is reflected in the social behavior of humans.
Approach: They compile a large collection of parallel community datasets comprising over 7 million posts and comments from Reddit and 200,000 posts and comment from Dread, a dark web discussion forum, covering similar topics.
Outcome: The results show that users on Reddit exhibit a stronger sense of community membership despite the dark web’s restricted accessibility.
Tox-BART: Leveraging Toxicity Attributes for Explanation Generation of Implicit Hate Speech (2024.findings-acl)

Copied to clipboard

Challenge: Existing language models to generate implicit hate explanations are lacking in many fields.
Approach: They propose to use language models to generate explicit hate posts to make it clear . they find that simpler models incorporating external toxicity signals outperform KG-infused models .
Outcome: The proposed setup produces more precise explanations than zero-shot GPT-3.5, highlighting the intricate nature of the task.
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)

Copied to clipboard

Challenge: polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks .
Approach: They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events.
Outcome: The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context.
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation (2024.lrec-main)

Copied to clipboard

Challenge: a number of languages are used in online conversations, resulting in code-mixing . the problem is largely unexplored due to the lack of annotated data and noise .
Approach: They propose a robust perturbation-based joint-training model that learns to handle noise in code-mixed text by parameter sharing across clean and noisy words.
Outcome: The proposed model learns to handle noise in the real-world code-mixed text by parameter sharing across clean and noisy words.
Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks (2026.eacl-long)

Copied to clipboard

Challenge: Existing multilingual long-context benchmarks are myopic and inherently limited, as successful recall alone does not indicate a model’s capacity to reason over extended contexts.
Approach: They propose a new synthetic benchmark for multilingual long-context reasoning that includes bAbI-style tasks that test multi-hop inference, aggregation, and epistemic reasoning.
Outcome: The proposed benchmarks are based on a multilingual long-context model and span seven languages.
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) and AI assistants are experiencing exponential growth in usage among expert and amateur users.
Approach: They propose to assess the reliability of current Large Language Models as science communicators . they use a dataset comprising 742 Yes/No queries embedded in complex scientific concepts .
Outcome: The proposed model outperforms open-access models in scientific question-answering tasks . the model outpersforms GPT-4 Turbo models in many evaluation aspects .
Language Models can Exploit Cross-Task In-context Learning for Data-Scarce Novel Tasks (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have transformed NLP with their remarkable In-context Learning capabilities.
Approach: They propose to use large language models to generalize from labeled examples of predefined tasks to novel tasks . they use biological neurons and the Transformer architecture to study the potential for information sharing across tasks.
Outcome: The proposed model can generalize from labeled examples of predefined tasks to novel tasks despite no examples from the target task in the context.
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking (2023.findings-emnlp)

Copied to clipboard

Challenge: Social media posts are noisy and pervasive, resulting in difficult to identify precise and prominent claims that require verification.
Approach: They propose a task called Claim Normalization that decomposes complex and noisy social media posts into more straightforward and understandable forms, termed normalized claims.
Outcome: The proposed model outperforms baselines across evaluation measures and errors.
Accuracy is not enough: Evaluating Personalization in Summarizers (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing accuracy measures cannot evaluate the degree of personalization of summarization models.
Approach: They propose to use a PENS dataset to analyze the degree of personalization of ten different summarization models.
Outcome: The proposed measure can evaluate the degree of personalization of summarization models using the PENS dataset.
Hate Personified: Investigating the role of LLMs in content moderation (2024.emnlp-main)

Copied to clipboard

Challenge: Our work provides preliminary guidelines and highlights the nuances of applying Large Language models in culturally sensitive cases.
Approach: They propose to use large language models to help with content moderation to assess how well the needs of diverse groups are reflected in annotated posts.
Outcome: The proposed model is able to leverage community-based flagging efforts and exposure to adversaries.
MOMENTA: A Multimodal Framework for Detecting Harmful Memes and Their Targets (2021.findings-emnlp)

Copied to clipboard

Challenge: a growing number of harmful memes are being used for trolling, cyberbullying and abuse . a new approach to detect harmful meme images and texts is emerging .
Approach: They propose a multimodal deep neural network that detects harmful memes . they extend the recently released HarMeme dataset with additional memes and a new topic .
Outcome: The proposed framework outperforms rival methods in detecting harmful memes and their target social entities.
Characterizing the Entities in Harmful Memes: Who is the Hero, the Villain, the Victim? (2023.eacl-main)

Copied to clipboard

Challenge: A common problem associated with meme comprehension lies in detecting the entities referenced and characterizing the role of each of these entities.
Approach: They propose to use a memes dataset on US Politics and Covid-19 memes to characterize the role of harmful entities in memes.
Outcome: The proposed model improves 4% over baseline and 1% over competing models.
Multilingual Needle in a Haystack: Investigating Long-Context Behavior of Multilingual Large Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Recent large language models demonstrate remarkable abilities in responding to queries in diverse languages, but their ability to handle long multilingual contexts is unexplored.
Approach: They propose a multilingual Needle-in-a-Haystack (MLNeedle) test to assess a model's ability to retrieve relevant information from a collection of multilingual distractor texts.
Outcome: The proposed model performance is the lowest when the needle is in a language outside the English language family and (ii) located in the middle of the input context.
Knowledge Planning in Large Language Models for Domain-Aligned Counseling Summarization (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit remarkable capabilities in various generative tasks, but their adaptation to domain-specific intricacies remains challenging.
Approach: They propose to use a planning engine to orchestrate structuring knowledge alignment to achieve high-order planning by encapsulating domain knowledge and leveraging sheaf convolution learning to enhance its understanding of the dialogue’s structural nuances.
Outcome: The proposed framework improves on existing LLMs and shows that it can generate better summaries with better quality and better execution.
On the Generalization vs Fidelity Paradox in Knowledge Distillation (2025.findings-acl)

Copied to clipboard

Challenge: Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance.
Approach: They propose to use knowledge distillation to compress large language models into smaller ones while preserving performance.
Outcome: The proposed technique improves the performance of smaller models by 10% while providing only marginal benefits for larger models.
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks emphasize final numerical answers while neglecting intermediate reasoning steps.
Approach: They propose a symbolic benchmark for verifiable Chain-of-Thought evaluation in finance . FINCHAIN spans 58 topics across 12 financial domains and three difficulty levels .
Outcome: The proposed benchmark aims to bridge symbolic reasoning and factual verification.
Multilingual LLMs are Better Cross-lingual In-context Learners with Alignment (2023.acl-long)

Copied to clipboard

Challenge: a handful of studies have explored ICL in a cross-lingual setting . emergence of large-scale, pretrained, Transformer-based language models has marked the commencement of an avant-garde era in NLP.
Approach: They propose a novel prompt construction strategy to bridge the gap between ICL and cross-lingual text classification.
Outcome: The proposed approach outperforms random prompt selection by a large margin across three tasks using 44 different cross-lingual pairs.
LM2: A Simple Society of Language Models Solves Complex Reasoning (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies show that providing guidance via decomposing the original question into multiple subproblems elicits more robustness in LLM reasoning.
Approach: They propose a language-based decomposition, solution and verification framework that modularizes the decomposer, solution, and verification into three different language models.
Outcome: The proposed model outperforms existing methods on in- and out-domain reasoning problems, outperforming the best baselines by 8.1% on MATH, 7.71% on JEEBench, and 9.7% on MedQA problems.
Manifold-Preserving Transformers are Effective for Short-Long Range Encoding (2023.findings-emnlp)

Copied to clipboard

Challenge: Multi-head self-attention-based Transformers have shown promise in different learning tasks . but encoders of Transformers and their variants fail to preserve layer-wise contextual information .
Approach: They propose an encoder model that guarantees a theoretical bound for layer-wise distance preservation between a pair of tokens.
Outcome: The proposed model preserves equivalence between tokens and performs better than Transformers.
MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing models lack cultural alignment across modalities and languages . a new framework to assess cultural awareness across linguistics and languages is needed .
Approach: They propose a framework that integrates tri-modally aligned cultural benchmarks and a five-dimensional evaluation protocol to assess cross-country awareness disparities.
Outcome: The proposed framework assesses cultural awareness disparities across modalities and languages . it is the first dataset aligned at the input level across text, image, and speech .
LESA: Linguistic Encapsulation and Semantic Amalgamation Based Generalised Claim Detection from Online Content (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on claim detection is built on the basis of a 'segregation' of claims across different domains.
Approach: They propose a generalized generalized model that captures syntactic features through part-of-speech and dependency embeddings, as well as contextual features through a fine-tuned language model.
Outcome: The proposed model outperforms baselines on six claim datasets by an average of 3 claim-F1 points and 2 claim-f1 points on the general-domain experiments.
Probing Critical Learning Dynamics of PLMs for Hate Speech Detection (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies on pretrained language models (PLMs) for hate speech detection have not investigated how their performance is affected by pretraining and finetuning.
Approach: They propose to compare pretrained language models, evaluate their seed robustness, finetuning settings, and the impact of pretraining data collection time.
Outcome: The proposed models show that they are more robust than other models and that they have a better chance of performing better than domain-specific models.
DISARM: Detecting the Victims Targeted by Harmful Memes (2022.findings-naacl)

Copied to clipboard

Challenge: DISARM is a framework that uses named-entity recognition and person identification to detect all entities a meme is referring to and then incorporates a novel contextualized deep neural network to classify whether the meme intends to harm these entities.
Approach: They propose a framework that uses named-entity recognition and person identification to detect all entities a meme is referring to and incorporates a novel contextualized deep neural network to classify whether the meme intends to harm them.
Outcome: The proposed framework outperforms 10 unimodal and multimodal systems and reduces error rate of harmful target identification by 9 % absolute over baseline systems.
Multilingual Language Models Encode Script Over Linguistic Structure (2026.acl-long)

Copied to clipboard

Challenge: a recent study suggests that multilingual language models organize representations around surface form, but the nature of this internal organization remains elusive.
Approach: They analyze language-associated units across different model families and scales . romanization induces near-disjoint representations that align with neither native-script inputs nor English .
Outcome: The results show that multilingual language models organize representations around surface form . romanization induces near-disjoint representations that align with neither native-script inputs nor English .
Domain-aware Self-supervised Pre-training for Label-Efficient Meme Analysis (2022.aacl-main)

Copied to clipboard

Challenge: Existing self-supervised learning strategies focus on uni-modal applications . a recent study shows that multimodality is a major challenge for multi-modal systems .
Approach: They propose two self-supervised pre-training methods that employ off-the-shelf multi-modal hate-speech data . they also incorporate multiple specialized pretext tasks to cater to complex multi-modity representation learning .
Outcome: The proposed methods outperform the baseline self-supervised learning strategies on the Memotion challenge and the HarMeme task.
Fingerprinting Fine-tuned Language Models in the Wild (2021.findings-acl)

Copied to clipboard

Challenge: Existing fingerprinting methods to fingerprint language models are limited to attributing organic text . however, fine-tuned LMs can generate long, coherent, and grammatically valid synthetic text.
Approach: They conduct extensive experiments to demonstrate the limitations of existing fingerprinting approaches.
Outcome: The proposed fingerprinting methods are limited to attributing synthetic text generated by 10 pre-trained LMs.
The Psychology of Falsehood: A Human-Centric Survey of Misinformation Detection (2025.emnlp-main)

Copied to clipboard

Challenge: a survey examines the interplay between factual accuracy and cognitive biases . misinformation is more than just the existence of incorrect information, it also entails complex relationships between the information and the entities that consume it.
Approach: They examine the interplay between traditional fact-checking and psychological concepts such as cognitive biases, social dynamics, and emotional responses.
Outcome: The findings highlight limitations of current methods and identify opportunities for improvement . they also outline future research directions to create more robust frameworks .
When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Indirect speech achieves a constellation of discourse goals in human communication, but it is challenging for AI agents to comprehend such idiosyncrasies.
Approach: They propose a task to generate natural language explanations of satirical conversations using a multimodal and code-mixed dataset to capture multimodality.
Outcome: The proposed task generates natural language explanations of satirical conversations in a multimodal and code-mixed setting and surpasses baselines on almost all metrics.
CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Current evaluation methods do not capture complex attributes of counterspeech quality, such as contextual relevance, aggressiveness, or argumentative coherence.
Approach: They propose to use a dataset and framework to evaluate counterspeech quality across four dimensions: contextual relevance, aggressiveness, argument-coherence, and suitability.
Outcome: The proposed method outperforms ROUGE, METEOR, and BertScore in correlating with human judgement, indicating a significant improvement in automated counterspeech evaluation.
SABER: Uncovering Vulnerabilities in Safety Alignment via Cross-Layer Residual Connection (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) with safe-alignment training are vulnerable to jailbreak attacks, causing malicious users to generate harmful outputs.
Approach: They propose a safe-alignment jailbreak method that bypasses the middle-to-late layers of large language models by a residual connection.
Outcome: The proposed method improves by 51% over the best performing baseline GCG on HarmBench test set.
Temporally Consistent Factuality Probing for Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are used as an alternative knowledge base for many tasks.
Approach: They propose a temporally consistent factuality probe task that extends the consistency probe in the temporal dimension.
Outcome: The proposed task extends the definitions of existing metrics to represent consistent factuality across temporal dimension.
PerDucer: Keyphrase-Driven Personalization Inducer for Summarization from User Histories (2026.findings-acl)

Copied to clipboard

Challenge: Prior work reported that prepending long interaction histories to LLMs leads to unstable personalization, especially for multi-aspect documents.
Approach: They propose a personalization inducer for frozen language models that maps latent preference signals to a small set of personalized keyphrases for the query document.
Outcome: The proposed model outperforms the strongest history-prompting LLMs and SLMs in the PENS and OpenAI-Reddit benchmarks.
Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech Generation (2023.acl-long)

Copied to clipboard

Challenge: a counterspeech with a certain intent may not be sufficient in every situation due to complex nature of hate speech . a novel framework for intent-conditioned counterseech generation is proposed to address the pervasive issue of hateful speech on the internet.
Approach: They propose a framework for intent-conditioned counterspeech generation that leverages intent-specific representations and a fusion module to incorporate intent-related information into the model.
Outcome: The proposed framework outperforms baselines by 10% across evaluation metrics.
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing efforts to ensure temporal consistency in large language models are lacking in time-sensitive fields . temporal reasoning is essential for time- sensitive fields such as finance and healthcare . a new benchmark aims to improve temporal referent consistency of LLMs .
Approach: They propose a temporal referential consistency benchmark with a resource TEMP-ReCon to assess LLMs across temporal references.
Outcome: The proposed model improves LLMs' temporal consistency by comparing them to baseline models.
Promoting Topic Coherence and Inter-Document Consorts in Multi-Document Summarization via Simplicial Complex and Sheaf Graph (2023.emnlp-main)

Copied to clipboard

Challenge: Existing systems that generate summaries from multiple sources often lack accuracy and accuracy due to the length of tokens used in encoding.
Approach: They propose a novel encoder-decoder model that uses pre-trained BART to analyze linguistic nuances, simplicial complex layer to apprehend inherent properties that transcend pairwise associations and sheaf graph attention to effectively capture heterophilic properties.
Outcome: The proposed model achieves consistent performance improvement across all evaluation metrics (syntactical, semantical and faithfulness).
Can Unsupervised Knowledge Transfer from Social Discussions Help Argument Mining? (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for argument mining are limited by the scarcity of manually annotated data and the highly domain-dependent nature of argumentation.
Approach: They propose a novel transfer learning strategy to fine tune pretrained Transformer-based Language Models on a selectively masked language modeling task and a new prompt-based strategy for inter-component relation prediction.
Outcome: The proposed method outperforms existing models on both within- and out-of-domain datasets while leveraging on the discourse context.
From Multilingual Complexity to Emotional Clarity: Leveraging Commonsense to Unveil Emotions in Code-Mixed Dialogues (2023.emnlp-main)

Copied to clipboard

Challenge: Understanding emotions during conversation is a fundamental aspect of human communication.
Approach: They propose an approach that integrates commonsense information with dialogue context to facilitate a deeper understanding of emotions.
Outcome: The proposed approach improves ERC for code-mixed conversations by integrating commonsense with dialogue context.
Do You Know About My Nation? Investigating Multilingual Language Models’ Cultural Literacy Through Factual Knowledge (2025.emnlp-main)

Copied to clipboard

Challenge: Multilingual question-answering benchmarks do not factor in regional diversity in the information they capture and tend to be Western-centric.
Approach: They propose to benchmark eight standard multilingual LLMs on XNationQA and evaluate them using two novel transference metrics.
Outcome: The proposed model shows greater knowledge of cultural information in English than in the dominant language of the respective culture.
Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning (2025.acl-long)

Copied to clipboard

Challenge: Existing methods to generate counterspeech based on intents are limited to single attributed . however, a holistic approach that considers multiple attributes simultaneously yields more nuanced and effective responses.
Approach: They propose a framework that leverages hierarchical prefix learning with preference optimization to generate more constructive counterspeech.
Outcome: The proposed framework improves intent conformity and emotion labels in 13,973 counterspeech instances.
Rank, Chunk and Expand: Lineage-Oriented Reasoning for Taxonomy Expansion (2025.findings-acl)

Copied to clipboard

Challenge: Existing taxonomy expansion methods struggle with representation limits and generalization, while generative methods process all candidates at once, introducing noise and exceeding context limits.
Approach: They propose a plug-and-play framework that combines discriminative ranking and generative reasoning for efficient taxonomy expansion.
Outcome: Experiments show that LORex improves accuracy by 12% and similarity by 5% over state-of-the-art methods.
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF (2024.naacl-long)

Copied to clipboard

Challenge: Existing systems that target hate speech with intent-conditioned counterspeech generate better results with longer contexts.
Approach: They propose a framework that enables counterspeech generation by modeling the pragmatic implications underlying social biases in hateful statements.
Outcome: The proposed framework outperforms existing benchmarks in intent-conditioned counterspeech generation.
Can LLMs Automate Fact-Checking Article Writing? (2026.tacl-1)

Copied to clipboard

Challenge: Existing tools for automatic fact-checking produce little or no justification for their assessments . 80% of american adults on major social media platforms regularly encounter news-related content .
Approach: They propose to extend automatic fact-checking pipeline with automatic generation of full fact- checking articles.
Outcome: The proposed framework outperforms existing frameworks but lags behind expert-written articles.
Empowering the Fact-checkers! Automatic Identification of Claim Spans on Twitter (2022.emnlp-main)

Copied to clipboard

Challenge: Current vogue is to employ manual fact-checkers to efficiently classify and verify such data to combat this avalanche of misinformation and fake news.
Approach: They propose a large-scale Twitter corpus with token-level claim spans on more than 7.5k tweets and a model that automatically detects and extracts the snippets of misinformation.
Outcome: The proposed model outperforms baseline systems on several evaluation metrics, improving by 1.5 points.
Nurse is Closer to Woman than Surgeon? Mitigating Gender-Biased Proximities in Word Embeddings (2020.tacl-1)

Copied to clipboard

Challenge: Existing methods for debiasing word embeddings lack gender-based debiases . Existing approaches only reduce gender-related proximity biases by at least 42.02% .
Approach: They propose a gender debiasing methodology that eliminates bias in word vectors and alters spatial distribution of neighboring vectors, achieving a bias-free setting while maintaining minimal semantic offset.
Outcome: The proposed method outperforms the state-of-the-art in reducing proximity bias by at least 42.02% and reduces direct bias, adding minimal semantic disturbance, and achieves the best performance in a downstream application task.
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception (2026.findings-acl)

Copied to clipboard

Challenge: Small Vision-Language Models (SVLMs) suffer from visual brittleness and poor tool orchestration.
Approach: They propose a supervision-free framework that bootstraps agentic capabilities via Coldstart Reinforcement Learning for SVLMs.
Outcome: The proposed framework improves task accuracy and tool efficiency by 5% and 9%.
MEMEX: Detecting Explanatory Evidence for Memes via Knowledge-Enriched Contextualization (2023.acl-long)

Copied to clipboard

Challenge: Besides digital archiving of memes and their metadata, there is no efficient way to deduce a meme’s context dynamically.
Approach: They propose a task to mine the context that succinctly explains the background of a meme and a related document to capture cross-modal semantic dependencies between the meme and the context.
Outcome: The proposed dataset outperforms existing systems and shows that it can capture cross-modal semantic dependencies between the meme and the context.
Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Recent attempts at prompt decomposition toward solving complex, multi-step reasoning problems depend on the ability of the LLM to simultaneously decompose and solve the problem.
Approach: They propose a decomposition generator that decomposes complex problems into subproblems that require fewer reasoning steps.
Outcome: The proposed method can produce competitive or even better performance compared to its larger successor, GPT-4.
Are Large Language Models In-Context Personalized Summarizers? Get an iCOPERNICUS Test Done! (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have succeeded in summarizing information in contexts but saliency is subject to user preferences.
Approach: They propose a framework that measures saliency using user reading histories and contrast in user profiles.
Outcome: The proposed framework evaluates state-of-the-art LLMs on their ICL performance and shows that they lack true ICPL.
Corpora Evaluation and System Bias Detection in Multi-document Summarization (2020.findings-emnlp)

Copied to clipboard

Challenge: Multi-document summarization (MDS) is a task of combining multiple documents into a concise text paragraph.
Approach: They propose to use a multi-document summarization task to reflect key points from any set of documents into a concise text paragraph.
Outcome: The proposed system performs better on a set of selected datasets than on the other ones.
Adding SPICE to Life: Speaker Profiling in Multiparty Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Prior studies assumed the speaker’s persona’s immediate availability, a premise not universally applicable.
Approach: They propose to synthesize persona attributes for each dialogue participant by combining three core tasks: persona discovery, persona-type identification, and persona value extraction.
Outcome: The proposed task synthesizes persona attributes for each dialogue participant . the resulting model is compared against a baseline model and the proposed model is robust.
Data-scarce Behavior Editing of Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Prior studies show that noisy neural circuitries coexist with generalizable abilities within LLMs.
Approach: a new method is proposed to improve the generalizability of large-scale web-based text models . a TaRot method is based on learnable rotation matrices optimized for Bayesian optimization .
Outcome: a new method for task adaptation improves on multiple classification and generation tasks . it improves upon zero- and few-shot performance, with average improvements of 14% and 15% .
Detecting Harmful Memes and Their Targets (2021.findings-acl)

Copied to clipboard

Challenge: a growing body of research on meme analysis has focused on detecting harmful memes and their social entities . a meme is a form of content that is often harmless and designed to look funny . but its multimodal nature and camouflaged semantics make its analysis challenging .
Approach: They propose to use multimodal models to detect harmful memes and identify social entities that harmful meme targets.
Outcome: The proposed model can detect harmful memes and the social entities they target . the proposed model lacks the appropriate contexts and is poorly validated .
MemeMQA: Multimodal Question Answering for Memes via Rationale-Based Inferencing (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies have focused on harms of memes in closed environments, such as hate speech and cyber-bullying.
Approach: They propose a multimodal question-answering framework that solicits accurate responses to structured questions while providing coherent explanations.
Outcome: The proposed framework outperforms existing frameworks in predicting answer prediction accuracy and text generation lead over a baseline.
Recent Advances in Online Hate Speech Moderation: Multimodality and the Role of Large Models (2024.findings-emnlp)

Copied to clipboard

Challenge: HS is any communication demeaning a person or a group based on social or ethnic characteristics that undermines social harmony and individual safety . the recent Israel-Hamas conflict has escalated both anti-Muslim and anti-Semitic sentiments worldwide .
Approach: They examine the role of large language models and large multimodal models in HS moderation . they examine how text, images, and audio interact to spread hate speech .
Outcome: The findings highlight the need for solutions in low-resource settings and highlight the gaps in existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations