Papers by Soroush Vosoughi

58 papers
Language Model Augmented Relevance Score (2021.acl-long)

Copied to clipboard

Challenge: Existing metrics that compare the candidate with the human reference do not consider the context, resulting in poor correlation with human judgements.
Approach: They propose a language model-aware metric that augments the human reference while considering the context to provide evaluation scores that correlate highly with human judgements.
Outcome: The proposed metric achieves higher correlation with human reference judgements and differentiates well-formed candidates from adversarial samples to a larger degree.
Addressing Healthcare-related Racial and LGBTQ+ Biases in Pretrained Language Models (2024.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models (PLMs) propagate social stigmas and stereotypes, a critical concern given their widespread use.
Approach: They adapt two intrinsic bias benchmarks to quantify racial and LGBTQ+ biases in prevalent PLMs and empirically evaluate the effectiveness of various debiasing methods in mitigating these biase.
Outcome: The proposed methods reduce biases without compromising performance in downstream tasks.
Pretrained Image-Text Models are Secretly Video Captioners (2025.naacl-short)

Copied to clipboard

Challenge: Current video captioning methods often incorporate intricate designs tailored to video inputs.
Approach: They adapt an image-based captioning model to address dynamic video sequences without modifications.
Outcome: The proposed model outperforms specialised captioning systems on major benchmarks.
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
Recontextualizing Revitalization: A Mixed Media Approach to Reviving the Nüshu Language (2025.emnlp-main)

Copied to clipboard

Challenge: Nüshu is an endangered language from Jiangyong County, Hunan, China, and the world’s only known writing system created and used exclusively by women.
Approach: They propose to use NüshuStrokes to record all 397 Unicode Nü Shu characters in sequential handwriting by an expert calligrapher.
Outcome: Evaluating five state-of-the-art Chinese Optical Character Recognition systems on NüshuVision lowers CER to 0.67, a modest but meaningful improvement over previous datasets.
Text Augmentation in a Multi-Task View (2021.eacl-main)

Copied to clipboard

Challenge: a multi-task view of data augmentation allows for a more robust performance than traditional augmentation.
Approach: They propose a multi-task view of data augmentation where original and augmented samples are weighted substantively during training.
Outcome: The proposed model improves on three benchmark text classification datasets.
Proto-lm: A Prototypical Network-Based Framework for Built-in Interpretability in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for interpreting LLMs are post hoc and focus on low-level features and lack of explainability at higher-level text units.
Approach: They propose a prototypical network-based white-box framework that allows LLMs to learn immediately interpretable embeddings during the fine-tuning stage while maintaining competitive performance.
Outcome: The proposed framework can learn interpretable embeddings during the fine-tuning stage while maintaining competitive performance.
Working Memory Identifies Reasoning Limits in Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using large language models, we examine the limitations of their cognitive capabilities and their working memory.
Approach: They examine the limitations of large language models from a scaling perspective . they also assess various prompting strategies, revealing their diverse impacts on LLM performance.
Outcome: The proposed models perform poorly on n-back tasks and on prompting strategies.
Linguistic Complexity Loss in Text-Based Therapy (2021.naacl-main)

Copied to clipboard

Challenge: linguistic complexity loss in text-based therapy can be used to identify patterns of mental health . authors: clients who reported more anxiety used less lexically diverse language .
Approach: They analyze linguistic complexity loss in online therapy conversations as it relates to mental health . they find that clients used less lexically diverse language when they were more anxious .
Outcome: The proposed analysis shows that therapists use more complex language when clients are anxious . the authors show that analyzing linguistic complexity can identify meaningful patterns in mental health .
Temporal Working Memory: Query-Guided Segment Refinement for Enhanced Multimodal Understanding (2025.findings-naacl)

Copied to clipboard

Challenge: Multimodal foundation models have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval.
Approach: They propose a specialized cognitive module, temporal working memory, which selectively retains task-relevant information across temporal dimensions.
Outcome: The module retains task-relevant information across temporal dimensions, ensuring that critical details are preserved throughout the processing of video and audio content.
KinyaProp: Fine-Grained Propaganda Annotation in Kinyarwanda (2026.acl-long)

Copied to clipboard

Challenge: Propaganda is a widely used approach for shaping public opinion and disseminating misinformation in news media.
Approach: They propose a fine-grained propaganda dataset for Kinyarwanda . they find that current LLMs are not reliable annotators in low resource settings .
Outcome: The proposed dataset shows that current LLMs perform poorly in low resource settings . the dataset shows they perform poorly on discourse-level techniques .
Deciphering Stereotypes in Pre-Trained Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Current approaches for examining stereotypes in PLMs require intricate human knowledge about these stereotypes and entail careful manual curation of examples.
Approach: They propose a framework for examining stereotype-encoding behavior of PLMs using model probing and textual analyses.
Outcome: The proposed approach can debiase PLMs without compromising their language modeling capabilities or performance.
Knowing More, Acting Better: Hierarchical Representation for Embodied Decision-Making (2025.findings-emnlp)

Copied to clipboard

Challenge: Modern embodied AI uses multimodal large language models as policy models, predicting actions from final-layer hidden states.
Approach: They propose a hierarchical action probing method that aggregates representations from all layers, mirroring the brain's multi-level organization.
Outcome: Experiments show that hierarchical probing improves on last-layer embodied models and achieves a 46.6% success rate and a 62.5% gain in spatial reasoning tasks.
A Generalizable Rhetorical Strategy Annotation Model Using LLM-based Debate Simulation and Labelling (2025.findings-emnlp)

Copied to clipboard

Challenge: Rhetorical strategies are important to persuasive communication, but their analysis relies on human annotation, which is costly, inconsistent and difficult to scale.
Approach: They propose a framework that leverages large language models to generate and label debate data . they fine-tune transformer-based classifiers on this dataset and validate it against human data a .
Outcome: The proposed model achieves high performance and strong generalization across topical domains.
Enhancing LLM-Based Persuasion Simulations with Cultural and Speaker-Specific Information (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to persuasive dialogue generation suffer from stance oscillation and low informativeness.
Approach: They propose reinforced instructional prompting, a method that ensures speaker characteristics consistently guide all stages of dialogue generation.
Outcome: The proposed method ensures speaker characteristics guide all stages of dialogue generation and aligns language use with speakers’ native languages to better capture cultural nuances.
What is it? Towards a Generalizable Native American Language Identification System (2025.naacl-srw)

Copied to clipboard

Challenge: Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems.
Approach: They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples.
Outcome: The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache .
Length Does Matter: Summary Length can Bias Summarization Metrics (2023.emnlp-main)

Copied to clipboard

Challenge: Existing summarization metrics favor shorter or longer summaries, but evaluations of these metrics are flawed.
Approach: They propose a Bayesian normalization technique that effectively diminishes this bias.
Outcome: The proposed method significantly improves the concordance between human annotators and most metrics in terms of summary coherence.
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation (2026.findings-acl)

Copied to clipboard

Challenge: Preprocessing-based methods for stereotype mitigation are widely used in NLP . preprocessing methods cause unintended shifts in attention flow, authors say .
Approach: They propose to use preprocessing-based methods to reduce stereotypes for targeted groups . they find that stereotyping or counter-stereotyping can increase for other demographics .
Outcome: The proposed methods often induce unintended shifts across demographics, the authors show . they show that such side effects are not accompanied by large changes in attention flow .
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs (2026.findings-acl)

Copied to clipboard

Challenge: Music audio-visual question answering presents unique challenges with dense audio-visual content, intricate temporal dynamics, and the need for domain-specific knowledge.
Approach: They analyze Music AVQA datasets and analyze their results to identify key design patterns . they propose concrete future directions for incorporating musical priors .
Outcome: The proposed architectures are critical for success in Music AVQA, the authors argue . they suggest concrete future directions for incorporating musical priors .
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures.
Approach: They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step.
Outcome: The proposed method outperforms slow-thinking methods while producing shorter responses.
Superficial Self-Improved Reasoners Benefit from Model Merging (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as data becomes scarce, model self-improvement offers a promising alternative.
Approach: They propose to merge the weights of original and self-improved LLMs to mitigate model collapse and improve generalized reasoning capability.
Outcome: The proposed model merge mitigates model collapse and improves generalized reasoning capability.
What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal Context (2026.acl-long)

Copied to clipboard

Challenge: Existing preference-alignment approaches rely on binary pairwise comparisons, overlooking preference intensity and temporal context.
Approach: They propose a unified preference optimization framework that maps both explicit and implicit feedback into a common preference signal and constructs adaptive reward margins that jointly account for preference intensity and interaction recency.
Outcome: The proposed framework outperforms state-of-the-art recommendations while maintaining behavioral patterns aligned with human decision-making.
Simulated Misinformation Susceptibility (SMISTS): Enhancing Misinformation Research with Large Language Model Simulations (2024.findings-acl)

Copied to clipboard

Challenge: Psychological inoculations have shown efficacy in curbing its spread and mitigating its adverse effects at early stages, but their design and optimization typically requires substantial human and financial resources due to the need for repeated experimental trials.
Approach: They propose to use large language models to simulate participant responses in misinformation studies to mitigate caricatures and stereotypes in the simulations.
Outcome: The proposed method mitigates caricatures and stereotypes in LLM simulations and enhances response diversity.
NüshuRescue: Reviving the Endangered Nüshu Language with AI (2025.coling-main)

Copied to clipboard

Challenge: Nüshu is a rare syllabic script used by Yao women in china for self-expression . a lack of data makes the reconstruction labor-intensive and costly .
Approach: They propose an AI-driven framework to train large corpora on endangered languages . Nüshu is a rare syllabic script used by Yao women in china for self-expression .
Outcome: NüshuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization.
Is It Navajo? Accurate Language Detection for Endangered Athabaskan Languages (2025.naacl-short)

Copied to clipboard

Challenge: Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization.
Approach: They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool.
Outcome: The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application.
The Computational Anatomy of Humility: Modeling Intellectual Humility in Online Public Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: enhancing the quality of online public discourse requires promoting foundational human virtues, such as “intellectual humility” (IH) . discourse on social media rewards forgetting our virtuous selves, embedding users within echo chambers and causing negative affect towards those who hold different beliefs.
Approach: They propose to use a codebook to measure "intellectual humility" they manually validated the codebook and used it to develop LLM-based models .
Outcome: The proposed model achieves a Macro-F1 score of 0.64 across labels and 0.70 when predicting IH/IA/Neutral at the coarse level.
Contributions of Transformer Attention Heads in Multi- and Cross-lingual Tasks (2021.acl-long)

Copied to clipboard

Challenge: Prior research has found that only a few attention heads are important in each mono-lingual NLP task and pruning the remaining heads leads to comparable or improved performance of the model.
Approach: They examine the relative importance of attention heads in Transformer-based models to aid their interpretability in cross-lingual and multi-lingual tasks.
Outcome: The proposed model performs better with the remaining heads pruned than with the other models, the authors show .
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction (2024.acl-long)

Copied to clipboard

Challenge: EVLGen is a framework for visual-language pre-training with high computational demands.
Approach: They propose a streamlined framework for the pre-training of visually conditioned language generation models with high computational demands.
Outcome: The proposed framework accelerates training of vision-language models by a factor of 5 without compromising performance.
EnCBP: A New Benchmark Dataset for Finer-Grained Cultural Background Prediction in English (2022.findings-acl)

Copied to clipboard

Challenge: Existing research on cultural background modeling is coarse-grained and does not examine cultural differences among speakers of the same language.
Approach: They use a news-based cultural background prediction dataset to annotate, validate and benchmark NLP models with cultural background features.
Outcome: The proposed model improves on nine syntactic, semantic, and psycholinguistic tasks while introducing cultural background information does not improve the Go-Emotions task due to text domain conflicts.
Multi-resolution Annotations for Emoji Prediction (2020.emnlp-main)

Copied to clipboard

Challenge: Emojis are able to express various linguistic components, such as emotions, sentiments, events, etc. emojis have the merit of preserving information more densely, compared to words, argues a new study.
Approach: They propose to use passage-level and aspect-level emoji annotations to predict the proper emmojis associated with text.
Outcome: The proposed method is heuristically generated and validated with a pre-trained BERT model.
Tailoring Memory Granularity for Multi-Hop Reasoning over Long Contexts (2026.findings-eacl)

Copied to clipboard

Challenge: Extensive experiments on long-context multi-hop question answering benchmarks show TAG achieves state-of-the-art performance.
Approach: They propose a framework that prestructures memory into diverse granularities and employs a reward-guided navigator to adaptively compose hybrid memory tailored to each query.
Outcome: Experiments on long-context multi-hop question answering show that the framework achieves state-of-the-art performance.
AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies combine LoRA with Mixture-of-Experts (MoE) to improve performance in Large Language Models.
Approach: They propose a method to combine LoRA and Mixture-of-Experts (MoE) to improve performance in Large Language Models.
Outcome: The proposed method reduces redundancy in LoRA experts within the MoE architecture, and improves training quality across layers.
Disordered-DABS: A Benchmark for Dynamic Aspect-Based Summarization in Disordered Texts (2024.findings-emnlp)

Copied to clipboard

Challenge: Current research focuses on predefined aspects within structured texts, neglecting complexities of dynamic and disordered environments.
Approach: They propose a benchmark for dynamic aspect-based summarization tailored to unstructured text.
Outcome: The proposed benchmark addresses the complexities of dynamic and disordered environments in unstructured text.
Contrastive Learning for Prompt-based Few-shot Language Learners (2022.naacl-main)

Copied to clipboard

Challenge: a recent study has shown that GPT-3 fine-tuning models with limited examples is effective . a contrastive learning framework clusters inputs from the same class under different augmented “views” and repels those from different classes.
Approach: They propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" they combine a contrastive loss with the standard masked language modeling loss in prompt-based few-shot learners .
Outcome: The proposed framework improves on the state-of-the-art methods in a diverse set of 15 language tasks.
Judging with Many Minds: Do More Perspectives Mean Less Prejudice? On Bias Amplification and Resistance in Multi-Agent Based LLM-as-Judge (2025.findings-emnlp)

Copied to clipboard

Challenge: LLM-as-Judge frameworks provide scalable alternative to human evaluation . but the question of how intrinsic biases manifest in these settings remains unexplored .
Approach: They conduct systematic analysis of four bias types in multi-agent LLM-as-Judge frameworks . they find debate framework amplifies biases sharply after initial debate .
Outcome: The proposed frameworks amplify biases after debate and show they are stronger in meta-judge scenarios.
TWEETSPIN: Fine-grained Propaganda Detection in Social Media Using Multi-View Representations (2022.naacl-main)

Copied to clipboard

Challenge: Recent studies on propaganda detection involve document and fragment-level analyses of news articles.
Approach: They propose a neural approach to detect and categorize propaganda tweets across fine-grained categories . they use a dataset containing tweets weakly annotated with different propaganda techniques .
Outcome: The proposed method outperforms benchmark methods and transfers knowledge to low-resource news domains.
Modulating Language Models with Emotions (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating context-aware language that embodies diverse emotions are dull or generic due to limited training data for diverse emotions.
Approach: They propose a modulated layer normalization technique that generates emotional responses using large pre-trained models.
Outcome: The proposed method outperforms baseline methods on the MojiTalk dataset while maintaining diversity, fluency, and coherence.
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored.
Approach: They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities.
Outcome: The proposed algorithm improves on the SoundMind benchmark.
Aligning Generative Language Models with Human Values (2022.findings-naacl)

Copied to clipboard

Challenge: Existing methods for learning human values do not consider contextual and abstract nature of human values.
Approach: They propose a reinforcement learning based method that embeds human values judgements into each step of language generation.
Outcome: The proposed method improves on human values judgements and shows higher alignment performance.
Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing methods for data augmentation produce low readability or semantic consistency.
Approach: They propose a framework which augments data through reinforcement learning guided conditional generation.
Outcome: The proposed framework improves F1 performance on three different classification tasks by 8.7% on average when given only 10% of the whole data for training.
Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches for detecting and mitigating embedded stereotypes rely on carefully annotated datasets like StereoSet and CrowS-Pairs, which are only in English and reflect stereotypes from a few English-speaking countries. Existing datasets, especially translation-based ones, often overlook such cultural distinctions.
Approach: They propose a cost-efficient human-LLM collaborative annotation framework to construct a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries.
Outcome: The proposed framework can identify nuanced, region-specific biases across Spanish-supporting LLMs and is adaptable to other languages and regions.
Capturing Topic Framing via Masked Language Modeling (2022.findings-emnlp)

Copied to clipboard

Challenge: a framework for measuring differential framing of issues is needed to address these issues . issue framers can be expressed explicitly with evaluative language or implicitly . quantitative methods have been used to measure issue framming .
Approach: They propose a framework for modeling the differential framing of issues through masked token prediction using large-scale fine-tuned language models.
Outcome: The proposed framework captures differential framing of issues with high reliability . it can be used to predict tone and word choices in written language .
MentalManip: A Dataset For Fine-grained Analysis of Mental Manipulation in Conversations (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on mental manipulation focus on context-free content and face challenges in identifying implicit toxicity.
Approach: They propose a dataset that analyzes mental manipulation and its components . they propose to use 4,000 fictional dialogues to identify the techniques utilized for manipulation .
Outcome: The proposed dataset enables a comprehensive analysis of mental manipulation . it shows that leading-edge models inadequately identify and categorize manipulative content .
Visibility as Survival: Generalizing NLP for Native Alaskan Language Identification (2025.findings-acl)

Copied to clipboard

Challenge: Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America.
Approach: They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data.
Outcome: The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data.
GradTS: A Gradient-Based Automatic Auxiliary Task Selection Method Based on Transformer Networks (2021.emnlp-main)

Copied to clipboard

Challenge: A key problem in multi-task learning (MTL) research is how to select high-quality auxiliary tasks automatically.
Approach: They propose an automatic auxiliary task selection method based on gradient calculation in Transformer-based models that improves MT-DNN performance.
Outcome: The proposed method improves MT-DNN performance on 8 natural language understanding (GLUE) tasks, while costing less than AUTOSEM and comparable GPU consumption.
Serial Position Effects of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Serial position effects (SPE) are well-documented cognitive biases in human behavior.
Approach: They propose to use binary choices instead of multiple choices where feasible . they also suggest limiting prompt length and placing crucial information at the beginning of prompts .
Outcome: The proposed framework shows that the effects are widespread across LLMs and the proposed mitigation methods are effective.
ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering (2025.emnlp-main)

Copied to clipboard

Challenge: Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify.
Approach: They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant.
Outcome: The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy.
Improving Syntactic Probing Correctness and Robustness with Control Tasks (2023.acl-short)

Copied to clipboard

Challenge: Syntactic probing methods are biased by the PLMs’ memorization of common word co-occurrences, even if they do not form syntactical relations.
Approach: They propose to use random word substitution and random label matching to reduce these biases and improve the robustness of syntactic probing methods.
Outcome: The proposed tasks improve probing results and consistency between probing methods and make them more generalizable to unseen text domains.
Communication Makes Perfect: Persuasion Dataset Construction via Multi-LLM Communication (2025.naacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown proficiency in generating persuasive dialogue, yet concerns about the fluency and sophistication of their outputs persist.
Approach: They propose a multi-LLM communication framework that facilitates the efficient production of high-quality, diverse linguistic content with minimal human oversight.
Outcome: The proposed framework excels in naturalness, linguistic diversity, and the strategic use of persuasion, even in complex scenarios involving social taboos.
Embedding Hallucination for Few-shot Language Fine-tuning (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning pre-trained language models can cause severe over-fitting.
Approach: They propose an Embedding Hallucination method which generates auxiliary embedding-label pairs to expand the fine-tuning dataset.
Outcome: The proposed method outperforms current fine-tuning methods in a wide range of language tasks.
Representing Lean Proofs as Trajectories in Latent Space (2026.acl-srw)

Copied to clipboard

Challenge: Lean proofs are built as sequences of tactic-induced state transitions, but learned models often represent proof steps through tactic strings or raw proof-state text.
Approach: They train an encoder-only Transformer to learn contextualized representations of Lean proof steps from state changes.
Outcome: The proposed model yields better held-out next-tactic retrieval than a surface-syntax control . the results provide a promising basis for future trajectory-aware theorem proving .
Is GPT-4V (ision) All You Need for Automating Academic Data Visualization? Exploring Vision-Language Models’ Capability in Reproducing Academic Charts (2024.findings-emnlp)

Copied to clipboard

Challenge: Using Vision-Language Models (VLMs) for data visualizations requires significant time and expertise in both data management and graphic design.
Approach: They propose a dataset comprising 2525 high-resolution data visualization figures with captions from AI conferences, extracted directly from source codes.
Outcome: The proposed model outperforms open-source models in reproducing complex charts while using Chain-of-Thought prompting.
Few-Shot Text Classification with Triplet Networks, Data Augmentation, and Curriculum Learning (2021.naacl-main)

Copied to clipboard

Challenge: a few-shot text classification task requires a large number of output classes, with few training examples per class.
Approach: They propose a data augmentation technique suitable for training with limited data for few-shot, highly-multiclass text classification scenarios.
Outcome: The proposed technique improves performance on four classification tasks by 3.0% on average.
What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Curriculum learning (CL) orders data corpus by difficulty, but prior work employs disparate difficulty metrics and training setups.
Approach: They propose a framework that decomposes curriculum difficulty into five dimensions: Problem Difficulty, Model Surprisal, Confidence Margin, Predictive Uncertainty and Decision Variability.
Outcome: The proposed framework decomposes curriculum difficulty into five dimensions . the results show that no curriculum strategy dominates universally .
Growing Through Experience: Scaling Episodic Grounding in Language Models (2025.acl-long)

Copied to clipboard

Challenge: Language models (LMs) require effective episodic grounding to perform well at physical planning tasks due to their limited ability to learn from and apply past experiences.
Approach: They propose a weak-to-strong episodic learning framework that integrates episodic memory into hierarchical representations and pre-trained knowledge to unlock larger LMs' potential for grounding.
Outcome: The proposed framework outperforms top proprietary LMs by 3.45% across diverse planning and question-answering tasks.
A Survey of Data Augmentation Approaches for NLP (2021.findings-acl)

Copied to clipboard

Challenge: Data augmentation is a field of research that has been underexplored due to the discrete nature of language data.
Approach: They present a comprehensive survey of data augmentation for NLP by summarizing the literature in a structured manner.
Outcome: The proposed methods are used for popular NLP applications and tasks and highlight current challenges and directions for future research.
Behavior Knowledge Merge in Reinforced Agentic Models (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for supervised fine-tuning (SFT) are suboptimal to preserve task-specific capabilities on RL-trained agentic models.
Approach: They propose a distribution-aware merging framework specifically designed for RL-trained agentic models that disentangles shared and task-specific unique parameter updates while selectively preserving and rescaling unique ones.
Outcome: Experiments across multiple agent domains and model architectures show that the proposed framework surpasses baselines and unlocks synergistic potential among agents.
MODABS: Multi-Objective Learning for Dynamic Aspect-Based Summarization (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating content specific summarization assume a fixed set of known aspects.
Approach: They propose a dynamic aspect-based summarization framework that optimizes aspect number prediction and minimizes disparity between generated and reference summaries.
Outcome: The proposed method outperforms baselines on three diverse datasets on different aspects of the input text.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations