Papers by Soroush Vosoughi
Copied to clipboard
| Challenge: | Existing metrics that compare the candidate with the human reference do not consider the context, resulting in poor correlation with human judgements. |
| Approach: | They propose a language model-aware metric that augments the human reference while considering the context to provide evaluation scores that correlate highly with human judgements. |
| Outcome: | The proposed metric achieves higher correlation with human reference judgements and differentiates well-formed candidates from adversarial samples to a larger degree. |
Copied to clipboard
| Challenge: | Pretrained language models (PLMs) propagate social stigmas and stereotypes, a critical concern given their widespread use. |
| Approach: | They adapt two intrinsic bias benchmarks to quantify racial and LGBTQ+ biases in prevalent PLMs and empirically evaluate the effectiveness of various debiasing methods in mitigating these biase. |
| Outcome: | The proposed methods reduce biases without compromising performance in downstream tasks. |
Copied to clipboard
| Challenge: | Current video captioning methods often incorporate intricate designs tailored to video inputs. |
| Approach: | They adapt an image-based captioning model to address dynamic video sequences without modifications. |
| Outcome: | The proposed model outperforms specialised captioning systems on major benchmarks. |
Copied to clipboard
| Challenge: | Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans . |
| Approach: | They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs. |
| Outcome: | The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs. |
Copied to clipboard
| Challenge: | Nüshu is an endangered language from Jiangyong County, Hunan, China, and the world’s only known writing system created and used exclusively by women. |
| Approach: | They propose to use NüshuStrokes to record all 397 Unicode Nü Shu characters in sequential handwriting by an expert calligrapher. |
| Outcome: | Evaluating five state-of-the-art Chinese Optical Character Recognition systems on NüshuVision lowers CER to 0.67, a modest but meaningful improvement over previous datasets. |
Copied to clipboard
| Challenge: | a multi-task view of data augmentation allows for a more robust performance than traditional augmentation. |
| Approach: | They propose a multi-task view of data augmentation where original and augmented samples are weighted substantively during training. |
| Outcome: | The proposed model improves on three benchmark text classification datasets. |
Copied to clipboard
| Challenge: | Existing methods for interpreting LLMs are post hoc and focus on low-level features and lack of explainability at higher-level text units. |
| Approach: | They propose a prototypical network-based white-box framework that allows LLMs to learn immediately interpretable embeddings during the fine-tuning stage while maintaining competitive performance. |
| Outcome: | The proposed framework can learn interpretable embeddings during the fine-tuning stage while maintaining competitive performance. |
Copied to clipboard
| Challenge: | Using large language models, we examine the limitations of their cognitive capabilities and their working memory. |
| Approach: | They examine the limitations of large language models from a scaling perspective . they also assess various prompting strategies, revealing their diverse impacts on LLM performance. |
| Outcome: | The proposed models perform poorly on n-back tasks and on prompting strategies. |
Copied to clipboard
| Challenge: | linguistic complexity loss in text-based therapy can be used to identify patterns of mental health . authors: clients who reported more anxiety used less lexically diverse language . |
| Approach: | They analyze linguistic complexity loss in online therapy conversations as it relates to mental health . they find that clients used less lexically diverse language when they were more anxious . |
| Outcome: | The proposed analysis shows that therapists use more complex language when clients are anxious . the authors show that analyzing linguistic complexity can identify meaningful patterns in mental health . |
Copied to clipboard
| Challenge: | Multimodal foundation models have demonstrated significant success in tasks such as visual captioning, question answering, and image-text retrieval. |
| Approach: | They propose a specialized cognitive module, temporal working memory, which selectively retains task-relevant information across temporal dimensions. |
| Outcome: | The module retains task-relevant information across temporal dimensions, ensuring that critical details are preserved throughout the processing of video and audio content. |
Copied to clipboard
| Challenge: | Propaganda is a widely used approach for shaping public opinion and disseminating misinformation in news media. |
| Approach: | They propose a fine-grained propaganda dataset for Kinyarwanda . they find that current LLMs are not reliable annotators in low resource settings . |
| Outcome: | The proposed dataset shows that current LLMs perform poorly in low resource settings . the dataset shows they perform poorly on discourse-level techniques . |
Copied to clipboard
| Challenge: | Current approaches for examining stereotypes in PLMs require intricate human knowledge about these stereotypes and entail careful manual curation of examples. |
| Approach: | They propose a framework for examining stereotype-encoding behavior of PLMs using model probing and textual analyses. |
| Outcome: | The proposed approach can debiase PLMs without compromising their language modeling capabilities or performance. |
Copied to clipboard
| Challenge: | Modern embodied AI uses multimodal large language models as policy models, predicting actions from final-layer hidden states. |
| Approach: | They propose a hierarchical action probing method that aggregates representations from all layers, mirroring the brain's multi-level organization. |
| Outcome: | Experiments show that hierarchical probing improves on last-layer embodied models and achieves a 46.6% success rate and a 62.5% gain in spatial reasoning tasks. |
Copied to clipboard
| Challenge: | Rhetorical strategies are important to persuasive communication, but their analysis relies on human annotation, which is costly, inconsistent and difficult to scale. |
| Approach: | They propose a framework that leverages large language models to generate and label debate data . they fine-tune transformer-based classifiers on this dataset and validate it against human data a . |
| Outcome: | The proposed model achieves high performance and strong generalization across topical domains. |
Copied to clipboard
| Challenge: | Existing approaches to persuasive dialogue generation suffer from stance oscillation and low informativeness. |
| Approach: | They propose reinforced instructional prompting, a method that ensures speaker characteristics consistently guide all stages of dialogue generation. |
| Outcome: | The proposed method ensures speaker characteristics guide all stages of dialogue generation and aligns language use with speakers’ native languages to better capture cultural nuances. |
Copied to clipboard
| Challenge: | Despite their cultural and historical significance, Native American languages remain unsupported by major commercial language identification systems. |
| Approach: | They propose to curate linguistic resources across all Native American languages for robust training and tailor data augmentation to generate synthetic yet linguistically coherent training samples. |
| Outcome: | The proposed system would be generalizable across all Native American languages . it would also generate coherent training samples for low-resource languages based on Plains Apache . |
Copied to clipboard
| Challenge: | Existing summarization metrics favor shorter or longer summaries, but evaluations of these metrics are flawed. |
| Approach: | They propose a Bayesian normalization technique that effectively diminishes this bias. |
| Outcome: | The proposed method significantly improves the concordance between human annotators and most metrics in terms of summary coherence. |
Copied to clipboard
| Challenge: | Preprocessing-based methods for stereotype mitigation are widely used in NLP . preprocessing methods cause unintended shifts in attention flow, authors say . |
| Approach: | They propose to use preprocessing-based methods to reduce stereotypes for targeted groups . they find that stereotyping or counter-stereotyping can increase for other demographics . |
| Outcome: | The proposed methods often induce unintended shifts across demographics, the authors show . they show that such side effects are not accompanied by large changes in attention flow . |
Copied to clipboard
| Challenge: | Music audio-visual question answering presents unique challenges with dense audio-visual content, intricate temporal dynamics, and the need for domain-specific knowledge. |
| Approach: | They analyze Music AVQA datasets and analyze their results to identify key design patterns . they propose concrete future directions for incorporating musical priors . |
| Outcome: | The proposed architectures are critical for success in Music AVQA, the authors argue . they suggest concrete future directions for incorporating musical priors . |
Copied to clipboard
| Challenge: | Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures. |
| Approach: | They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step. |
| Outcome: | The proposed method outperforms slow-thinking methods while producing shorter responses. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) rely heavily on large-scale reasoning data, but as data becomes scarce, model self-improvement offers a promising alternative. |
| Approach: | They propose to merge the weights of original and self-improved LLMs to mitigate model collapse and improve generalized reasoning capability. |
| Outcome: | The proposed model merge mitigates model collapse and improves generalized reasoning capability. |
Copied to clipboard
| Challenge: | Existing preference-alignment approaches rely on binary pairwise comparisons, overlooking preference intensity and temporal context. |
| Approach: | They propose a unified preference optimization framework that maps both explicit and implicit feedback into a common preference signal and constructs adaptive reward margins that jointly account for preference intensity and interaction recency. |
| Outcome: | The proposed framework outperforms state-of-the-art recommendations while maintaining behavioral patterns aligned with human decision-making. |
Copied to clipboard
| Challenge: | Psychological inoculations have shown efficacy in curbing its spread and mitigating its adverse effects at early stages, but their design and optimization typically requires substantial human and financial resources due to the need for repeated experimental trials. |
| Approach: | They propose to use large language models to simulate participant responses in misinformation studies to mitigate caricatures and stereotypes in the simulations. |
| Outcome: | The proposed method mitigates caricatures and stereotypes in LLM simulations and enhances response diversity. |
Copied to clipboard
| Challenge: | Nüshu is a rare syllabic script used by Yao women in china for self-expression . a lack of data makes the reconstruction labor-intensive and costly . |
| Approach: | They propose an AI-driven framework to train large corpora on endangered languages . Nüshu is a rare syllabic script used by Yao women in china for self-expression . |
| Outcome: | NüshuRescue automates evaluation and expands target corpora to accelerate linguistic revitalization. |
Copied to clipboard
| Challenge: | Endangered languages are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and revitalization. |
| Approach: | They propose a random forest classifier trained on Navajo and 20 erroneously suggested languages by Google's Language Identification tool. |
| Outcome: | The proposed classifier achieves near-perfect accuracy across other Athabaskan languages suggesting its potential for broader application. |
Copied to clipboard
| Challenge: | enhancing the quality of online public discourse requires promoting foundational human virtues, such as “intellectual humility” (IH) . discourse on social media rewards forgetting our virtuous selves, embedding users within echo chambers and causing negative affect towards those who hold different beliefs. |
| Approach: | They propose to use a codebook to measure "intellectual humility" they manually validated the codebook and used it to develop LLM-based models . |
| Outcome: | The proposed model achieves a Macro-F1 score of 0.64 across labels and 0.70 when predicting IH/IA/Neutral at the coarse level. |
Copied to clipboard
| Challenge: | Prior research has found that only a few attention heads are important in each mono-lingual NLP task and pruning the remaining heads leads to comparable or improved performance of the model. |
| Approach: | They examine the relative importance of attention heads in Transformer-based models to aid their interpretability in cross-lingual and multi-lingual tasks. |
| Outcome: | The proposed model performs better with the remaining heads pruned than with the other models, the authors show . |
Copied to clipboard
| Challenge: | EVLGen is a framework for visual-language pre-training with high computational demands. |
| Approach: | They propose a streamlined framework for the pre-training of visually conditioned language generation models with high computational demands. |
| Outcome: | The proposed framework accelerates training of vision-language models by a factor of 5 without compromising performance. |
Copied to clipboard
| Challenge: | Existing research on cultural background modeling is coarse-grained and does not examine cultural differences among speakers of the same language. |
| Approach: | They use a news-based cultural background prediction dataset to annotate, validate and benchmark NLP models with cultural background features. |
| Outcome: | The proposed model improves on nine syntactic, semantic, and psycholinguistic tasks while introducing cultural background information does not improve the Go-Emotions task due to text domain conflicts. |
Copied to clipboard
| Challenge: | Emojis are able to express various linguistic components, such as emotions, sentiments, events, etc. emojis have the merit of preserving information more densely, compared to words, argues a new study. |
| Approach: | They propose to use passage-level and aspect-level emoji annotations to predict the proper emmojis associated with text. |
| Outcome: | The proposed method is heuristically generated and validated with a pre-trained BERT model. |
Copied to clipboard
| Challenge: | Extensive experiments on long-context multi-hop question answering benchmarks show TAG achieves state-of-the-art performance. |
| Approach: | They propose a framework that prestructures memory into diverse granularities and employs a reward-guided navigator to adaptively compose hybrid memory tailored to each query. |
| Outcome: | Experiments on long-context multi-hop question answering show that the framework achieves state-of-the-art performance. |
Copied to clipboard
| Challenge: | Recent studies combine LoRA with Mixture-of-Experts (MoE) to improve performance in Large Language Models. |
| Approach: | They propose a method to combine LoRA and Mixture-of-Experts (MoE) to improve performance in Large Language Models. |
| Outcome: | The proposed method reduces redundancy in LoRA experts within the MoE architecture, and improves training quality across layers. |
Copied to clipboard
| Challenge: | Current research focuses on predefined aspects within structured texts, neglecting complexities of dynamic and disordered environments. |
| Approach: | They propose a benchmark for dynamic aspect-based summarization tailored to unstructured text. |
| Outcome: | The proposed benchmark addresses the complexities of dynamic and disordered environments in unstructured text. |
Copied to clipboard
| Challenge: | a recent study has shown that GPT-3 fine-tuning models with limited examples is effective . a contrastive learning framework clusters inputs from the same class under different augmented “views” and repels those from different classes. |
| Approach: | They propose a supervised contrastive framework that clusters inputs from the same class under different augmented "views" they combine a contrastive loss with the standard masked language modeling loss in prompt-based few-shot learners . |
| Outcome: | The proposed framework improves on the state-of-the-art methods in a diverse set of 15 language tasks. |
Copied to clipboard
| Challenge: | LLM-as-Judge frameworks provide scalable alternative to human evaluation . but the question of how intrinsic biases manifest in these settings remains unexplored . |
| Approach: | They conduct systematic analysis of four bias types in multi-agent LLM-as-Judge frameworks . they find debate framework amplifies biases sharply after initial debate . |
| Outcome: | The proposed frameworks amplify biases after debate and show they are stronger in meta-judge scenarios. |
Copied to clipboard
| Challenge: | Recent studies on propaganda detection involve document and fragment-level analyses of news articles. |
| Approach: | They propose a neural approach to detect and categorize propaganda tweets across fine-grained categories . they use a dataset containing tweets weakly annotated with different propaganda techniques . |
| Outcome: | The proposed method outperforms benchmark methods and transfers knowledge to low-resource news domains. |
Copied to clipboard
| Challenge: | Existing methods for generating context-aware language that embodies diverse emotions are dull or generic due to limited training data for diverse emotions. |
| Approach: | They propose a modulated layer normalization technique that generates emotional responses using large pre-trained models. |
| Outcome: | The proposed method outperforms baseline methods on the MojiTalk dataset while maintaining diversity, fluency, and coherence. |
Copied to clipboard
| Challenge: | Recent large language models have demonstrated impressive reasoning abilities, but their extension to the audio modality remains underexplored. |
| Approach: | They propose a rule-based reinforcement learning algorithm to equip LALMs with robust reasoning capabilities. |
| Outcome: | The proposed algorithm improves on the SoundMind benchmark. |
Copied to clipboard
| Challenge: | Existing methods for learning human values do not consider contextual and abstract nature of human values. |
| Approach: | They propose a reinforcement learning based method that embeds human values judgements into each step of language generation. |
| Outcome: | The proposed method improves on human values judgements and shows higher alignment performance. |
Copied to clipboard
| Challenge: | Existing methods for data augmentation produce low readability or semantic consistency. |
| Approach: | They propose a framework which augments data through reinforcement learning guided conditional generation. |
| Outcome: | The proposed framework improves F1 performance on three different classification tasks by 8.7% on average when given only 10% of the whole data for training. |
Copied to clipboard
| Challenge: | Existing approaches for detecting and mitigating embedded stereotypes rely on carefully annotated datasets like StereoSet and CrowS-Pairs, which are only in English and reflect stereotypes from a few English-speaking countries. Existing datasets, especially translation-based ones, often overlook such cultural distinctions. |
| Approach: | They propose a cost-efficient human-LLM collaborative annotation framework to construct a Spanish-language stereotype dataset spanning multiple Spanish-speaking countries. |
| Outcome: | The proposed framework can identify nuanced, region-specific biases across Spanish-supporting LLMs and is adaptable to other languages and regions. |
Copied to clipboard
| Challenge: | a framework for measuring differential framing of issues is needed to address these issues . issue framers can be expressed explicitly with evaluative language or implicitly . quantitative methods have been used to measure issue framming . |
| Approach: | They propose a framework for modeling the differential framing of issues through masked token prediction using large-scale fine-tuned language models. |
| Outcome: | The proposed framework captures differential framing of issues with high reliability . it can be used to predict tone and word choices in written language . |
Copied to clipboard
| Challenge: | Existing studies on mental manipulation focus on context-free content and face challenges in identifying implicit toxicity. |
| Approach: | They propose a dataset that analyzes mental manipulation and its components . they propose to use 4,000 fictional dialogues to identify the techniques utilized for manipulation . |
| Outcome: | The proposed dataset enables a comprehensive analysis of mental manipulation . it shows that leading-edge models inadequately identify and categorize manipulative content . |
Copied to clipboard
| Challenge: | Indigenous languages are largely invisible in commercial language identification systems, a stark reality exemplified by Google Translate’s LangID tool, which excludes all 150 Indigenous languages of North America. |
| Approach: | They propose a framework that shows how large language models and specialized classifiers can effectively identify these languages with minimal data. |
| Outcome: | The proposed framework shows that large language models and specialized classifiers can effectively identify these languages with minimal data. |
Copied to clipboard
| Challenge: | A key problem in multi-task learning (MTL) research is how to select high-quality auxiliary tasks automatically. |
| Approach: | They propose an automatic auxiliary task selection method based on gradient calculation in Transformer-based models that improves MT-DNN performance. |
| Outcome: | The proposed method improves MT-DNN performance on 8 natural language understanding (GLUE) tasks, while costing less than AUTOSEM and comparable GPU consumption. |
Copied to clipboard
| Challenge: | Serial position effects (SPE) are well-documented cognitive biases in human behavior. |
| Approach: | They propose to use binary choices instead of multiple choices where feasible . they also suggest limiting prompt length and placing crucial information at the beginning of prompts . |
| Outcome: | The proposed framework shows that the effects are widespread across LLMs and the proposed mitigation methods are effective. |
Copied to clipboard
| Challenge: | Visual Question Answering (VQA) is increasingly used in diverse applications where models must provide accurate answers and explanations that humans can easily understand and verify. |
| Approach: | They propose a unified prototypical framework that learns question-aware prototypes that serve as reasoning anchors and applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant. |
| Outcome: | The proposed framework yields faithful, fine-grained explanations while maintaining competitive accuracy. |
Copied to clipboard
| Challenge: | Syntactic probing methods are biased by the PLMs’ memorization of common word co-occurrences, even if they do not form syntactical relations. |
| Approach: | They propose to use random word substitution and random label matching to reduce these biases and improve the robustness of syntactic probing methods. |
| Outcome: | The proposed tasks improve probing results and consistency between probing methods and make them more generalizable to unseen text domains. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown proficiency in generating persuasive dialogue, yet concerns about the fluency and sophistication of their outputs persist. |
| Approach: | They propose a multi-LLM communication framework that facilitates the efficient production of high-quality, diverse linguistic content with minimal human oversight. |
| Outcome: | The proposed framework excels in naturalness, linguistic diversity, and the strategic use of persuasion, even in complex scenarios involving social taboos. |
Copied to clipboard
| Challenge: | Existing methods for fine-tuning pre-trained language models can cause severe over-fitting. |
| Approach: | They propose an Embedding Hallucination method which generates auxiliary embedding-label pairs to expand the fine-tuning dataset. |
| Outcome: | The proposed method outperforms current fine-tuning methods in a wide range of language tasks. |
Copied to clipboard
| Challenge: | Lean proofs are built as sequences of tactic-induced state transitions, but learned models often represent proof steps through tactic strings or raw proof-state text. |
| Approach: | They train an encoder-only Transformer to learn contextualized representations of Lean proof steps from state changes. |
| Outcome: | The proposed model yields better held-out next-tactic retrieval than a surface-syntax control . the results provide a promising basis for future trajectory-aware theorem proving . |
Copied to clipboard
| Challenge: | Using Vision-Language Models (VLMs) for data visualizations requires significant time and expertise in both data management and graphic design. |
| Approach: | They propose a dataset comprising 2525 high-resolution data visualization figures with captions from AI conferences, extracted directly from source codes. |
| Outcome: | The proposed model outperforms open-source models in reproducing complex charts while using Chain-of-Thought prompting. |
Copied to clipboard
| Challenge: | a few-shot text classification task requires a large number of output classes, with few training examples per class. |
| Approach: | They propose a data augmentation technique suitable for training with limited data for few-shot, highly-multiclass text classification scenarios. |
| Outcome: | The proposed technique improves performance on four classification tasks by 3.0% on average. |
Copied to clipboard
| Challenge: | Curriculum learning (CL) orders data corpus by difficulty, but prior work employs disparate difficulty metrics and training setups. |
| Approach: | They propose a framework that decomposes curriculum difficulty into five dimensions: Problem Difficulty, Model Surprisal, Confidence Margin, Predictive Uncertainty and Decision Variability. |
| Outcome: | The proposed framework decomposes curriculum difficulty into five dimensions . the results show that no curriculum strategy dominates universally . |
Copied to clipboard
| Challenge: | Language models (LMs) require effective episodic grounding to perform well at physical planning tasks due to their limited ability to learn from and apply past experiences. |
| Approach: | They propose a weak-to-strong episodic learning framework that integrates episodic memory into hierarchical representations and pre-trained knowledge to unlock larger LMs' potential for grounding. |
| Outcome: | The proposed framework outperforms top proprietary LMs by 3.45% across diverse planning and question-answering tasks. |
Copied to clipboard
| Challenge: | Data augmentation is a field of research that has been underexplored due to the discrete nature of language data. |
| Approach: | They present a comprehensive survey of data augmentation for NLP by summarizing the literature in a structured manner. |
| Outcome: | The proposed methods are used for popular NLP applications and tasks and highlight current challenges and directions for future research. |
Copied to clipboard
| Challenge: | Existing methods for supervised fine-tuning (SFT) are suboptimal to preserve task-specific capabilities on RL-trained agentic models. |
| Approach: | They propose a distribution-aware merging framework specifically designed for RL-trained agentic models that disentangles shared and task-specific unique parameter updates while selectively preserving and rescaling unique ones. |
| Outcome: | Experiments across multiple agent domains and model architectures show that the proposed framework surpasses baselines and unlocks synergistic potential among agents. |
Copied to clipboard
| Challenge: | Existing methods for generating content specific summarization assume a fixed set of known aspects. |
| Approach: | They propose a dynamic aspect-based summarization framework that optimizes aspect number prediction and minimizes disparity between generated and reference summaries. |
| Outcome: | The proposed method outperforms baselines on three diverse datasets on different aspects of the input text. |