Papers by Sarvnaz Karimi

11 papers
Presentation Matters: How to Communicate Science in the NLP Venues and in the Wild? (2024.acl-tutorials)

Copied to clipboard

Challenge: a tutorial on communication skills is being proposed to help early career researchers . the tutorial would cover writing, oral presentation and social media presence .
Approach: a tutorial on communication skills is proposed to help early career researchers . the tutorial would cover writing, oral presentation and social media presence .
Outcome: a new tutorial will cover communication skills, including writing, oral presentation and social media presence . the tutorial will allow attendees to ask questions and clarify their research .
Question Answering in Climate Adaptation for Agriculture: Model Development and Evaluation with Expert Feedback (2025.findings-acl)

Copied to clipboard

Challenge: Existing domain-specific question answering systems have generative capabilities, but their ability to answer climate adaptation questions remains unclear.
Approach: They propose an iterative framework that enables LLMs to dynamically aggregate information from heterogeneous sources, such as climate literature and structured tabular climate data from climate model projections and historical observations.
Outcome: The proposed framework enables LLMs to dynamically aggregate information from heterogeneous sources, such as text from climate literature and structured tabular climate data from climate model projections and historical observations.
NNE: A Dataset for Nested Named Entity Recognition in English Newswire (P19-1)

Copied to clipboard

Challenge: Named entity recognition (NER) is widely used in downstream tasks but most tools focus on flat mention structure over coarse schemas.
Approach: They describe a fine-grained, nested named entity dataset over the Wall Street Journal portion of the Penn Treebank.
Outcome: The proposed dataset comprises 279,795 mentions of 114 entity types with up to 6 layers of nesting.
Understanding Faithfulness and Reasoning of Large Language Models on Plain Biomedical Summaries (2024.findings-emnlp)

Copied to clipboard

Challenge: Generating plain biomedical summaries with Large Language Models (LLMs) can enhance access to biomedically knowledge.
Approach: They propose a benchmark dataset with expert-annotated Faithfulness and Reasoning on plain biomedical summaries.
Outcome: The proposed dataset shows that LLMs perform poorly in generating faithful biomedical summaries and that abstractiveness and faithfulness are negatively correlated.
Born Differently Makes a Difference: Counterfactual Study of Bias in Biography Generation from a Data-to-Text Perspective (2024.acl-short)

Copied to clipboard

Challenge: Current research shows that biographies reflect bias from society such as gender and religions.
Approach: They propose a method that manipulates the personal attributes of interest while keeping the co-occurring attributes unchanged.
Outcome: The proposed method expands the analysis of gender-centered bias in text generation.
Using Similarity Measures to Select Pretraining Data for NER (N19-1)

Copied to clipboard

Challenge: Existing studies on how to select appropriate data to pretrain word vectors or LMs are lacking.
Approach: They propose to quantify aspects of similarity between pretraining and target data.
Outcome: The proposed measures are good predictors of the usefulness of pretrained models for Named Entity Recognition over 30 data pairs.
Figurative Usage Detection of Symptom Words to Improve Personal Health Mention Detection (P19-1)

Copied to clipboard

Challenge: Past work in personal health mention detection uses classification-based methods with human-engineered features or word embedding-based features.
Approach: They propose to combine a pipeline-based and a feature augmentation-based approach to combine personal health mention detection with figurative usage detection to improve the accuracy of the prediction.
Outcome: The proposed method improves the F-score of personal health mention detection by 2.21% over the pipeline-based approach and feature augmentation-based approaches.
An Effective Transition-based Model for Discontinuous NER (2020.acl-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) data sets often contain mentions consisting of discontinuous spans.
Approach: They propose a transition-based model with generic neural encoding for discontinuous NER that can recognize discontinuous mentions without sacrificing the accuracy on continuous mentions.
Outcome: The proposed model can recognize discontinuous mentions without sacrificing accuracy on continuous mentions.
My Climate CoPilot: A Question Answering System for Climate Adaptation in Agriculture (2025.acl-demo)

Copied to clipboard

Challenge: Accurately answering climate science questions requires scientific literature and climate data.
Approach: They propose to provide climate adaptation experts with information on adaptation practices relevant to a specific commodity and location.
Outcome: My Climate CoPilot is a platform that assists users to mitigate and adapt to projected climate change by providing answers grounded in evidence.
A Critical Look at Meta-evaluating Summarisation Evaluation Metrics (2024.findings-emnlp)

Copied to clipboard

Challenge: Effective summarisation evaluation metrics enable researchers and practitioners to compare different summarization systems efficiently.
Approach: They argue that evaluation metrics are primarily meta-evaluated on news summarisation datasets and that there has been a noticeable shift in research focus towards evaluating the faithfulness of generated summaries.
Outcome: The evaluation metrics are primarily meta-evaluated on news summarisation datasets and there has been a noticeable shift in research focus towards evaluating the faithfulness of generated summaries.
Cost-effective Selection of Pretraining Data: A Case Study of Pretraining BERT on Social Media (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that domain-specific BERT models can be improved when in-domain data is used for pretraining.
Approach: They propose to use Twitter and forum text as pretraining sources for two BERT models and use similarity measures to nominate in-domain data for pretraining.
Outcome: The proposed method can be used to improve performance on downstream tasks by using in-domain data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations