Papers by Karën Fort

12 papers
“Women do not have heart attacks!” Gender Biases in Automatically Generated Clinical Cases in French (2025.findings-naacl)

Copied to clipboard

Challenge: Healthcare professionals are increasingly including Language Models (LMs) in clinical practice.
Approach: They propose to use LMs to generate clinical cases in french and an automatic linguistic gender detection tool to measure gender biases.
Outcome: The proposed model over-generates cases describing male patients, creating synthetic corpora that are not consistent with documented prevalence for these disorders.
Quantification Annotation in ISO 24617-12, Second Draft (2022.lrec-1)

Copied to clipboard

Challenge: a project aimed at establishing an interoperable annotation schema for quantification phenomena was relaunched in early 2022 due to the Covid-19 pandemic .
Approach: This paper describes the continuation of a project that aims at establishing an interoperable annotation schema for quantification phenomena as part of the ISO suite of semantic annotation standards.
Outcome: The proposed schema is part of the ISO suite of semantic annotation standards known as the Semantic Annotation Framework (SemAF).
French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English (2022.acl-long)

Copied to clipboard

Challenge: We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age.
Approach: They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures.
Outcome: The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language.
Reviewing Natural Language Processing Research (2021.eacl-tutorials)

Copied to clipboard

Challenge: a tutorial on reviewing is a useful tool for researchers who are new to the field of NLP.
Approach: this tutorial provides an opportunity to learn the basics of reviewing . more experienced researchers might find this tutorial interesting to revise their reviewing procedure.
Outcome: This tutorial teaches researchers how to revise their reviewing procedure .
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)

Copied to clipboard

Challenge: Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones.
Approach: They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises .
Outcome: The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs.
CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives (2022.lrec-1)

Copied to clipboard

Challenge: Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluation.
Approach: They propose to annotate sentences in French using a definition of similarity guided by clinical facts and use it to evaluate the corpus.
Outcome: The proposed model can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set.
Rigor Mortis: Annotating MWEs with a Gamified Platform (2020.lrec-1)

Copied to clipboard

Challenge: gamification of the platform should be improved, in order to attract and retain more players.
Approach: They propose to use a gamified crowdsourcing platform to evaluate the intuition of speakers and then train them to annotate multi-word expressions in French corpora.
Outcome: The proposed platform evaluates the speakers' intuition and trains them to annotate multi-word expressions in French corpora.
Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing (L18-1)

Copied to clipboard

Challenge: Using a crowdsourcing platform, we collected 18,917 annotations for a less-resourced French regional language, Alsatian.
Approach: They developed a platform that allows people to gather part-of-speech annotations on a variety of corpora and train a first tagger specific to Alsatian.
Outcome: The proposed method is valid for Alsatian and can be adapted to other languages.
Navigating Ethical Challenges in NLP: Hands-on strategies for students and researchers (2025.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review .
Approach: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review .
Outcome: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review .
Do we Name the Languages we Study? The #BenderRule in LREC and ACL articles (2022.lrec-1)

Copied to clipboard

Challenge: Using the #BenderRule, we examine the number and which languages are studied in two NLP conferences.
Approach: They examine the application of the #BenderRule in NLP articles by inspecting 14,000 articles over a period of time ranging from 2000 to 2020 for LREC and 1979 to 2020 respectively.
Outcome: The authors examine the application of the #BenderRule in natural language processing articles over a period of time ranging from 2000 to 2020 for LREC and ACL.
Understanding Ethics in NLP Authoring and Reviewing (2023.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues .
Approach: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . the methodology is interactive and participatory, including case studies and working in groups .
Outcome: This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . the methodology is interactive and participatory, including case studies and working in groups.
Reviewing Natural Language Processing Research (2020.acl-tutorials)

Copied to clipboard

Challenge: a tutorial on reviewing research in natural language processing will cover the theory and practice of reviewing research.
Approach: tutorial covers the theory and practice of reviewing research in natural language processing . authors say reviewers should be more aware of "false negatives"
Outcome: tutorial covers the theory and practice of reviewing research in natural language processing . authors say their reviews leave something to be desired .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations