Papers by Karën Fort
“Women do not have heart attacks!” Gender Biases in Automatically Generated Clinical Cases in French (2025.findings-naacl)
Copied to clipboard
| Challenge: | Healthcare professionals are increasingly including Language Models (LMs) in clinical practice. |
| Approach: | They propose to use LMs to generate clinical cases in french and an automatic linguistic gender detection tool to measure gender biases. |
| Outcome: | The proposed model over-generates cases describing male patients, creating synthetic corpora that are not consistent with documented prevalence for these disorders. |
Quantification Annotation in ISO 24617-12, Second Draft (2022.lrec-1)
Copied to clipboard
Harry Bunt, Maxime Amblard, Johan Bos, Karën Fort, Bruno Guillaume, Philippe de Groote, Chuyuan Li, Pierre Ludmann, Michel Musiol, Siyana Pavlova, Guy Perrier, Sylvain Pogodalla
| Challenge: | a project aimed at establishing an interoperable annotation schema for quantification phenomena was relaunched in early 2022 due to the Covid-19 pandemic . |
| Approach: | This paper describes the continuation of a project that aims at establishing an interoperable annotation schema for quantification phenomena as part of the ISO suite of semantic annotation standards. |
| Outcome: | The proposed schema is part of the ISO suite of semantic annotation standards known as the Semantic Annotation Framework (SemAF). |
French CrowS-Pairs: Extending a challenge dataset for measuring social bias in masked language models to a language other than English (2022.acl-long)
Copied to clipboard
| Challenge: | We introduce 1,679 sentence pairs in French that cover stereotypes in ten types of bias like gender and age. |
| Approach: | They build on the US-centered CrowS-pairs dataset to create a multilingual stereotypes dataset that allows for comparability across languages and cultures. |
| Outcome: | The proposed dataset allows for comparability across languages while characterizing biases that are specific to each country and language. |
Reviewing Natural Language Processing Research (2021.eacl-tutorials)
Copied to clipboard
| Challenge: | a tutorial on reviewing is a useful tool for researchers who are new to the field of NLP. |
| Approach: | this tutorial provides an opportunity to learn the basics of reviewing . more experienced researchers might find this tutorial interesting to revise their reviewing procedure. |
| Outcome: | This tutorial teaches researchers how to revise their reviewing procedure . |
Creating Expert Knowledge by Relying on Language Learners: a Generic Approach for Mass-Producing Language Resources by Combining Implicit Crowdsourcing and Language Learning (2020.lrec-1)
Copied to clipboard
Lionel Nicolas, Verena Lyding, Claudia Borg, Corina Forascu, Karën Fort, Katerina Zdravkova, Iztok Kosem, Jaka Čibej, Špela Arhar Holdt, Alice Millour, Alexander König, Christos Rodosthenous, Federico Sangati, Umair ul Hassan, Anisia Katinskaia, Anabela Barreiro, Lavinia Aparaschivei, Yaakov HaCohen-Kerner
| Challenge: | Lack of wide-coverage and high-quality LRs is a longstanding issue in natural language processing (NLP) however, there are no large initiatives of similar scale for creating new LR or improving existing ones. |
| Approach: | They propose a generic approach to combine implicit crowdsourcing and language learning to mass-produce language resources (LRs) they describe its core paradigm that consists in pairing specific types of LRs with specific exercises . |
| Outcome: | The proposed approach can be used in several learning scenarios to produce a multitude of NLP resources and alleviate the long-standing issue of the lack of LRs. |
CLISTER : A Corpus for Semantic Textual Similarity in French Clinical Narratives (2022.lrec-1)
Copied to clipboard
| Challenge: | Modern Natural Language Processing relies on the availability of annotated corpora for training and evaluation. |
| Approach: | They propose to annotate sentences in French using a definition of similarity guided by clinical facts and use it to evaluate the corpus. |
| Outcome: | The proposed model can capture similarity with state-of-the-art performance on the DEFT STS shared task evaluation data set. |
Rigor Mortis: Annotating MWEs with a Gamified Platform (2020.lrec-1)
Copied to clipboard
| Challenge: | gamification of the platform should be improved, in order to attract and retain more players. |
| Approach: | They propose to use a gamified crowdsourcing platform to evaluate the intuition of speakers and then train them to annotate multi-word expressions in French corpora. |
| Outcome: | The proposed platform evaluates the speakers' intuition and trains them to annotate multi-word expressions in French corpora. |
Toward a Lightweight Solution for Less-resourced Languages: Creating a POS Tagger for Alsatian Using Voluntary Crowdsourcing (L18-1)
Copied to clipboard
| Challenge: | Using a crowdsourcing platform, we collected 18,917 annotations for a less-resourced French regional language, Alsatian. |
| Approach: | They developed a platform that allows people to gather part-of-speech annotations on a variety of corpora and train a first tagger specific to Alsatian. |
| Outcome: | The proposed method is valid for Alsatian and can be adapted to other languages. |
Navigating Ethical Challenges in NLP: Hands-on strategies for students and researchers (2025.acl-tutorials)
Copied to clipboard
Luciana Benotti, Fanny Ducel, Karën Fort, Guido Ivetta, Zhijing Jin, Min-Yen Kan, Seunghun J. Lee, Minzhi Li, Margot Mieskes, Adriana Pagano
| Challenge: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review . |
| Approach: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review . |
| Outcome: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . participants will gain practical experience on when to flag a paper for ethics review . |
Do we Name the Languages we Study? The #BenderRule in LREC and ACL articles (2022.lrec-1)
Copied to clipboard
| Challenge: | Using the #BenderRule, we examine the number and which languages are studied in two NLP conferences. |
| Approach: | They examine the application of the #BenderRule in NLP articles by inspecting 14,000 articles over a period of time ranging from 2000 to 2020 for LREC and 1979 to 2020 respectively. |
| Outcome: | The authors examine the application of the #BenderRule in natural language processing articles over a period of time ranging from 2000 to 2020 for LREC and ACL. |
Understanding Ethics in NLP Authoring and Reviewing (2023.eacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . |
| Approach: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . the methodology is interactive and participatory, including case studies and working in groups . |
| Outcome: | This tutorial will equip participants with basic guidelines for thinking deeply about ethical issues . the methodology is interactive and participatory, including case studies and working in groups. |
Reviewing Natural Language Processing Research (2020.acl-tutorials)
Copied to clipboard
| Challenge: | a tutorial on reviewing research in natural language processing will cover the theory and practice of reviewing research. |
| Approach: | tutorial covers the theory and practice of reviewing research in natural language processing . authors say reviewers should be more aware of "false negatives" |
| Outcome: | tutorial covers the theory and practice of reviewing research in natural language processing . authors say their reviews leave something to be desired . |