Papers by Jean-Philippe Goldman

7 papers
NeuroTrialNER: An Annotated Corpus for Neurological Diseases and Therapies in Clinical Trial Registries (2024.emnlp-main)

Copied to clipboard

Challenge: Despite substantial investment, developing new treatments for neurological conditions is a challenging and often unsuccessful endeavour.
Approach: They propose a corpus for named entity recognition that is annotated clinical trial summaries from ClinicalTrials.gov.
Outcome: The proposed corpus is annotated for neurological diseases, therapeutic interventions, and control treatments and achieves a close-to-human performance.
MIAPARLE: Online training for the discrimination of stress contrasts (L18-1)

Copied to clipboard

Challenge: Second language learners tend to imprint the prosody of their mother language onto the second language (L2) . this can hamper communication between learners and natives, and can also affect the credibility of learners and how they are evaluated by others.
Approach: They propose a tool that focuses on stress perception for speakers whose L1 is a fixed-stress language, such as French.
Outcome: The tool is particularly useful for speakers whose L1 is a fixed-stress language, such as French.
Strategies and Challenges for Crowdsourcing Regional Dialect Perception Data for Swiss German and Swiss French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing project in the field of Swiss German dialects and Swiss French accents collects linguistic data.
Approach: a gamified crowdsourcing platform was set up to collect linguistic data on Swiss German and Swiss French accents.
Outcome: a gamified crowdsourcing platform collects linguistic data on Swiss German and Swiss French accents . the platform has provided 470,000 localizations, with 7,500 registered users and 30,000 anonymous visitors .
FRASIMED: A Clinical French Annotated Resource Produced through Crosslingual BERT-Based Annotation Projection (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for generating annotations for large datasets are time-consuming and resource-intensive.
Approach: They propose a method for generating translated versions of annotated datasets through crosslingual annotation projection.
Outcome: The proposed method shows that it is efficient and high-quality in the resulting dataset.
UniMorph 4.0: Universal Morphology (2022.lrec-1)

Copied to clipboard

Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
Challenge: The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages.
Approach: They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema.
Outcome: The proposed schema has added 66 new languages, including 24 endangered languages.
Crowdsourcing Regional Variation Data and Automatic Geolocalisation of Speakers of European French (L18-1)

Copied to clipboard

Challenge: a crowdsourcing platform is used to collect linguistic data and document language use, with a focus on regional variation in European French.
Approach: They propose a crowdsourcing platform to collect linguistic data and document language use with a special focus on regional variation in European French.
Outcome: The proposed platform collects linguistic data and documents language use with a special focus on regional variation in European French.
Bratly: A Python Extension for BRAT Functionalities (2025.emnlp-demos)

Copied to clipboard

Challenge: BRAT is a widely used web-based text annotation tool, but lacks robust Python support for effective annotation management and processing.
Approach: They propose an open-source extension of BRAT that introduces a solid Python backend and enables advanced annotation functions such as annotation typings, collection typings with statistical insights, corpus and annotation handling, object modifications, and entity-level evaluation.
Outcome: The proposed extension streamlines annotation workflows, improves usability, and facilitates high-quality NLP research.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations