Papers by Cyril Grouin
QA Analysis in Medical and Legal Domains: A Survey of Data Augmentation in Low-Resource Settings (2025.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized natural language processing, but their success remains limited to high-resource domains. |
| Approach: | They analyze the coverage and representativeness of specialized-domain QA datasets against large-scale reference datasets. |
| Outcome: | The proposed methods and evaluations highlight the challenges faced by LLMs in low-resource domains. |
Evaluating Tokenizers Impact on OOVs Representation with Transformers Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Pre-trained Transformer models have proven their effectiveness in adapting to multiple NLP tasks and domains. |
| Approach: | They evaluated three categories of out-of-vocabulary words using three French domain-specific datasets on the legal, medical, and energetical domains to robustly analyze these categories. |
| Outcome: | The proposed models can create new representations for out-of-vocabulary words by adding external morpho-syntactic context rather than improving the semantic understanding of the words directly. |
Does the structure of textual content have an impact on language models for automatic summarization? (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing models for automatic summarization of long sequences suffer from context limitation. |
| Approach: | They propose to take into account textual information coming from distinct passages from the long texts to be summarized. |
| Outcome: | The proposed model improves on the performance of LongFormer on English. |
Enriching a Time-Domain Astrophysics Corpus with Named Entity, Coreference and Astrophysical Relationship Annotations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing corpora for astrophysical natural language processing are limited to Named Entity Recognition tasks, leaving a gap in resource diversity. |
| Approach: | They propose to expand astroECR to cover named entities, coreferences, annotations related to aastrphysical relationships, and normalizing celestial object names. |
| Outcome: | The proposed model extends the time-domain astrophysics corpus to include named entities, coreferences, and annotations related to aastrphysical relationships. |
Three Dimensions of Reproducibility in Natural Language Processing (L18-1)
Copied to clipboard
K. Bretonnel Cohen, Jingbo Xia, Pierre Zweigenbaum, Tiffany Callahan, Orin Hargraves, Foster Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, Lawrence E. Hunter
| Challenge: | a recent editorial on reproducibility in language processing defined three dimensions of reproducibility . authors had already submitted a correction, but there is no consensus on the definitions . |
| Approach: | They propose an ontology of reproducibility in natural language processing to address these problems . they propose to analyze three dimensions of reproducible in natural languages papers . authors propose to use a 'replicability' term to describe the reproducibility of a conclusion, finding, value . |
| Outcome: | The proposed ontology aims to enhance future research and communication about the topic and retrospective meta-analyses. |
Inference Annotation of a Chinese Corpus for Opinion Mining (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing tools for opinion mining can accurately predict the writer's attitude in simple explicit sentences. |
| Approach: | They propose to define inference, classify different types and provide an annotation framework to analyze the annotation results. |
| Outcome: | The proposed framework defines inference type, polarity and topic and analyzes the results. |
A Dataset for Pharmacovigilance in German, French, and Japanese: Annotating Adverse Drug Reactions across Languages (2024.lrec-main)
Copied to clipboard
Lisa Raithel, Hui-Syuan Yeh, Shuntaro Yada, Cyril Grouin, Thomas Lavergne, Aurélie Névéol, Patrick Paroubek, Philippe Thomas, Tomohiro Nishiyama, Sebastian Möller, Eiji Aramaki, Yuji Matsumoto, Roland Roller, Pierre Zweigenbaum
| Challenge: | Existing clinical corpora mostly revolves around scientific articles in English . existing literature is limited to only a few scientific articles . |
| Approach: | They propose to use user-generated data sources to uncover adverse drug reactions . existing clinical corpora mostly revolves around scientific articles in english . authors provide statistics to highlight certain challenges associated with the corpus . |
| Outcome: | The proposed corpus includes 12 entity types, four attribute types, and 13 relation types . it provides strong baselines for extracting entities and relations between entities . |