Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
Similar Papers
DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical domains (2023.acl-long)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Richard Dufour, Mickael Rouvier, Emmanuel Morin, Béatrice Daille, Pierre-Antoine Gourraud
| Challenge: | Recent studies have shown that pre-trained language models improve performance on a wide range of NLP tasks. |
| Approach: | They propose to use pre-trained language models to train medical domains on French language to compare performance with specialized ones. |
| Outcome: | The proposed models can take advantage of existing biomedical models in a foreign language by further pre-training them on our targeted data. |
DrBenchmark: A Large Language Understanding Evaluation Benchmark for French Biomedical Domain (2024.lrec-main)
Copied to clipboard
Yanis Labrak, Adrien Bazoge, Oumaima El Khettari, Mickael Rouvier, Pacome Constant Dit Beaufils, Natalia Grabar, Béatrice Daille, Solen Quiniou, Emmanuel Morin, Pierre-Antoine Gourraud, Richard Dufour
| Challenge: | Existing benchmarks for pre-trained language models are limited to only a few languages . a limited number of tasks are evaluated on non-standardized protocols . |
| Approach: | They propose to aggregate diverse downstream tasks into a benchmark to assess PLMs' qualities . they evaluate 8 pre-trained masked language models on general and biomedical-specific data . |
| Outcome: | The proposed benchmark assesses pre-trained language models on 20 diversified tasks. |
Exploiting Language Characteristics for Legal Domain-Specific Language Model Pretraining (2023.findings-eacl)
Copied to clipboard
| Challenge: | Pretraining large language models has resulted in tremendous performance improvement for many natural language processing tasks. |
| Approach: | They propose to incorporate pretraining objectives that explicitly exploit domain specific language characteristics into the model. |
| Outcome: | The proposed objectives target token-level feature representation and incorporate sentence level semantics. |
Cross-domain Analysis on Japanese Legal Pretrained Language Models (2022.findings-aacl)
Copied to clipboard
| Challenge: | Existing studies do not care the performance of domain-adapted PLMs for a generic domain. |
| Approach: | They propose to use pretraining strategies to build pretrained language models specialised in the legal domain to improve their performance. |
| Outcome: | The pretrained language models can learn domain-specific and general word meanings simultaneously and can distinguish them. |
AdminSet and AdminBERT: a Dataset and a Pre-trained Language Model to Explore the Unstructured Maze of French Administrative Documents (2025.coling-main)
Copied to clipboard
| Challenge: | Pre-trained language models are used to analyze documents but administrative texts are unstructured and do not perform well. |
| Approach: | They propose a French pre-trained language model for the administrative domain . they compare it with a general domain language model and a large language model . |
| Outcome: | The proposed model improves performance on administrative and general domains. |
Recent Advances in Pre-trained Language Models: Why Do They Work and How Do They Work (2022.aacl-tutorials)
Copied to clipboard
| Challenge: | Pre-trained language models are language models that are pre-taught on large-scaled corpora in a self-supervised fashion. |
| Approach: | This tutorial provides a broad and comprehensive introduction to pre-trained language models . it focuses on emerging methods that enable PLMs to perform diverse downstream tasks . |
| Outcome: | This tutorial focuses on the benefits of pre-trained language models and how to use them in NLP tasks. |
Evaluating Pretraining Strategies for Clinical BERT Models (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing generic language models in specialized domains may be sub-optimal due to domain differences. |
| Approach: | They propose various strategies for adapting a generic language model to the target domain and various forms of vocabulary modifications to fine-tune it. |
| Outcome: | The proposed strategies outperform a general-domain language model but little difference in performance between the models. |
LaoPLM: Pre-trained Language Models for Lao (2022.lrec-1)
Copied to clipboard
| Challenge: | Pre-trained language models (PLMs) can capture different levels of concepts in context . previous work on Lao has been hampered by the lack of annotated datasets . |
| Approach: | They construct a text classification dataset to alleviate the resource-scarce situation of Lao . they evaluate them on two downstream tasks: part-of-speech tagging and text classification . |
| Outcome: | The proposed model can capture different levels of concepts in context and generate universal language representations. |
mDAPT: Multilingual Domain Adaptive Pretraining in a Single Model (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing domain-specific multilingual pretraining data is difficult to obtain due to regulations, legislation, or simply a lack of language- and domain- specific text. |
| Approach: | They propose to continue pretraining a language model on domain-specific unlabelled text . this allows for better modelling of text for downstream tasks within the domain . |
| Outcome: | The proposed approach outperforms the general multilingual model and performs close to its monolingual counterpart. |
Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing (2022.emnlp-main)
Copied to clipboard
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais
| Challenge: | Existing pre-trained language models are not well-explored and are not reproducible in the literature. |
| Approach: | They propose to improve existing Arabic language pre-trained language models using a more methodical approach. |
| Outcome: | The proposed models outperform existing models on ALUE, a leaderboard-powered benchmark for Arabic NLU and NLG tasks. |