Papers by Loïc Grobol
ARBRES Kenstur: A Breton-French Parallel Corpus Rooted in Field Linguistics (2024.lrec-main)
Copied to clipboard
| Challenge: | ARBRES is a project documenting the Breton language and state of research and engineering in linguistics and NLP. |
| Approach: | ARBRES is an ongoing project of open science documenting the Breton language and state of research and engineering in linguistics and NLP. |
| Outcome: | ARBRES Kenstur is a project documenting the Breton language and state of research and engineering in linguistics and NLP. |
Automatic Period Segmentation of Oral French (2020.lrec-1)
Copied to clipboard
| Challenge: | Analor is a semi-automatic tool for speech segmentation in periods but it only takes into account prosodic characteristics of speech. |
| Approach: | They propose to use a Fribourg model of macro-syntax to detect periods in syntactic and prosodic terms to develop an automatic tool for automatic segmentation of linguistic units. |
| Outcome: | The proposed tool is compared with an existing tool Analor which divides speech into smaller segments and that CRF models detect larger segments rather than macro-syntactic periods. |
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)
Copied to clipboard
| Challenge: | a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages. |
| Approach: | They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data . |
| Outcome: | The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality . |
Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages (2024.naacl-long)
Copied to clipboard
Nathaniel Robinson, Raj Dabre, Ammon Shurtz, Rasul Dent, Onenamiyi Onesi, Claire Monroc, Loïc Grobol, Hasan Muhammad, Ashi Garg, Naome Etori, Vijay Murari Tiyyala, Olanrewaju Samuel, Matthew Stutzman, Bismarck Odoom, Sanjeev Khudanpur, Stephen Richardson, Kenton Murray
| Challenge: | Creole languages are used in much of Latin America, Africa and the Caribbean . a large multilingual bitext like ours has potential to build the best yet or first ever MT models for many languages . |
| Approach: | They present the largest cumulative dataset to date for Creole language MT . they provide MT models supporting all 41 Creoles in 172 translation directions . |
| Outcome: | The proposed model outperforms a genre-specific Creole MT model on its own benchmark for 23 of 34 translation directions. |
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)
Copied to clipboard
| Challenge: | ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones. |
| Approach: | They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions. |
| Outcome: | The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format. |