Papers by Loïc Grobol

5 papers
ARBRES Kenstur: A Breton-French Parallel Corpus Rooted in Field Linguistics (2024.lrec-main)

Copied to clipboard

Challenge: ARBRES is a project documenting the Breton language and state of research and engineering in linguistics and NLP.
Approach: ARBRES is an ongoing project of open science documenting the Breton language and state of research and engineering in linguistics and NLP.
Outcome: ARBRES Kenstur is a project documenting the Breton language and state of research and engineering in linguistics and NLP.
Automatic Period Segmentation of Oral French (2020.lrec-1)

Copied to clipboard

Challenge: Analor is a semi-automatic tool for speech segmentation in periods but it only takes into account prosodic characteristics of speech.
Approach: They propose to use a Fribourg model of macro-syntax to detect periods in syntactic and prosodic terms to develop an automatic tool for automatic segmentation of linguistic units.
Outcome: The proposed tool is compared with an existing tool Analor which divides speech into smaller segments and that CRF models detect larger segments rather than macro-syntactic periods.
BERTrade: Using Contextual Embeddings to Parse Old French (2022.lrec-1)

Copied to clipboard

Challenge: a growing interest in digital humanities for automatic processing and annotation of historical texts is generating new models for historical languages.
Approach: They use POS-tagging and dependency parsing to evaluate contextual word embedding models . Old French is one of the historical languages for which they have the largest amount of syntactically annotated data .
Outcome: The proposed model can be used to improve performance in Old French, the authors show . they use POS-tagging and dependency parsing to evaluate the model's quality .
Kreyòl-MT: Building MT for Latin American, Caribbean and Colonial African Creole Languages (2024.naacl-long)

Copied to clipboard

Challenge: Creole languages are used in much of Latin America, Africa and the Caribbean . a large multilingual bitext like ours has potential to build the best yet or first ever MT models for many languages .
Approach: They present the largest cumulative dataset to date for Creole language MT . they provide MT models supporting all 41 Creoles in 172 translation directions .
Outcome: The proposed model outperforms a genre-specific Creole MT model on its own benchmark for 23 of 34 translation directions.
ANCOR-AS: Enriching the ANCOR Corpus with Syntactic Annotations (L18-1)

Copied to clipboard

Challenge: ANCOR-AS is an enriched version of the ANCor corpus that adds syntactic annotations in addition to the existing coreference and speech transcription ones.
Approach: They propose to use syntactic annotations in addition to existing coreference and speech transcription annotations to improve detection of mentions.
Outcome: The proposed version adds syntactic annotations to existing coreference and speech transcription annotations and is released in a new TEI-compliant XML format.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations