Papers by Delphine Charlet

5 papers
A Multimodal Educational Corpus of Oral Courses: Annotation, Analysis and Case Study (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of spontaneous speech is being developed for educational use . the dataset will be freely available to the research community .
Approach: They propose to use a French speech educational corpus to explore synchronous speech transcription and application in teaching situations.
Outcome: The proposed corpus includes 10 hours of lectures, manually transcribed and segmented . the dataset will be freely available to the research community .
CALOR-QUEST : generating a training corpus for Machine Reading Comprehension models from shallow semantic annotations (D19-58)

Copied to clipboard

Challenge: Recent large corpora of triplets have opened the door to supervised machine learning approaches for Question-Answering.
Approach: They propose to generate questions from the semantic Frame analysis of large corpora using a CALOR-QUEST resource in French and use it to improve machine reading comprehension.
Outcome: The proposed method generates questions from the semantic Frame analysis of large corpora and then tests them on the CALOR-QUEST resource in French.
Handling Normalization Issues for Part-of-Speech Tagging of Online Conversational Text (L18-1)

Copied to clipboard

Challenge: a new approach to POS tagging noisy user generated text is proposed . word embeddings are trained on a noisy corpus to address both normalization and POS.
Approach: They propose to use word embeddings to normalize text before tagging it, while a gated neural network based tagger handles the remaining errors.
Outcome: The proposed approach normalizes some errors before tagging, while a gated neural network handles the remaining errors.
Cross-lingual and Cross-domain Evaluation of Machine Reading Comprehension with Squad and CALOR-Quest Corpora (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has shown that language mismatch and domain mismatch can affect performance of a machine reading task . a factor between language mismatched and domain-mismatched has the strongest influence on performance .
Approach: They compare the cross-language and cross-domain capabilities of BERT on a machine reading comprehension task on two corpora: SQuAD and a new French Machine Reading dataset.
Outcome: The proposed model matches human performance on a machine reading comprehension task with BERT on Chinese and French documents with interesting results.
FrNewsLink : a corpus linking TV Broadcast News Segments and Press Articles (L18-1)

Copied to clipboard

Challenge: a corpus of TV Broadcast News resources is proposed to address several applicative tasks.
Approach: They propose to use a corpus to address several applicative tasks that are made public . they propose to gather TVBN shows and press articles and use them to study semantic similarity .
Outcome: The proposed corpus is based on 112 TVBN shows and press articles . it allows to study semantic similarity and multimedia News linking .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations