Papers by Iris Eshkol-Taravella
Automatic Period Segmentation of Oral French (2020.lrec-1)
Copied to clipboard
| Challenge: | Analor is a semi-automatic tool for speech segmentation in periods but it only takes into account prosodic characteristics of speech. |
| Approach: | They propose to use a Fribourg model of macro-syntax to detect periods in syntactic and prosodic terms to develop an automatic tool for automatic segmentation of linguistic units. |
| Outcome: | The proposed tool is compared with an existing tool Analor which divides speech into smaller segments and that CRF models detect larger segments rather than macro-syntactic periods. |
An Empirical Examination of Online Restaurant Reviews (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for opinion mining and sentiment analysis focus on extracting either positive or negative opinions from texts and determining the targets of these opinions. |
| Approach: | They propose a corpus-based scheme that detects evaluative language at a finer-grained level. |
| Outcome: | The proposed scheme classifies each sentence into one of four evaluation types based on the proposed scheme. |
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains (2024.lrec-main)
Copied to clipboard
Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab
| Challenge: | Pretrained language models are the de facto backbone of most state-of-the-art NLP systems. |
| Approach: | They propose a family of domain-specific pretrained PLMs for French focusing on three important domains: transcribed speech, medicine, and law. |
| Outcome: | The proposed models perform better on transcribed speech, medicine, and law domains than state-of-the-art models on a diverse set of tasks and datasets. |
Chunk Different Kind of Spoken Discourse: Challenges for Machine Learning (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing chunkers for spoken data are based on a corpus composed of monologues and spontaneous talk in interaction. |
| Approach: | They propose to use CRFs to develop a chunker for spoken data . the chunker is based on a small corpus composed of two kinds of discourse . |
| Outcome: | The proposed chunker is based on a spoken corpus composed of monologue and spontaneous talk in interaction. |
What Speakers really Mean when they Ask Questions: Classification of Intentions with a Supervised Approach (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on hidden intentions of speakers in questions during meals is based on written or oral data, which are less easy to interpret. |
| Approach: | They propose a typology of hidden intentions in questions asked during meals . they implement an automatic classification model based on annotated data and selected linguistic features. |
| Outcome: | The proposed model is based on annotated data and features and evaluates its performance. |