Automatic Period Segmentation of Oral French (2020.lrec-1)

Copied to clipboard

Challenge: Analor is a semi-automatic tool for speech segmentation in periods but it only takes into account prosodic characteristics of speech.
Approach: They propose to use a Fribourg model of macro-syntax to detect periods in syntactic and prosodic terms to develop an automatic tool for automatic segmentation of linguistic units.
Outcome: The proposed tool is compared with an existing tool Analor which divides speech into smaller segments and that CRF models detect larger segments rather than macro-syntactic periods.

Similar Papers

Prosodic segmentation for parsing spoken dialogue (2021.acl-long)

Copied to clipboard

Challenge: Existing parsers struggle to parse spoken dialogue because of disfluencies and unmarked boundaries between sentence-like units (SUs).
Approach: They hypothesize that prosody affects a parser that receives an entire dialogue turn as input, instead of gold standard pre-segmented SUs.
Outcome: The proposed model performs better than the SU-based model on the English Switchboard corpus despite performing two tasks rather than one, and pitch and intensity features are the most important for this corpus.
Developing Resources for Automated Speech Processing of Quebec French (2020.lrec-1)

Copied to clipboard

Challenge: acoustic models for automatic segmentation of Quebec French are not available for all languages . linguistic resources are developed to perform phonetic annotations in Quebec French . physical characteristics of speech can be observed in the production of sounds .
Approach: They propose to use a French lexicon to train automatic QF segmentation models . they adapt existing pronunciation dictionary and acoustic model from existing ones .
Outcome: The proposed tools perform the full process of speech segmentation in Quebec French.
A Multimodal Corpus of Expert Gaze and Behavior during Phonetic Segmentation Tasks (L18-1)

Copied to clipboard

Challenge: Phonetic segmentation is the process of splitting speech into distinct phonetic units . methods for automatic segmentation are not always accurate enough .
Approach: They propose to model phonetic segmentation as close as possible to manual segmentation by recording experts performing a segmentation task.
Outcome: This corpus captures human segmentation behavior by recording experts performing a segmentation task.
Subword Segmentation in LLMs: Looking at Inflection and Consistency (2024.emnlp-main)

Copied to clipboard

Challenge: Subword segmentation is not linguistically guided and is not currently well understood in LLMs.
Approach: They group words according to their segmentation properties and compare how well a model can solve a linguistic task for these groups using two criteria: adherence to morpheme boundaries and segmentation consistency of inflected forms of a lemma.
Outcome: The results show that the criterion of segmentation consistency can predict the model’s ability to recognize and generate the lemma from an inflected form, providing evidence that subword segmentation is relevant.
The MonPaGe_HA Database for the Documentation of Spoken French Throughout Adulthood (L18-1)

Copied to clipboard

Challenge: Existing studies on life-span changes in the speech of adults are mainly based on English speakers and few studies have compared more than two extreme age groups.
Approach: They describe a MonPaGe_HealthyAdults database of spoken french with 405 speakers aged from 20 to 93 years old.
Outcome: The proposed database includes 405 speakers aged 20 to 93 years old and includes 4 regiolects.
An Automatic Tool For Language Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: standardized tests are used to assess and screen developmental language impairments but require manual laborious transcription, annotation and calculation.
Approach: They propose to use the correct sentence and the sentence produced by patients to evaluate the level of verbal production and return a score.
Outcome: The proposed system evaluates the level of the verbal production and returns a score.
Joint Learning of Syntactic Features Helps Discourse Segmentation (2020.lrec-1)

Copied to clipboard

Challenge: Discourse segmentation is a task of fragmenting text into minimal disjoint chunks of text called Elementary Discourse Units (EDUs).
Approach: They propose a framework for multi-lingual discourse segmentation with BERT . they cast the problem as a token classification problem and jointly learn syntactic features like part-of-speech tags and dependency relations.
Outcome: Experiments in English, Dutch, German, Portuguese Brazilian and Basque show that the proposed model performs better across languages.
Improving Text Readability through Segmentation into Rheses (2024.lrec-main)

Copied to clipboard

Challenge: a new study examines the segmentation of sentences into rheses to improve readability for dyslexics . short lines of text can be beneficial for dyslexia sufferers as it limits attention span . however, random line splits can be confusing than helpful .
Approach: They propose to segment sentences into rhythmic and semantic units to improve comprehension . they also use a bilingual dataset to evaluate the efficiency of their approach .
Outcome: The proposed approach achieves an F1 score of 90.0% in English and 91.3% in French . the proposed approach also demonstrates the potential of leveraging prosodic elements .
Segmenting Natural Language Sentences via Lexical Unit Analysis (2021.findings-emnlp)

Copied to clipboard

Challenge: Recent work on sequence segmentation models suffer from invalid predictions and a lack of consistency.
Approach: They propose a unified span-based model that embeds every span and computes a score for each segmentation candidate.
Outcome: The proposed model achieves state-of-the-art on 6 of the 3 tasks tested.
Visualizing the “Dictionary of Regionalisms of France” (DRF) (L18-1)

Copied to clipboard

Challenge: a corpus of regionalisms, parts of speech and recognition rates is published in the Dictionnaire des Régionalismes de France.
Approach: They propose to curate and analyze the corpus of regionalisms published in the Dictionnaire des Régionalismes de France.
Outcome: The corpus contains all entries in the DRF for which recognition rates were recorded . the analysis compares with previous work on regionalalisms and atlas .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations