Challenge: Language documentation often requires segmenting transcriptions of utterances into words and morphemes . a long tradition of nonparametric Bayesian models is used to handle these tasks .
Approach: They propose a Bayesian model for simultaneously segmenting utterances at two levels . they use two under-resourced languages to better understand the value of weak supervision .
Outcome: The proposed model can be used to identify language documents with weak supervision.

Similar Papers

Weakly Supervised Word Segmentation for Computational Language Documentation (2022.acl-long)

Copied to clipboard

Challenge: a recent paper aims to improve the effectiveness of unsupervised language analysis techniques in low resource settings.
Approach: They propose to use a weak supervision to improve linguistic segmentation in low resource languages . they propose to provide linguists with LTs that can be used to create interactive annotation tools .
Outcome: The proposed models can be used to improve the quality of language segmentation in low resource languages.
Massively Multilingual Joint Segmentation and Glossing (2026.acl-long)

Copied to clipboard

Challenge: Existing models generate morpheme-level glosses but assign them to whole words without predicting the actual morphological boundaries, making them less interpretable and therefore untrustworthy to human annotators.
Approach: They propose to use neural networks to predict interlinear glosses and morphological segmentation from raw text.
Outcome: The proposed model outperforms GlossLM on glossing and beats open-source models on segmentation, glossing, and alignment.
Joint Dialogue Topic Segmentation and Categorization: A Case Study on Clinical Spoken Conversations (2023.emnlp-industry)

Copied to clipboard

Challenge: Utilizing natural language processing in clinical conversations is effective to improve the efficiency of workflows for medical staff and patients.
Approach: They propose a model for dialogue segmentation and topic categorization that integrates natural language processing techniques into a joint model.
Outcome: The proposed model improves on follow-up calls for diabetes management and reduces computational complexity and cost.
A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images (D18-1)

Copied to clipboard

Challenge: Existing approaches combine language and perception to infer word embeddings . however, the embeddables produced by such models do not reflect the actual word representations.
Approach: They propose a probabilistic model that integrates linguistic and perceptual inputs to explain observed word-context pairs in a text corpus.
Outcome: The proposed model achieves competitive or stronger results on tasks of assessing pairwise word similarity and image/caption retrieval compared to other state-of-the-art models.
LLMSegm: Surface-level Morphological Segmentation Using Large Language Model (2024.lrec-main)

Copied to clipboard

Challenge: Existing approaches to morphological segmentation split word into its morphemes . LLMSegm is applicable in low-data settings and low-resourced languages .
Approach: They propose a novel approach to surface-level morphological segmentation leveraging large language models.
Outcome: The proposed method is applicable in low-data settings and low-resource languages.
A Joint Model for Document Segmentation and Segment Labeling (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to text segmentation focus on document segmentation and segment labeling separately.
Approach: They propose a method for jointly segmenting a document and labeling segments . they show that S-LSTM reduces segmentation error by 30% on average .
Outcome: The proposed method reduces segmentation error by 30% while improving segment labeling.
Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies using LLMs on psycholinguistic data have gone unverified . a growing body of research is using word-level prediction as a computational proxy .
Approach: They compare morphological, morphologic, and BPE tokenization estimates with reading time data.
Outcome: The proposed method could be used to evaluate morphological prediction.
Joint Learning of Syntactic Features Helps Discourse Segmentation (2020.lrec-1)

Copied to clipboard

Challenge: Discourse segmentation is a task of fragmenting text into minimal disjoint chunks of text called Elementary Discourse Units (EDUs).
Approach: They propose a framework for multi-lingual discourse segmentation with BERT . they cast the problem as a token classification problem and jointly learn syntactic features like part-of-speech tags and dependency relations.
Outcome: Experiments in English, Dutch, German, Portuguese Brazilian and Basque show that the proposed model performs better across languages.
Tackling the Low-resource Challenge for Canonical Segmentation (2020.emnlp-main)

Copied to clipboard

Challenge: morphological segmentation is a task of dividing words into their constituting morphemes . we compare two new approaches for the task when training data is limited .
Approach: They propose to use an LSTM pointer-generator and a sequence-to-sequence model to perform canonical segmentation when training data is limited.
Outcome: The proposed models outperform existing models on German, English, and Indonesian in low-resource scenarios by 11.4% accuracy.
Learning to Discover, Ground and Use Words with Segmental Neural Language Models (P19-1)

Copied to clipboard

Challenge: Existing models of word learning do not account for the long-range dependencies manifest in language and that are easily captured by recurrent neural networks.
Approach: They propose a segmental neural language model that unifies word discovery, learning how words fit together to form sentences, and by conditioning the model on visual context, how words’ meanings ground in representations of nonlinguistic modalities.
Outcome: The proposed model learns predictive distributions better than character LSTM models, discovers words competitively with nonparametric Bayesian word segmentation models, and improves on both.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations