| Challenge: | a system for automatic diacritization of Hebrew Text is available for both casual and expert users. |
| Approach: | They propose a system for automatic diacritization of Hebrew Text . the system combines declarative linguistic knowledge with machine learning models . |
| Outcome: | The proposed system is available for both casual and expert users. |
Similar Papers
Restoring Hebrew Diacritics Without a Dictionary (2022.findings-naacl)
Copied to clipboard
| Challenge: | a number of modern Hebrew texts are written in a letter-only version of the Hebrew script, which omits the diacritics present in the full diacritized, or dotted variant. |
| Approach: | They propose a character-level LSTM that can accurately diacritize Hebrew script without human-curated resources. |
| Outcome: | The proposed model performs on par with complex curation-dependent systems across a diverse array of modern Hebrew sources. |
Neural Arabic Text Diacritization: State of the Art Results and a Novel Approach for Machine Translation (D19-52)
Copied to clipboard
| Challenge: | a number of Arabic text diacritizers use diacritics to convey information about meaning of a word . Arabic text to speech (TTS) requires a complex process to determine the correct diacritical for each character . |
| Approach: | They propose to use Arabic diacritization to enhance machine translation models . they propose to build automatic Arabic text diacritics using two approaches . |
| Outcome: | The proposed models are either better or on par with other models, which require language-dependent post-processing steps, unlike ours. |
Advancing Arabic Diacritization: Improved Datasets, Benchmarking, and State-of-the-Art Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Arabic diacritics are typically omitted in written Arabic, leading to ambiguity . authors propose a methodology to analyze and refine a large diacritized corpus . |
| Approach: | They propose a methodology to analyze and refine a large diacritized corpus to improve training quality. |
| Outcome: | The proposed model achieves state-of-the-art results with 3.12% and 2.70% WER on WikiNews-2014 and Wikinews-2024. |
What’s Wrong with Hebrew NLP? And How to Make it Right (D19-3)
Copied to clipboard
| Challenge: | Sub-optimal performance of many morphologically rich languages (MRLs) is due to errors in early morphology disambiguation decisions, that cannot be recovered later on in the pipeline, yielding incoherent annotations on the whole. |
| Approach: | They propose to use a joint morpho-syntactic infrastructure for processing Modern Hebrew texts to provide rich and expressive annotations. |
| Outcome: | The proposed pipelines are based on a morpho-syntactic infrastructure for processing Modern Hebrew texts. |
AlephBERT: Language Model Pre-training and Evaluation from Sub-Word to Sentence Level (2022.acl-long)
Copied to clipboard
| Challenge: | a recent study shows that large pre-trained language models are not sufficient for Hebrew. |
| Approach: | They propose a large pre-trained language model for Hebrew that recovers morphological segments encoded in contextualized embedding vectors. |
| Outcome: | The proposed model obtains state-of-the-art on all tasks beyond contemporary Hebrew baselines. |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |
Don’t Touch My Diacritics (2025.naacl-short)
Copied to clipboard
| Challenge: | a recent paper examines the effects of preprocessing text with diacritics on model performance . we show that inconsistent encoding of diacritized characters and removing diacritical characters can have detrimental downstream effects . |
| Approach: | They propose to improve the handling of diacritized text by preserving diacritics and removing them altogether. |
| Outcome: | The proposed approach reduces the number of errors in the preprocessing process, the authors argue . they show that the proposed approach can reduce the number and complexity of errors . |
Highly Effective Arabic Diacritization using Sequence to Sequence Modeling (N19-1)
Copied to clipboard
| Challenge: | Arabic text is written without short vowels (or diacritics) their presence is essential for properly verbalizing Arabic . |
| Approach: | They propose a character-level sequence-to-sequence deep learning model that recovers both types of diacritics without the use of explicit feature engineering. |
| Outcome: | The proposed model outperforms all previous state-of-the-art models on overlapping windows of words . it achieves a word error rate (WER) of 4.49% compared to the state- of-the art systems . |
The Hebrew Essay Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | Annotated corpus of argumentative essays authored by prospective higher-education students . corpus includes essays by native speakers and essays by non-native speakers . |
| Approach: | They propose to use an annotated corpus of Hebrew argumentative essays to analyze non-native language use. |
| Outcome: | The proposed corpus includes essays by native speakers and essays authored by non-native speakers with three different native languages. |
Better Together: Modern Methods Plus Traditional Thinking in NP Alignment (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study shows that end-to-end systems are not structurally free. |
| Approach: | They propose to use dictionary- and word vector-based baselines to align NPs in the bitext . they argue that alignment of NP's in MT can be improved by using old-fashioned methods . |
| Outcome: | a new study shows that alignment of NPs in the bitext is relevant even in an end-to-end paradigm . the proposed system can be improved by bringing in old-fashioned methods, the authors argue . |