Towards the First Machine Translation System for Sumerian Transliterations (2020.coling-main)
Copied to clipboard
| Challenge: | Sumerian cuneiform script was invented more than 5,000 years ago and is one of the oldest in history. |
| Approach: | They propose to translate Sumerian texts into English automatically using supervised, phrase-based, and transfer learning techniques. |
| Outcome: | The proposed method accelerates the costly and time-consuming manual translation process and helps researchers better explore the relationships between Sumerian and Mesopotamian culture. |
Similar Papers
A Repository of Corpora for Summarization (L18-1)
Copied to clipboard
| Challenge: | Summarization corpora are numerous but fragmented, making it difficult to pinpoint corporata best suited for a given summarization task. |
| Approach: | They propose a repository containing corpora available to train and evaluate automatic summarization systems. |
| Outcome: | The proposed system is based on a repository of corpora available for summarization tasks. |
How Low is Too Low? A Computational Perspective on Extremely Low-Resource Languages (2021.acl-srw)
Copied to clipboard
| Challenge: | Sumerian is one of the world’s oldest written languages attested from at least the beginning of the 3rd millennium BC. |
| Approach: | They propose to use interpretLR to train attention-based deep learning models in a low-resource language, Sumerian cuneiform, which includes part-of-speech tagging, named entity recognition, and machine translation. |
| Outcome: | The proposed pipeline outperforms the large language model RoBERTa for POS Tagging and NER. |
The WAW Corpus: The First Corpus of Interpreted Speeches and their Translations for English and Arabic (L18-1)
Copied to clipboard
| Challenge: | Using the corpus, we study the characteristics of interpreters' work and train machine translation systems. |
| Approach: | They propose to build an interpreting corpus for Arabic and an Arabic corpus to study interpreters' work. |
| Outcome: | The proposed corpus can be used for teaching interpreters and to train machine translation systems. |
At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging (2024.lrec-main)
Copied to clipboard
| Challenge: | cuneiform texts are dominated by multiple languages and language families . the most dominant language written in cuniform is the Semitic Akkadian . existing cnl models are not suitable for digital editions of Akkadi . |
| Approach: | They focus on letters written in the Semitic Akkadian, a cuneiform language dominated by cuniform texts . they propose to use pre-trained embeddings, sentence segmentation and cnl to fine-tune language models . |
| Outcome: | The dominant language written in cuneiform is the Semitic Akkadian . the paper examines the input material and tries to initiate a discussion about best-practices . |
A Large-Scale Comparison of Historical Text Normalization Systems (N19-1)
Copied to clipboard
| Challenge: | a large study of historical text normalization is done on eight languages . there is no consensus on the state-of-the-art approach to normalization . |
| Approach: | They present a large study of historical text normalization done on eight languages . they evaluate four different systems based on supervised learning on datasets from eight different languages based in the literature . |
| Outcome: | The proposed methods are based on supervised learning and are available online. |
Automated Phonological Transcription of Akkadian Cuneiform Text (2020.lrec-1)
Copied to clipboard
| Challenge: | Akkadian was an east-semitic language spoken in ancient Mesopotamia . cuneiform text does not mark the inflection for logograms, so the inflected form needs to be inferred from the sentence context. |
| Approach: | They propose to automate phonological transcription of the transliterated Akkadian corpora . transcription is normalized according to the grammatical description of a given dialect . they find that cuneiform text does not mark the inflection for logograms . |
| Outcome: | The proposed transcriptions show the Akkadian renderings for Sumerian logograms, while the logogram transcription is more challenging. |
The DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World’s Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Until recently, language descriptions were available in paper form only, with indexes as the only search aid. |
| Approach: | They propose to digitize a multilingual corpus of language descriptions and annotate it with various meta, word, and text attributes to make searching and analysis easier and more useful. |
| Outcome: | The proposed corpus is searchable through a couple of well-established corpus infrastructures. |
A Modular Tool for Automatic Summarization (P19-3)
Copied to clipboard
| Challenge: | Abstractive automatic summarization methods are supervized, but they require large corpora to perform tasks. |
| Approach: | They propose to use a modular tool for automatic summarization that is as simple as possible for end-users. |
| Outcome: | The proposed tool is open source and written in Java . it could be used as a baseline for future work and evaluate methods on different corpora. |
From FreEM to D’AlemBERT: a Large Corpus and a Language Model for Early Modern French (2022.lrec-1)
Copied to clipboard
Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot
| Challenge: | Anguage models for historical states of language are becoming more complex to process and more scarce in the corpora available. |
| Approach: | They propose to use a contextualised language model to analyse historical states of language in French. |
| Outcome: | The proposed model is based on a corpus of historical texts and is evaluated with an NLP task. |
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)
Copied to clipboard
| Challenge: | We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script. |
| Approach: | They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic. |
| Outcome: | The proposed training recipe improves multilingual transliteration for Indic languages. |