Papers by Umut Sulubacak

2 papers
Normalizing Non-canonical Turkish Texts Using Machine Translation Approaches (P19-2)

Copied to clipboard

Challenge: a study using non-canonical text normalization shows that it can surpass the current best performing system by a large margin.
Approach: They propose a fully automated, context-aware machine translation approach with fewer stages of processing.
Outcome: The proposed approach surpasses the current best-performing system by a large margin . the proposed method is more data-hungry and more data sensitive than other methods .
OpusTools and Parallel Corpus Diagnostics (2020.lrec-1)

Copied to clipboard

Challenge: Currently OPUS contains 57 released corpora covering over 700 languages and language variants creating more than 70,000 bitexts in the sense of aligned language pairs across all corporata.
Approach: They introduce OpusTools, a package for downloading and processing parallel corpora in OPUS . the package implements tools for accessing compressed data in their archived release format . they show how they can be used in parallel corpus creation and data diagnostics .
Outcome: The proposed tools can be used in parallel corpus creation and data diagnostics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations