Papers by Mirna Adriani

    1 papers
    Normalization of Indonesian-English Code-Mixed Twitter Data (D19-55)

    Copied to clipboard

    Challenge: Twitter is an excellent source of textual data for NLP researches, but it is noisy and often contains typos, slang terms, and non-standard abbreviations.
    Approach: They propose a standardization system for Indonesian-English code-mixed Twitter data that includes tokenization, language identification, lexical normalization, and translation.
    Outcome: The proposed standardization system is based on four modules for tokenization, language identification, lexical normalization, and translation.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations