Open Vocabulary Learning for Neural Chinese Pinyin IME (P19-1)

Copied to clipboard

Challenge: Pinyin-to-character conversion is the core component of pinyin based Chinese input method engine (IME).
Approach: They propose a neural P2C conversion model augmented by an online updated vocabulary to support open vocabulary learning during IME working.
Outcome: The proposed model outperforms commercial IMEs and state-of-the-art models on standard corpus and true inputting history dataset in terms of multiple metrics and the online updated vocabulary helps it follow user inputting behavior.

Similar Papers

Chinese Pinyin Aided IME, Input What You Have Not Keystroked Yet (D18-1)

Copied to clipboard

Challenge: Chinese pinyin input method engine (IME) converts pinyine into character based on its core component, pinyan-to-character conversion (P2C).
Approach: They propose a sequence-to-sequence model with gated-attention mechanism for Chinese IMEs.
Outcome: The proposed model improves on existing models in benchmark datasets showing great user experience improvement compared to traditional models.
Moon IME: Neural-based Chinese Pinyin Aided Input Method with Customizable Association (P18-4)

Copied to clipboard

Challenge: a pinyin input method engine (IME) allows users to input Chinese into a computer by typing pinyan through the common keyboard.
Approach: They present a pinyin IME that integrates neural machine translation and IR to offer amusive and customizable association ability.
Outcome: The Moon IME integrates neural machine translation and IR to offer amusive association ability.
Enabling Real-time Neural IME with Incremental Vocabulary Selection (N19-2)

Copied to clipboard

Challenge: Input method editor (IME) converts sequential alphabet key inputs to words in a target language.
Approach: They propose a neural-based language model that incrementally builds a subset vocabulary from the word lattice.
Outcome: The proposed approach achieves 50x speedup on Japanese IME benchmark without losing conversion accuracy.
Exploring Conditional Variational Mechanism to Pinyin Input Method for Addressing One-to-Many Mappings in Low-Resource Scenarios (2024.acl-short)

Copied to clipboard

Challenge: Experimental results demonstrate the superior performance of our method.
Approach: They propose to leverage conditional variational mechanism to simplify pinyin IME . they employ a strategy that facilitates interaction between pinyan and Chinese character information .
Outcome: The proposed method improves the performance of pinyin input method engine (IME) under low-resource conditions.
Correcting Chinese Word Usage Errors for Learning Chinese as a Second Language (C18-1)

Copied to clipboard

Challenge: a word usage error is the most common error type in Chinese, according to the HSK dynamic composition corpus . a system that considers both target erroneous token and context can generate a correction vector .
Approach: They propose a neural network model that considers target erroneous token and context to generate a correction vector and compare it against a candidate vocabulary to propose suitable corrections.
Outcome: The proposed model can detect 91% of the cases and propose suitable corrections within a list of five candidates.
Exploring and Adapting Chinese GPT to Pinyin Input Method (2022.acl-long)

Copied to clipboard

Challenge: a frozen GPT can generate state-of-the-art performance on perfect pinyin, but performance drops when input includes abbreviated pinyan, which links to even larger number of Chinese characters.
Approach: They propose to use Chinese GPT to generate fluent sentences using abbreviated pinyin.
Outcome: The proposed approach improves on abbreviated pinyin across all domains.
State-of-the-art Chinese Word Segmentation with Bi-LSTMs (D18-1)

Copied to clipboard

Challenge: A wide variety of neural-network architectures have been proposed for the task of Chinese word segmentation.
Approach: They propose a bidirectional LSTM model with standard deep learning techniques and best practices for the task of Chinese word segmentation.
Outcome: The proposed model outperforms models based on standard deep learning techniques and best practices on Chinese word segmentation datasets.
Self-Vocabularizing Training for Neural Machine Translation (2025.naacl-srw)

Copied to clipboard

Challenge: Past vocabulary learning techniques identify relevant vocabulary before training, relying on corpus statistics or frequency counts without considering contextual information or the model's ability to represent it.
Approach: They propose a method that self-vocabularizes a smaller, more optimal vocabulary by pairing source sentences with the model's predictions to define a new vocabulary.
Outcome: The proposed method produces a 1.49 BLEU improvement in the simulated model and an increase in unique token usage and a 6–8% reduction in vocabulary size.
VCWE: Visual Character-Enhanced Word Embeddings (N19-1)

Copied to clipboard

Challenge: Currently, word embeddings are playing a pivotal role in many natural language processing tasks.
Approach: They propose a model to learn Chinese word embeddings via three-level composition . they use convolutional neural network to extract intra-character compositionality from character shape .
Outcome: The proposed model performs better on word similarity, sentiment analysis, named entity recognition and part-of-speech tagging tasks.
Revisiting Pre-Trained Models for Chinese Natural Language Processing (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing pre-trained language models have shown tremendous improvements across various NLP tasks.
Approach: They propose to revisit Chinese pre-trained language models to examine their effectiveness in a non-English language and release the Chinese pretrained model series to the community.
Outcome: The proposed model improves on RoBERTa in several ways, especially the masking strategy that adopts MLM as correction (Mac).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations