Papers by Ryuichi Kiryo
Language-Independent Tokenisation Rivals Language-Specific Tokenisation for Word Similarity Prediction (2020.lrec-1)
Copied to clipboard
| Challenge: | Language-independent tokenisation (LIT) methods that do not require labelled language resources or lexicons have gained popularity because of their compactness and ability to handle unseen or rare words. |
| Approach: | They empirically compare language-independent tokenisation methods with language-specific tokenisation (LST) methods using carefully created lexicons and training resources. |
| Outcome: | The proposed methods outperform LIT and LST on evaluation tasks across eight languages. |
I Wish I Would Have Loved This One, But I Didn’t – A Multilingual Dataset for Counterfactual Detection in Product Review (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using machine translation, counterfactual statements are often found in natural languages. |
| Approach: | They annotate a multilingual CFD dataset from Amazon product reviews covering counterfactuals written in English, German, and Japanese languages. |
| Outcome: | The proposed dataset is robust against selection biases due to cue phrase-based sentence selection. |