Papers by Shu Okabe

7 papers
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data (2026.acl-long)

Copied to clipboard

Pedro Ortiz Suarez, Laurie Burchell, Catherine Arnett, Rafael Mosquera, Sara Hincapié Monsalve, Thom Vaughan, Damian Stewart, Malte Ostendorff, Idris Abdulmumin, Vukosi Marivate, Shamsuddeen Hassan Muhammad, Atnafu Lambebo Tonja, Hend Al-Khalifa, Nadia Ghezaiel Hammouda, Verrah Akinyi Otiende, Tack Hwa Wong, Jakhongir Saydaliev, Melika Nobakhtian, Muhammad Ravi Shulthan Habibi, Chalamalasetti Kranti, Carol Muchemi, Khang Nguyen, Faisal Muhammad Adam, Luis Frentzen Salim, Reem Alqifari, Cynthia Jayne Amol, Joseph Marvin Imperial, Ilker Kesen, Ahmad Mustafid, Pavel Stepachev, Leshem Choshen, David Anugraha, Hamada Nayel, Seid Muhie Yimam, Vallerie Alexandra Putra, My Chiffon Nguyen, Azmine Toushik Wasi, Gouthami Vadithya, Rob Van Der Goot, Lanwenn ar C’horr, Karan Dua, Andrew Yates, Mithil Bangera, Yeshil Bangera, Hitesh Laxmichand Patel, Shu Okabe, Fenal Ashokbhai Ilasariya, Dmitry Gaynullin, Genta Indra Winata, Yiyuan Li, Juan Pablo Martínez, Amit Agarwal, Ikhlasul Akmal Hanif, Raia Abu Ahmad, Esther Adenuga, Filbert Aurelian Tjiaranata, Weerayut Buaphet, Michael Anugraha, Sowmya Vajjala, Benjamin L Rice, Azril Hafizi Amirudin, Jesujoba Oluwadara Alabi, Srikant Panda, Yassine Toughrai, Bruhan Kyomuhendo, Daniel Ruffinelli, null Akshata, Manuel Goulão, Ej Zhou, Ingrid Gabriela Franco Ramirez, Cristina Aggazzotti, Konstantin Dobler, Jun Kevin, Quentin Pagès, Nicholas Andrews, Nuhu Ibrahim, Mattes Ruckdeschel, Amr Keleg, Mike Zhang, Casper Rufaro Muziri, Saron Samuel, Sotaro Takeshita, Kun Kerdthaisong, Luca Foppiano, Rasul Dent, Tommaso Green, Ahmad Mustapha Wali, Kamohelo Makaaka, Vicky Feliren, Inshirah Idris, Hande Celikkanat, Abdulhamid Abubakar, Jean Maillard, Benoît Sagot, Thibault Clérice, Kenton Murray, Sarah K. K. Luger
Challenge: Language identification (LID) is a fundamental step in curating multilingual corpora.
Approach: They introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages.
Outcome: The proposed benchmark covers 109 languages and shows that existing evaluations overestimate accuracy for many languages in the web domain.
Weakly Supervised Word Segmentation for Computational Language Documentation (2022.acl-long)

Copied to clipboard

Challenge: a recent paper aims to improve the effectiveness of unsupervised language analysis techniques in low resource settings.
Approach: They propose to use a weak supervision to improve linguistic segmentation in low resource languages . they propose to provide linguists with LTs that can be used to create interactive annotation tools .
Outcome: The proposed models can be used to improve the quality of language segmentation in low resource languages.
Improving Parallel Sentence Mining for Low-Resource and Endangered Languages (2025.acl-short)

Copied to clipboard

Challenge: Parallel sentence mining is a technique used to find matching sentence pairs from a source and target language.
Approach: They propose a benchmark dataset for parallel sentence mining on three low-resource languages . they apply alignment post-processing and cluster-based isotropy enhancement techniques to one of them .
Outcome: The proposed datasets show better mining quality overall for low-resource languages . the proposed methods are crucial for optimizing parallel data extraction for low resource languages - a new study shows.
Joint Word and Morpheme Segmentation with Bayesian Non-Parametric Models (2023.findings-eacl)

Copied to clipboard

Challenge: Language documentation often requires segmenting transcriptions of utterances into words and morphemes . a long tradition of nonparametric Bayesian models is used to handle these tasks .
Approach: They propose a Bayesian model for simultaneously segmenting utterances at two levels . they use two under-resourced languages to better understand the value of weak supervision .
Outcome: The proposed model can be used to identify language documents with weak supervision.
Multimodal Quality Estimation for Machine Translation (2020.acl-main)

Copied to clipboard

Challenge: Existing work has only explored textual context.
Approach: They propose to use visual and text modalities to explore Quality Estimation for Machine Translation and integrate them into multimodal QE frameworks.
Outcome: The proposed approaches improve on sentence-level and document-level predictions using visual features extracted from images.
Optical Character Recognition for the International Phonetic Alphabet (2026.eacl-short)

Copied to clipboard

Challenge: Grammar books are increasingly used as additional reference resources for low-resource languages . a significant portion of these documents come from scans and require an OCR tool .
Approach: They compare two neural OCR frameworks and a large vision-language model with a synthetic dataset based on Wiktionary to study the International Phonetic Alphabet (IPA).
Outcome: The proposed model improves on the International Phonetic Alphabet (IPA) character set.
Towards Multilingual Interlinear Morphological Glossing (2023.findings-emnlp)

Copied to clipboard

Challenge: Interlinear Morphological Glosses are annotations produced in the context of language documentation.
Approach: They propose to use a conditional random field to label morphs in L1 and then align them to L2 words to facilitate the process.
Outcome: The proposed method outperforms baselines in several under-resourced languages and is effective and data-efficient.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations