Papers by Hiromitsu Nishizaki

5 papers
Presentation Slide Translation and Layout Error Correction by LLMs (2026.acl-srw)

Copied to clipboard

Challenge: Existing translation tools suffer from layout errors due to text expansion during translation . a new approach to translating Japanese slides into English is proposed to overcome this issue .
Approach: They propose a framework to translate Japanese slides into English and correct layout errors by using multimodal LLMs with slide images and XML structures.
Outcome: The proposed method outperforms baselines and achieves 4.1% layout error rate and over 80% success rate.
Improving Speech Recognition for the Elderly: A New Corpus of Elderly Japanese Speech and Investigation of Acoustic Modeling for Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: In an aging society, a highly accurate speech recognition system is needed for use in electronic devices for the elderly but this cannot be achieved using conventional speech recognition systems due to the unique features of the speech of elderly people.
Approach: They construct a new corpus of elderly Japanese speech from existing Japanese speech corpora and train them using existing data.
Outcome: The proposed models achieve word error rates (WER) as low as 13.38%, exceeding the results of the previous study.
Handwritten Character Generation using Y-Autoencoder for Character Recognition Model Training (2022.lrec-1)

Copied to clipboard

Challenge: re-emergence of deep learning since third winter of artificial intelligence has led to mainstreaming of deep-learning systems that use large amounts of data to train a model.
Approach: They propose a Y-Autoencoder-based handwritten character generator to generate Japanese Hiragana characters with a single image to increase the amount of data needed for character recognition.
Outcome: The proposed system generates Japanese Hiragana characters with a single image . the results show that the Y-AE-based generator produces an improved F1 score .
Semi-Automatic Construction and Refinement of an Annotated Corpus for a Deep Learning Framework for Emotion Classification (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for emotion classification are expensive and require a large corpus of data.
Approach: They propose a method for creating a semi-automatically constructed emotion corpus by correcting errors in the corpus.
Outcome: The proposed method improves the quality of the emotion labels by correcting errors.
Integrating Disfluency-based and Prosodic Features with Acoustics in Automatic Fluency Evaluation of Spontaneous Speech (2020.lrec-1)

Copied to clipboard

Challenge: acoustics, prosody, and disfluency-based features are used to evaluate fluent/disfluent speech . filling pauses and word fragments are used for automatic fluency evaluation .
Approach: They integrate acoustics, prosody, and disfluency-based features into an automatic fluency evaluation task.
Outcome: The proposed model improves when integrated with prosodic features, but not when disfluent speech is detected.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations