Papers by Seokhwan Jang
Language, OCR, Form Independent (LOFI) pipeline for Industrial Document Information Extraction (2024.emnlp-industry)
Copied to clipboard
| Challenge: | Existing models for low-resource language (LRL) documents are limited in semantic entity extraction, and there are limitations in SER from word level results. |
| Approach: | They propose a pipeline for Document Information Extraction (DIE) in low-resource language (LRL) business documents that solves language, Optical Character Recognition (OCR), and form dependencies through flexible model architecture, a token-level box split algorithm, and the SPADE decoder. |
| Outcome: | Experiments on Korean and Japanese documents show that the pipeline performs well in the Semantic Entity Recognition task without pre-training. |