Papers by A.k.m Mahamud
Gold Standard Bangla OCR Dataset: An In-Depth Look at Data Preprocessing and Annotation Processes (2023.emnlp-industry)
Copied to clipboard
| Challenge: | Existing datasets designed specifically for the Bengali language have been limited. |
| Approach: | They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR. |
| Outcome: | The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images. |