Papers by A.k.m Mahamud

1 papers
Gold Standard Bangla OCR Dataset: An In-Depth Look at Data Preprocessing and Annotation Processes (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing datasets designed specifically for the Bengali language have been limited.
Approach: They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR.
Outcome: The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations