Papers by Fuad Rahman

6 papers
BornoDrishti: Leveraging Vision Encoders and Domain-Adaptive Learning for Bangla OCR on Diverse Documents (2026.eacl-industry)

Copied to clipboard

Challenge: Existing solutions for OCR for Bangla scripts are limited to single-domain processing.
Approach: They propose a unified OCR system that recognizes both printed and handwritten Bangla scripts within a single model.
Outcome: The proposed system achieves competitive accuracy across both domains and surpasses specialized uni-domain systems.
Adaptive Weighted Proxy Tuning: Efficient Gray-Box Steering for Image Captioning. (2026.acl-industry)

Copied to clipboard

Challenge: Proxy tuning is a decoding-time approach that fails to account for instance-specific variations in model certainty and domain shift.
Approach: They propose a gray-box steering framework that dynamically modulates the logit contributions of a large base model, a fine-tuned expert, and an untune .
Outcome: Adaptive Weighted Proxy Tuning achieves performance parity with fine-tuned models while remaining parameter-free.
BANMIME : Misogyny Detection with Metaphor Explanation on Bangla Memes (2025.emnlp-main)

Copied to clipboard

Challenge: Existing studies have explored hate speech and general meme classification, but the nuanced identification of misogyny in Bangla memes remains underexplored.
Approach: They propose a Bangla misogynistic meme dataset that includes misos, humor, metaphors and detailed human-written explanations.
Outcome: The proposed dataset is the first comprehensive dataset of misogynistic Bangla memes . it includes misos, humor categories, metaphor localization, and detailed human-written explanations based on 2,000 culturally grounded samples .
BanHADEX: Towards Explainable HAte Speech Detection in Bangla Using Human Annotated EXplanation (2026.acl-long)

Copied to clipboard

Challenge: Existing studies in Bangla focus on hate classification while overlooking interpretability.
Approach: They propose to create a dataset with human-annotated labels for banla that contains 19,203 YouTube comments spanning April 2024–June 2025.
Outcome: The proposed dataset outperforms existing datasets on open and closed-source LLMs on interpretability and better understanding of hate speech in linguistically rich yet under-resourced languages.
AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning for Small and Large Language Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models excel at complex reasoning tasks, yet their performance hinges on the quality of their prompts and pipeline structures.
Approach: They propose a framework that fully automates large language models' pipeline construction using reinforcement learning.
Outcome: Experimental results show that autoDSPy outperforms DSPy benchmarks in accuracy gains and time.
Gold Standard Bangla OCR Dataset: An In-Depth Look at Data Preprocessing and Annotation Processes (2023.emnlp-industry)

Copied to clipboard

Challenge: Existing datasets designed specifically for the Bengali language have been limited.
Approach: They propose to use a large collection of labeled Bangla text image datasets to improve the performance of Bangla OCR.
Outcome: The proposed system is the most extensive gold standard corpus for Bangla characters and words, comprising over 4 million human-annotated images.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations