Papers by Sheikh Shafayat

3 papers
BEnQA: A Question Answering Benchmark for Bengali and English (2024.findings-acl)

Copied to clipboard

Challenge: a dataset of parallel Bengali and English exam questions is used to compare LLMs in low-resource languages.
Approach: They introduce BEnQA, a dataset comprising parallel Bengali and English exam questions . they benchmark several Large Language Models with their parallel dataset and observe performance disparity .
Outcome: The proposed dataset consists of 5K questions covering several subjects in science . the authors find that the models perform poorly in Bengali and English .
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)

Copied to clipboard

Challenge: a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment.
Approach: They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation .
Outcome: The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks.
LangBridge: Multilingual Reasoning Without Multilingual Supervision (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to adapt language models for multilingual reasoning tasks require multilingual supervision.
Approach: They propose a zero-shot approach to adapt language models for multilingual reasoning tasks without multilingual supervision by bridging two models by introducing minimal trainable parameters between them.
Outcome: The proposed approach significantly improves multilingual reasoning capabilities on low-resource languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations