Papers by Fahim Faisal

5 papers
Phylogeny-Inspired Adaptation of Multilingual Models to New Languages (2022.aacl-main)

Copied to clipboard

Challenge: Large pretrained multilingual models have delivered promising results due to cross-lingual learning capabilities on a variety of language tasks.
Approach: They propose to use language phylogenetic information to improve cross-lingual transfer by leveraging closely related languages in a structured, linguistically-informed manner.
Outcome: The proposed model significantly improves on the baseline model on languages unseen during training.
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)

Copied to clipboard

Challenge: linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages .
Approach: They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions .
Outcome: The proposed model is representative of the language diversity and coverage of natural language processing systems.
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)

Copied to clipboard

Challenge: despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages.
Approach: They propose to use global benchmarks to track progress on all NLP datasets in all languages.
Outcome: a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench.
SD-QA: Spoken Dialectal Question Answering for the Real World (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing QA benchmarks do not account for errors that speech recognition models might introduce . evaluating production-ready QA systems on data that is not representative of real-world inputs is problematic .
Approach: They construct a multi-dialect, spoken QA benchmark on five languages with 68k audio prompts in 24 dialects from 255 speakers.
Outcome: The proposed model is based on 68k audio prompts in 24 dialects from 255 speakers.
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties (2025.findings-emnlp)

Copied to clipboard

Challenge: toxicity detection of modern LLMs is underexplored due to dialectal differences.
Approach: They evaluate toxicity detection by using LLMs as evaluators across diverse dialects . they create a multi-dialect dataset using synthetic transformations and human-assisted translations based on human-aided translations.
Outcome: The proposed model shows that LLMs are sensitive to dialectal shifts and low-resource multilingual variation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations