Papers by Fahim Faisal
Phylogeny-Inspired Adaptation of Multilingual Models to New Languages (2022.aacl-main)
Copied to clipboard
| Challenge: | Large pretrained multilingual models have delivered promising results due to cross-lingual learning capabilities on a variety of language tasks. |
| Approach: | They propose to use language phylogenetic information to improve cross-lingual transfer by leveraging closely related languages in a structured, linguistically-informed manner. |
| Outcome: | The proposed model significantly improves on the baseline model on languages unseen during training. |
Dataset Geography: Mapping Language Data to Language Users (2022.acl-long)
Copied to clipboard
| Challenge: | linguistic diversity and coverage of natural language processing systems is a key factor in determining quality of data available in the language field . lack of linguistic, typological, and geographical diversity is acknowledged and documented . but, the advent of massively multilingual models presents opportunity and hope for under-represented languages . |
| Approach: | They analyze the geographical representativeness of NLP datasets to determine their utility . they also explore economic and geographical factors that may explain the observed distributions . |
| Outcome: | The proposed model is representative of the language diversity and coverage of natural language processing systems. |
GlobalBench: A Benchmark for Global Progress in Natural Language Processing (2023.emnlp-main)
Copied to clipboard
Yueqi Song, Simran Khanuja, Pengfei Liu, Fahim Faisal, Alissa Ostapenko, Genta Winata, Alham Aji, Samuel Cahyawijaya, Yulia Tsvetkov, Antonios Anastasopoulos, Graham Neubig
| Challenge: | despite advances in NLP, significant disparities in performance across languages still exist . prior benchmarks focused on a limited number of tasks and languages, but now GlobalBench tracks progress on all languages. |
| Approach: | They propose to use global benchmarks to track progress on all NLP datasets in all languages. |
| Outcome: | a new tool tracks progress on all NLP datasets in all languages and tracks per-speaker utility and equity . globalbench is designed to identify the most under-served languages and reward research efforts . a globalbech is available at https://github.com/neulab/globalbench. |
SD-QA: Spoken Dialectal Question Answering for the Real World (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing QA benchmarks do not account for errors that speech recognition models might introduce . evaluating production-ready QA systems on data that is not representative of real-world inputs is problematic . |
| Approach: | They construct a multi-dialect, spoken QA benchmark on five languages with 68k audio prompts in 24 dialects from 255 speakers. |
| Outcome: | The proposed model is based on 68k audio prompts in 24 dialects from 255 speakers. |
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties (2025.findings-emnlp)
Copied to clipboard
| Challenge: | toxicity detection of modern LLMs is underexplored due to dialectal differences. |
| Approach: | They evaluate toxicity detection by using LLMs as evaluators across diverse dialects . they create a multi-dialect dataset using synthetic transformations and human-assisted translations based on human-aided translations. |
| Outcome: | The proposed model shows that LLMs are sensitive to dialectal shifts and low-resource multilingual variation. |