Papers by Muhammad Adilazuarda

5 papers
Towards Measuring and Modeling “Culture” in LLMs: A Survey (2024.emnlp-main)

Copied to clipboard

Challenge: Existing models are biased towards Western, Anglocentric or American cultures, a problem that is arguably detrimental to the performance of LLMs.
Approach: They analyze more than 90 recent papers that aim to study cultural representation and inclusion in large language models.
Outcome: The proposed models are biased towards Western, Anglocentric or American cultures, despite their diversity and their robustness.
NusaCrowd: Open Source Initiative for Indonesian NLP Resources (2023.findings-acl)

Copied to clipboard

Challenge: Existing NLP research in Indonesian languages has been held back by factors such as language diversity, orthographic variation, resource limitation and other societal challenges.
Approach: They present a collaborative initiative to collect and unify existing resources for Indonesian languages and open access to previously non-public resources.
Outcome: The results show that the datasets are highly reliable and can be used to generate the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia.
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization (2024.findings-emnlp)

Copied to clipboard

Challenge: Pretrained language models have shown remarkable generalization toward multiple tasks and languages, but their generalization towards unseen languages is poor.
Approach: They propose a regularization technique that incorporates various aspects of languages to better characterize linguistics constraints.
Outcome: The proposed technique improves accuracy of mBERT and XLM-R on unseen languages by 18% and 2% compared to fully finetuned models.
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting (2024.emnlp-main)

Copied to clipboard

Challenge: Socio-demographic prompting is a commonly employed approach to study cultural biases in LLMs as well as for aligning models to certain cultures.
Approach: They propose to use socio-demographic prompting to probe four LLMs with culturally sensitive and non-sensitive cues on datasets that are supposed to be culturally neutral or sensitive.
Outcome: The proposed model shows significant differences in responses on both kinds of datasets, casting doubt on its robustness.
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Southeast Asia (SEA) is home to over 1,300 indigenous languages and 671 million people . prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA .
Approach: They propose to provide a resource center that provides standardized corpora in nearly 1,000 SEA languages across three modalities.
Outcome: a new benchmark assesses the quality of AI models on 36 SEA languages across 13 tasks . the results highlight the importance of SEA as a culturally diverse region .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations