Papers by Manohar Swaminathan

2 papers
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data (2024.emnlp-main)

Copied to clipboard

Challenge: Evaluation of multilingual Large Language Models is challenging due to a variety of factors including the lack of benchmarks with sufficient linguistic diversity, contamination of popular benchmarks into LLM pre-training data and lack of local, cultural nuances in translated benchmarks.
Approach: They evaluate 30 models across 10 Indic languages by conducting 90K human evaluations and 30K LLM-based evaluations.
Outcome: The proposed models perform best in most Indic languages, while the agreement drops for direct assessment especially for Bengali and Odia.
“#DisabledOnIndianTwitter” : A Dataset towards Understanding the Expression of People with Disabilities on Indian Twitter (2022.findings-aacl)

Copied to clipboard

Challenge: a majority of disabled Indians exist at the margins of society with little to no access to social media . as access to ICTs and high-speed internet grows, Indian Twitter's user base is expanding to include disability influencers, activists, and everyday disabled users.
Approach: They propose a hierarchical annotation taxonomy to classify tweets into various themes including discrimination, advocacy, and self-identification.
Outcome: The proposed taxonomy classifies 2,384 tweets into various themes including discrimination, advocacy, and self-identification.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations