Papers by Nishanth Dikkala

2 papers
On the Benefits of Learning to Route in Mixture-of-Experts Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing Mixture-of-Expert (MoE) models allow us to scale up model sizes while keeping the amount of compute time fixed.
Approach: They propose to use a router to route inputs to experts in a layer to scale up model sizes while keeping the amount of compute time fixed.
Outcome: The proposed model scales up with the help of a router that routes input tokens to experts in a layer and shows that it is more efficient than a non-trainable router.
BIG-Bench Extra Hard (2025.acl-long)

Copied to clipboard

Challenge: Current benchmarks for large language model reasoning focus on math and coding abilities, leaving a gap in evaluating broader reasoning proficiencies.
Approach: They propose a benchmark to evaluate general reasoning in large language models . they use BIG-Bench and its harder version BIG-Benefit Hard to assess general reasoning .
Outcome: The new benchmark pushes the boundaries of LLM reasoning evaluation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations