Papers by Siddharth Jain

4 papers
A Large-scale Evaluation of Neural Machine Transliteration for Indic Languages (2021.eacl-main)

Copied to clipboard

Challenge: We analyze multilingual transliteration for Indic languages using scripts derived from the ancient Brahmi script.
Approach: They propose a multilingual training recipe for Indic languages that utilizes orthographic similarity between English and Indic.
Outcome: The proposed training recipe improves multilingual transliteration for Indic languages.
Learning to Speak and Act in a Fantasy Text Adventure Game (D19-1)

Copied to clipboard

Challenge: Existing studies on grounded dialogue use only statistical regularities of text data, without explicit understanding of the world that the text describes.
Approach: They propose a large-scale crowdsourced text adventure game as a research platform for studying grounded dialogue.
Outcome: The proposed game allows agents to perceive, emote, and act whilst conducting dialogue with other agents.
InfoSync: Information Synchronization across Multilingual Semi-structured Tables (2023.findings-acl)

Copied to clipboard

Challenge: Information Synchronization of semi-structured data across languages is challenging . culture differences, topic preferences, and editing inconsistency lead to information mismatches .
Approach: They propose a method for tabular synchronization that uses information from Wikipedia tables in one language with tables in another language.
Outcome: The proposed method achieves an acceptance rate of 77.28% on Wikipedia . english articles across the web are more timely updated than other languages .
Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)’s influential work on the New Yorker Cartoon Caption Contest.
Approach: They propose to decompose humor understanding into three components and improve each by enhancing visual understanding through improved annotation and utilizing LLM-generated humor reasoning and explanations.
Outcome: The proposed approach achieves 82.4% accuracy in caption ranking, significantly better than the previous 67% benchmark and matches the performance of world-renowned human experts in this domain.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations