Papers by Jaydeep Sen

11 papers
Benchmarking and Building Zero-Shot Hindi Retrieval Model with Hindi-BEIR and NLLB-E5 (2025.naacl-long)

Copied to clipboard

Challenge: Existing benchmarks for evaluating retrieval models in Hindi are lacking . despite efforts to build multilingual retrieval systems, this is still a work in progress .
Approach: They evaluate Hindi retrieval models on the Hindi-BEIR benchmark and introduce a multilingual model that leverages a zero-shot approach to support Hindi without the need for Hindi training data.
Outcome: The proposed model leverages a zero-shot approach to support Hindi without the need for Hindi training data.
MILU: A Multi-task Indic Language Understanding Benchmark (2025.naacl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on English, leaving substantial gaps in assessing LLM capabilities in low-resource and linguistically diverse languages.
Approach: They propose a multi-task indic language understanding benchmark to assess LLMs in low-resource languages.
Outcome: The new benchmark spans 8 domains and 41 subjects across 11 Indic languages, reflecting general and culturally specific knowledge.
AIT-QA: Question Answering Dataset over Complex Tables in the Airline Industry (2022.naacl-industry)

Copied to clipboard

Challenge: Table Question Answering (Table QA) systems have been shown to be highly accurate when trained and tested on open-domain datasets built on top of Wikipedia tables.
Approach: They propose a domain-specific Table QA test dataset to test Table Question Answering systems on open-domain datasets built on top of Wikipedia tables.
Outcome: The proposed methods are highly accurate when tested on open-domain datasets built on top of Wikipedia tables.
Multi-Row, Multi-Span Distant Supervision For Table+Text Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Existing question answering systems for tables and linked text are relatively unexplored.
Approach: They propose a transformer-based question answering system that copes with distant supervision along both axes of the question and answer.
Outcome: The proposed system beats baselines for HybridQA and OTT-QA with best EM and F1 scores on a held out test set.
Topic Transferable Table Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Weakly-supervised table question-answering (TableQA) models have achieved state-of-art performance by using pre-trained BERT transformer to jointly encoding a question and a table to produce structured query for the question.
Approach: They propose a framework for TableQA that incorporates topic-specific vocabulary injection into BERT, a novel text-to-text transformer generator and a logical form re-ranker.
Outcome: The proposed framework provides a reasonably good baseline for topic shift benchmarks.
PrimeQA: The Prime Repository for State-of-the-Art Multilingual Question Answering Research and Development (2023.acl-demo)

Copied to clipboard

Challenge: Question Answering (QA) is a major area of research in Natural Language Processing (NLP)
Approach: They propose a one-stop and open-source QA repository for question answering . it supports core QA functionalities like retrieval and reading comprehension . they say it will facilitate easy replication of state-of-the-art (SOTA) QA methods .
Outcome: The proposed framework enables easy replication of state-of-the-art (SOTA) QA methods.
ODASim: Ordered, Distinctive and Absolute Semantic Similarity for Code Explanation Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for code explanations fail to distinguish correct from partially or fully incorrect explanations and their similarity scores are poorly calibrated.
Approach: They propose a model-agnostic graded fine-tuning framework that learns calibrated similarity representations between code and explanations to support fine-grained supervision and evaluation.
Outcome: The proposed framework improves F1 score and ECE scores on two embedding models and reduces expected calibration error.
UR2N: Unified Retriever and ReraNker (2025.coling-industry)

Copied to clipboard

Challenge: XTR-style retrieval on top of trained Mono-T5 reranker is suboptimal for two-stage retrieval, arguing that it is sub-optimal.
Approach: They propose a unified encoder-decoder architecture with a novel training regimen which enables the encoder representation to be used for retrieval and the decoder for re-ranking within a single unified model.
Outcome: The proposed architecture outperforms ColBERT, XTR, and even serves as a superior reranker compared to the Mono-T5 re-ranker.
XTR meets ColBERTv2: Adding ColBERTv2 Optimizations to XTR (2025.coling-industry)

Copied to clipboard

Challenge: XTR eliminates the need for multi-stage retrieval, but doesn't incorporate efficiency optimizations from ColBERTv2 which improve indexing and retrieval speed.
Approach: They propose a multi-vector retrieval method that simplifies retrieval into a single stage through a modified learning objective.
Outcome: The proposed method eliminates the need for multistage retrieval but doesn't incorporate efficiency optimizations from ColBERTv2 which improve indexing and retrieval speed.
Schema Aware Semantic Reasoning for Interpreting Natural Language Queries in Enterprise Settings (2020.coling-main)

Copied to clipboard

Challenge: Using ontology reasoning to understand natural language is a challenge for QA systems . a recent study shows that ontologies can improve natural language understanding .
Approach: They propose to use ontology reasoning to translate natural language interpretation into a sequence of solvable tasks by an ontologist.
Outcome: The proposed framework achieves better natural language understanding with a 30% accuracy improvement over the current state of natural language query interfaces.
INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages (2025.findings-naacl)

Copied to clipboard

Challenge: Large Language Models perform well on unseen tasks in English, but their abilities in non-English languages are less explored due to limited benchmarks and training data.
Approach: They propose to release a large dataset for context-grounded question answering in 11 major Indian languages.
Outcome: The Indic-QA Benchmark compared large datasets of large LLMs on extractive and abstractive tasks in 11 major Indian languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations