Papers by Amit Kumar

5 papers
Counter Turing Test (CT2): AI-Generated Text Detection is Not as Easy as You May Think - Introducing AI Detectability Index (ADI) (2023.emnlp-main)

Copied to clipboard

Challenge: a number of issues have arisen regarding the risk and consequences of AI-generated text detection.
Approach: They propose a counter-turing test to evaluate the robustness of existing AGTD methods . they propose ADI, a quantifiable spectrum to assess detectability of LLMs .
Outcome: The proposed method evaluates the robustness of existing AGTD methods . it shows that larger LLMs tend to have lower ADI, indicating they are less detectable .
Gated Transformer for Robust De-noised Sequence-to-Sequence Modelling (2021.findings-emnlp)

Copied to clipboard

Challenge: Noisy texts are common in user-generated texts that appear abundant in social media platforms like SMS, online chat, email, blogs, wikis etc.
Approach: They propose a sequence-to-sequence architecture that uses a gating mechanism to detect types of corrections required from English texts.
Outcome: The proposed architecture performs better than non-gated models on machine translation and Summarization tasks.
MVTamperBench: Evaluating Robustness of Vision-Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) have been a key advance in video understanding but their vulnerability to adversarial tampering remains underexplored.
Approach: They evaluate MLLMs against five prevalent tampering techniques to assess their robustness . they use a tampered video format to examine the vulnerability of ML models .
Outcome: The benchmark evaluates MLLMs against five prevalent tampering techniques based on 19 video manipulation tasks.
NLPRL at WAT2019: Transformer-based Tamil – English Indic Task Neural Machine Translation System (D19-52)

Copied to clipboard

Challenge: a majority of Asians speak low to medium resource languages . lack of resources poses a challenge, which requires innovative solutions .
Approach: They propose a Neural Machine Translation system for Tamil-English Indic Task . they train a system for both Tamil-to-English and English-to Tamil pairs .
Outcome: The proposed system is based on a Transformer-based architecture and is not very innovative, but can be treated as an incremental step in this direction.
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly being used for communication tasks across different regions.
Approach: They propose a benchmark to evaluate whether Large Language Models are ethically aligned and can be used in real-world situations.
Outcome: The proposed benchmark evaluates whether LLMs comply with or resist swearing instructions and assesses their alignment with ethical frameworks, cultural nuances, and language comprehension capabilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations