Papers by Sanjan Baitalik

2 papers
Confidence as a Tie-Breaker: Reassessing Multilingual Hedging Bias in LLM-as-a-Judge Evaluation (2026.acl-srw)

Copied to clipboard

Challenge: LLM judges are often used to score generated answers, but their decisions may be affected by surface style rather than semantic correctness.
Approach: They propose a benchmark to examine multilingual hedging effects in LLM evaluation . they find that when two answers are equally correct, the judge prefers the assertive answer .
Outcome: The proposed benchmark shows that judge preferences are based on style and style . the results highlight that benchmark auditing is a key requirement for judge-bias research .
Garden Path Recovery in Causal and Masked Language Models (2026.acl-srw)

Copied to clipboard

Challenge: a linguistics study of garden-path sentences shows that recovery dynamics are important for linguistic evaluation . causal models show larger within-model disambiguation effects than masked models overall .
Approach: They propose to compare garden-path recovery in causal and masked language models . they use 100 English garden- path/control pairs spanning three constructions .
Outcome: The proposed model shows that decoder-only models exhibit sharper disruption at the point of syntactic revision, while encoders appear comparatively buffered at the disambiguator due to right-context access.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations