Papers by Sanjan Baitalik
Confidence as a Tie-Breaker: Reassessing Multilingual Hedging Bias in LLM-as-a-Judge Evaluation (2026.acl-srw)
Copied to clipboard
| Challenge: | LLM judges are often used to score generated answers, but their decisions may be affected by surface style rather than semantic correctness. |
| Approach: | They propose a benchmark to examine multilingual hedging effects in LLM evaluation . they find that when two answers are equally correct, the judge prefers the assertive answer . |
| Outcome: | The proposed benchmark shows that judge preferences are based on style and style . the results highlight that benchmark auditing is a key requirement for judge-bias research . |
Garden Path Recovery in Causal and Masked Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | a linguistics study of garden-path sentences shows that recovery dynamics are important for linguistic evaluation . causal models show larger within-model disambiguation effects than masked models overall . |
| Approach: | They propose to compare garden-path recovery in causal and masked language models . they use 100 English garden- path/control pairs spanning three constructions . |
| Outcome: | The proposed model shows that decoder-only models exhibit sharper disruption at the point of syntactic revision, while encoders appear comparatively buffered at the disambiguator due to right-context access. |