Papers by Anurag Das
More Images, More Problems? A Controlled Analysis of VLM Failure Modes. (2026.findings-acl)
Copied to clipboard
Anurag Das, Adrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Bernt Schiele, Georgios Tzimiropoulos, Brais Martinez
| Challenge: | Existing evaluations of large vision language models lack a comprehensive analysis of their weaknesses and causes. |
| Approach: | They propose a new benchmark to evaluate multi-image capabilities of Large Vision Language Models. |
| Outcome: | The proposed model outperforms existing benchmarks on multi-image models. |
Information Extraction from Visually Rich Documents using LLM-based Organization of Documents into Independent Textual Segments (2025.acl-long)
Copied to clipboard
| Challenge: | Specialized non-LLM NLP-based solutions lack reasoning and are not able to infer values not explicitly present in documents. |
| Approach: | They propose a novel LLM-based approach that organizes VRDs into localized semantic textual segments called semantic blocks. |
| Outcome: | The proposed approach outperforms the state-of-the-art on public VRD benchmarks by 1-3% in F1 scores and is resilient to document formats previously not encountered. |