Papers by Rachneet Kaur

7 papers
LAW: Legal Agentic Workflows for Custody and Fund Services Contracts (2025.coling-industry)

Copied to clipboard

Challenge: Currently, there are limited resources available to build a legal domain-specific Large Language Model (LLM) however, legal contracts are highly varied not only in terms of semantics but also accessibility.
Approach: They propose a Large Language Model (LLM) that integrates multiple specialized agents and text agents to respond to user queries.
Outcome: The proposed model outperforms the baseline model in complex tasks such as calculating a contract’s termination date by 92.9% points.
MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A (2026.acl-industry)

Copied to clipboard

Challenge: Recent advances in multimodal retrieval-augmented generation (MM-RAG) have shifted toward minimal parsing, relying on page-level images for producing retriever embeddings and answer generation.
Approach: They propose a document structure-aware split that extracts and represents document structure via a structure-based split that dynamically routes documents through orientation-specific ingestion pipelines.
Outcome: The proposed model outperforms state-of-the-art vision-centric baselines by up to 32% points and achieves strong gains on report-style layouts.
SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding (2026.acl-long)

Copied to clipboard

Challenge: Multimodal large language models (MLLMs) are a promising tool for document understanding, but they are not able to handle complex multi-page visual documents.
Approach: They propose a flexible agentic framework for understanding multi-modal, multi-page, and multi-layout documents . SlideAgent employs specialized agents and decomposes reasoning into three specialized levels .
Outcome: a new agentic framework improves accuracy over open-source and proprietary models . it decomposes reasoning into three levels to capture themes and visual cues . the framework is based on a multimodal large language model and a MLLM .
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations (2025.acl-long)

Copied to clipboard

Challenge: State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs).
Approach: They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability.
Outcome: The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations.
Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are a critical tool for time series analysis and reporting in many fields, including healthcare, finance, climate, and many more.
Approach: They propose a framework for rigorously evaluating the capabilities of Large Language Models (LLMs) on time series understanding, encompassing both univariate and multivariate forms.
Outcome: The proposed framework delineates various characteristics inherent in time series data.
LETS-C: Leveraging Text Embedding for Time Series Classification (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in language modeling have shown promising results when applied to time series data.
Approach: They propose a method to fine-tune large language models for time series classification tasks using text embedding models and a simple classification head.
Outcome: The proposed model outperforms the current SOTA model on a time series classification benchmark and uses only 14.5% of the trainable parameters.
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering (2026.acl-long)

Copied to clipboard

Challenge: Recent multimodal LLMs have shown promise in chart-based visual question answering, but their performance declines sharply on unannotated charts.
Approach: They propose a novel agentic framework that explicitly performs visual reasoning directly within the chart’s spatial domain.
Outcome: The proposed framework achieves state-of-the-art accuracy on the ChartBench and ChartX benchmarks surpassing prior methods by up to 16.07% absolute gain overall and 17.31% on numerically intensive queries.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations