Challenge: Major scandals in corporate history have urged the need for regulatory compliance, where organizations need to ensure that their controls (processes) comply with relevant laws, regulations, and policies.
Approach: They introduce regulatory information retrieval (REG-IR) an application of document-to-document information retrievals where the query is an entire document making the task more challenging than traditional IR where the queries are short.
Outcome: The proposed approach is more challenging than traditional IR where the query is an entire document making the task more challenging.

Similar Papers

Regulation and NLP (RegNLP): Taming Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: polarization in AI safety and ethics debates are swaying political agendas on AI regulation and governance . regulation studies are rich source of knowledge on how to systematically deal with risk and uncertainty .
Approach: They argue that NLP research can benefit from proximity to regulatory studies . they argue that regulation studies should focus on linking scientific knowledge to regulatory processes .
Outcome: The proposed research space should focus on linking scientific knowledge to regulatory processes based on systematic methodologies.
From Complexity to Clarity: AI/NLP’s Role in Regulatory Compliance (2025.findings-acl)

Copied to clipboard

Challenge: Recent advances in natural language processing have demonstrated remarkable capabilities in text analysis and reasoning.
Approach: They propose to use standardized evaluation frameworks and balanced human-AI collaboration to address these challenges.
Outcome: The proposed research will focus on standardized evaluation frameworks and balanced human-AI collaboration to address these challenges.
Towards Automated Extraction of Business Constraints from Unstructured Regulatory Text (C18-2)

Copied to clipboard

Challenge: a system for machine-driven annotations of legal documents is currently undergoing user trials within our organization.
Approach: a system for machine-driven annotations of legal documents is presented . the system is currently undergoing user trials within our organization.
Outcome: the proposed system is currently undergoing user trials within our organization.
Proceedings of the Second Workshop on Economics and Natural Language Processing (D19-51)

Copied to clipboard

Challenge: ECONLP 2019 will focus on the many ways natural language processing influences business relations and procedures .
Approach: a talk will discuss use-cases of natural language processing to aid in regulatory workflows . a workshop will focus on the many ways how NLP influences business relations and procedures .
Outcome: This talk covers use-cases of natural language processing to aid in regulatory workflows . it also discusses shortcomings of current NLP technologies for financial regulation .
RegNLI: Detecting Online Product Misbranding through Legal and Linguistic Alignment (2026.eacl-industry)

Copied to clipboard

Challenge: Existing approaches to claim verification focus on keyword matching or generic text classification . misbranding involves deceptive labeling or advertising that misleads consumers about a product's nature or quality .
Approach: They propose a framework that formulates misbranding detection as an inference task between product claims and regulatory provisions.
Outcome: The proposed framework outperforms baselines in misbranding detection and regulation alignment metrics.
BERT-QE: Contextualized Query Expansion for Document Re-ranking (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to expand query use pseudo relevance feedback (PRF) but they are under-equipped to evaluate the relevance of information pieces used for expansion.
Approach: They propose a query expansion model that leverages the BERT model to select relevant document chunks for expansion.
Outcome: The proposed model significantly outperforms existing models on the TREC Robust04 and GOV2 test collections.
CLERC: A Dataset for U. S. Legal Case Retrieval and Retrieval-Augmented Analysis Generation (2025.findings-naacl)

Copied to clipboard

Challenge: a dataset of case law is used to train and evaluate models for writing legal analyses . current approaches struggle to find relevant cases and generate legal analyses, authors say .
Approach: They build a dataset of case law to support information retrieval and retrieval-augmented generation.
Outcome: The proposed dataset supports two important backbone tasks: retrieval (IR) and retrieval-augmented generation (RAG).
Extraction of Information Provision Activity Requirements from EU Acquis (2025.emnlp-industry)

Copied to clipboard

Challenge: Using knowledge-, classical ML-, transformer-, and generative AI-based approaches, we extract structured information from EU acquis documents.
Approach: They propose a task of Information Provision Activity Requirement Extraction to identify text fragments that introduce an obligation to provide information and the extraction of structured information about the key entities involved.
Outcome: The proposed task is based on knowledge-, classical ML-, transformer-, and generative AI-based approaches.
Populating Legal Ontologies using Semantic Role Labeling (2020.lrec-1)

Copied to clipboard

Challenge: This paper is concerned with the ‘resource consumption bottleneck’ of creating semantic technologies manually.
Approach: They propose to combine general-purpose NLP modules with pre- and post-processing using rules based on domain knowledge to solve the acquisition paradox.
Outcome: The proposed system extracts norms from legislation and represents them as structured norms in legal ontologies.
Towards Operationalizing Right to Data Protection (2025.naacl-long)

Copied to clipboard

Challenge: Recent work introduces the concept of generating unlearnable datasets (by adding imperceptible spurious correlations to the clean data) this approach is limited by several practical constraints like requiring knowledge of the target model.
Approach: They propose a framework that injects imperceptible spurious correlations into natural language datasets, rendering them unlearnable without affecting semantic content.
Outcome: The proposed framework can restrict newer models like GPT-4o and Llama from learning on generated data, resulting in a drop in test accuracy compared to their zero-shot performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations