Challenge: Existing datasets for incident management tasks are labor-intensive and time-consuming.
Approach: They propose a new IncidentAI dataset for safety prevention that includes three tasks . they argue that NLP techniques are beneficial for analyzing incident reports .
Outcome: The proposed dataset shows that NLP techniques are beneficial for analyzing incident reports to prevent future failures.

Similar Papers

AEGIS2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails (2025.naacl-long)

Copied to clipboard

Challenge: Existing safety-related content safety models are not well-suited for commercial use.
Approach: They propose a taxonomy that can be used to categorize safety risks . it combines human annotations with a multi-LLM "jury" system to assess safety . they plan to open-source Aegis2.0 data and models to aid in safety guardrailing .
Outcome: The proposed taxonomy can be used to assess the safety of human-LLM interactions . it can be trained on large, non-commercial datasets and is open-source .
OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset (2026.findings-acl)

Copied to clipboard

Challenge: Existing LLM safety datasets rely on ad-hoc taxonomies and lack rule-grounded, real-world cases.
Approach: They construct a rule-grounded, real-world case dataset OmniCompliance-100K from a compliance perspective using a powerful web-searching agent.
Outcome: The proposed dataset spans 74 regulations and policies across a wide range of domains including security and privacy regulations, content safety and user data privacy policies, financial security requirements, medical device risk management standards, educational integrity guidelines, and protections of fundamental human rights.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
A Chinese Dataset for Evaluating the Safeguards in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks.
Approach: They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples .
Outcome: The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset.
PHEE: A Dataset for Pharmacovigilance Event Extraction from Text (2022.emnlp-main)

Copied to clipboard

Challenge: Using NLP methods to discover and extract adverse drug events from unstructured textual data is difficult because it requires time-consuming manual curation.
Approach: They propose to use a hierarchical event schema to extract annotated events from medical case reports and biomedical literature to analyze patient data.
Outcome: The proposed dataset is the largest public dataset to date and contains over 5000 events from medical case reports and biomedical literature.
How to Contextualize Empirical Data for Risk Analysis with LLMs: A Case Study of Power Outages (2026.findings-eacl)

Copied to clipboard

Challenge: Large language models (LLMs) are increasingly being considered for high-stakes decision-making, yet their application in statistical risk analysis remains largely underexplored.
Approach: They propose a method for extracting key information from raw data and translating it into structured contextual input within the LLM prompt.
Outcome: The proposed approach significantly improves the LLM’s performance in risk assessment tasks.
Chinese SafetyQA: A Safety Short-form Factuality Benchmark for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models have created significant safety concerns . factuality ability is crucial in determining whether they can be deployed and applied safely and compliantly within specific regions.
Approach: They propose a benchmark to evaluate the factuality of large language models in China . they evaluate the models' ability to provide accurate and reliable information .
Outcome: The proposed benchmark evaluates the factuality abilities of existing LLMs and compares them to LLM abilities.
Semantic Annotation for Improved Safety in Construction Work (2020.lrec-1)

Copied to clipboard

Challenge: a number of documents provide evidence of previous incidents and mitigation strategies . but information about previous projects with similar attributes is often hidden within . a new named entity annotation scheme is being developed for construction safety .
Approach: a team of four health and safety experts have developed a named entity annotation scheme for construction safety documents.
Outcome: a new named entity annotation scheme annotates 600 sentences from accident reports . the scheme has an average agreement rate of 0.79 F-Score .
SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems (2022.acl-long)

Copied to clipboard

Challenge: Several studies discuss the potential harms and benefits of large language models (LLMs) large neural models can replicate and even amplify negative, stereotypical, and derogatory associations in the data.
Approach: They propose to use a first aid kit to assess the safety of conversational AI in various settings . they propose several future directions and discuss ethical considerations .
Outcome: The proposed tools can provide estimates of the relative safety of systems in various settings, but they still have several shortcomings.
Forecasting Future International Events: A Reliable Dataset for Text-Based Event Modeling (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches for text-based event prediction are limited in quality due to dynamic nature of international relations and conflicting economic dynamics.
Approach: They propose a novel dataset that leverages the advanced reasoning capabilities of large-language models to address these limitations.
Outcome: The proposed dataset features high-quality scoring labels generated through advanced prompt modeling and rigorously validated by domain experts in political science.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations